Reactive Voice Device Management

JP7901424B2Active Publication Date: 2026-08-06INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2022-09-29
Publication Date
2026-08-06

Smart Images

  • Figure 0007901424000003
    Figure 0007901424000003
  • Figure 0007901424000004
    Figure 0007901424000004
  • Figure 0007901424000005
    Figure 0007901424000005
Patent Text Reader

Abstract

One or more user interactions directed to a set of one or more voice controlled devices in the environment are received by a first connected device. A first input to a first voice controlled device of the set of voice controlled devices is detected based on the user interactions. In response to the first input, a potential second input to the set of voice controlled devices is determined based on the activity model. In response to the first input, the user interactions are monitored for deviations from the potential second input. Based on the monitoring, anomalies in activity in the environment are identified. In response to the anomalies in activity, corrective action is performed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to voice-based devices, and more particularly, to the management of voice-based devices using artificial intelligence technology.

Background Art

[0002] Voice-based devices can operate based on an audio interface. Voice-based devices can receive commands from a user. Voice-based devices may operate based on only exact commands or specific commands.

Summary of the Invention

[0003] According to an embodiment, a method, a system, and a computer program product are disclosed.

[0004] One or more user interactions directed to a set of one or more voice control devices in an environment are received by a first connection device. A first input to a first voice control device of the set of voice control devices is detected based on the user interaction. A potential second input to the set of voice control devices is determined based on an activity model in response to the first input. A deviation from the potential second input is monitored from the user interaction in response to the first input. Based on the monitoring, an abnormality in the activity in the environment is identified. In response to the abnormality in the activity, a corrective action is executed.

[0005] The above summary is not intended to explain each illustrated embodiment or all embodiments of the present disclosure.

[0006] Hereinafter, preferred embodiments of the present invention will be described by way of example only, with reference to the following drawings.

Brief Description of the Drawings

[0007] [Figure 1]The following are representative key components of exemplary computer systems that may be used in some embodiments of this disclosure. [Figure 2] This document shows a cloud computing environment using one embodiment of the present invention. [Figure 3] An abstraction model layer according to one embodiment of the present invention is shown. [Figure 4] The following are representative neural network examples of one or more artificial neural networks capable of performing reactive speech device management ("RVM") on one or more speech devices in an environment consistent with the embodiments of this disclosure. [Figure 5] An exemplary system for voice device management, consistent with some embodiments of this disclosure, is shown. [Figure 6A] A first portion of training data for a first machine learning model of the system for identifying abnormal inputs is shown, consistent with some embodiments of the present disclosure. [Figure 6B] A second portion of the training data for a first machine learning model of the system, consistent with some embodiments of this disclosure, is shown. [Figure 7] The following is a portion of the training data for a second machine learning model of the system for identifying abnormal inputs, consistent with some embodiments of this disclosure. [Figure 8] The following is a portion of the training data for a third machine learning model of the system for identifying abnormal inputs, consistent with some embodiments of this disclosure. [Figure 9A] The first portion of training data for a fourth machine learning model of the system is shown to provide precise corrective behavior consistent with some embodiments of the present disclosure. [Figure 9B] A second portion of the training data for a fourth machine learning model of the system, consistent with some embodiments of this disclosure, is shown. [Figure 10] A method 1000 for performing a corrective action on an audio device, consistent with some embodiments of the present disclosure, is shown. [Modes for carrying out the invention]

[0008] While the present invention is subject to various modifications and alternative forms, specific embodiments are illustrated in the drawings and described in detail. However, it should be understood that the invention is not intended to be limited to the specific embodiments described. Rather, the intention is to cover all modifications, equivalents, and alternatives that fall within the scope of the invention.

[0009] Aspects of this disclosure relate to voice-based devices, and more specifically, to the management of voice-based devices using artificial intelligence technologies. While this disclosure is not necessarily limited to such uses, various aspects of this disclosure can be understood through discussions of various examples using this context.

[0010] A voice-controlled client device (alternatively, a voice-controlled device, or voice-based device) ("voice device") may be a computing device that operates based on voice input from a user, such as a voice. Voice devices may be becoming increasingly popular due to one or more factors. One factor is that voice devices may facilitate user use that may be perceived as intuitive. Users may be accustomed to speaking commands or questions to a computer because the voice device may be programmed to respond to similar commands (e.g., "What's the weather like today?"). Another factor is that voice devices may enable multitasking in real-world scenarios (e.g., a user may be able to operate the voice device while their hands are occupied with other tasks). For example, a user may be able to instruct the voice device to turn on the lights while walking their dog.

[0011] Another factor may be that the cost of computer components (such as memory, processors, and voice transceivers) has fallen to a level where voice devices are readily available or integrated into all kinds of devices. These voice devices include computer devices (such as laptops and desktops), portable electronic devices (such as smartphones and tablets), or wearable client devices (such as augmented reality glasses), or a combination thereof. Voice devices may also be consumer electronics (e.g., smart refrigerators, voice-controlled washing machines, and beverage machines with voice-based interfaces). In some cases, voice devices may include devices that operate solely on receiving voice or verbal input from the user (e.g., voice-based assistants that do not have a touchscreen, physical buttons, or other means of receiving input other than voice commands).

[0012] Voice devices may have limited functionality in certain scenarios, making them difficult or impossible to operate. Specifically, voice devices may only respond to rigid or fixed commands. For example, a voice device installed in a home office might respond to "What's the temperature?" but not to "What's the weather like now?" or "Is it raining here?" Another issue may be limited knowledge of how to respond. Limited knowledge may include the voice device responding to the user by playing a limited range of voice samples or overly general voice samples (e.g., an audio file saying "Invalid command," a sound wave saying "Something's wrong," or a visible on-screen message saying "I can't answer").

[0013] Existing solutions may not be suitable for all scenarios. One existing solution might involve connecting the device to a network, such as the internet, and performing additional processing (e.g., by a computer or technical support user). Such solutions could infringe on user privacy, such as monitoring voice input to a voice device in a home environment. Furthermore, constant monitoring may not be possible at certain times. For example, a user might travel to a remote location with a smartphone acting as a voice device. The smartphone may lack network connectivity, rendering network-based voice listening and processing completely ineffective.

[0014] Another existing solution might involve generating a vast number of command and response options in an attempt to cover every possible scenario with a voice device. Specifically, a voice device could contain dozens or even hundreds of commands that a user might potentially provide. This existing solution also has its drawbacks. For example, it might require a large memory or storage subsystem in the voice device to provide audio recordings of all potential audible questions and answers. Furthermore, because language is not fixed but constantly evolving, a relatively large dataset may not be able to cover future ways of communication.

[0015] Another drawback is that users are not perfect and cannot always operate voice devices in a predictable or understandable way. Specifically, a given user may forget to send a voice command even though the voice device expects one for the correct operation. Users may be busy with other activities and forget to speak a voice command. For example, a user may be putting tea in a smart microwave. The smart microwave may be a voice device configured to heat items after receiving a voice command. After placing the tea cup in the smart microwave appliance, the user may normally need to speak a specific command to start heating the tea. Also, a user may be actively caring for a child and may accidentally forget to speak a specific language command.

[0016] In the second example, the first user might be putting tea in a microwave oven at home. The first user might accidentally utter an unexpected, unwanted, or unusual command. For example, the user might be talking on the phone with someone else while the tea is heating up. The other person on the phone might ask the first user how long it will take to get ready to leave, and the first user might reply, "I'll be ready to leave in 45 minutes." The microwave oven might misinterpret this as 45 minutes to heat the tea.

[0017] Reactive Voice Device Management (RVM) can be operative to detect user input and responsively determine potential additional input for beneficial processing and operation (e.g., management) of voice devices. RVM can operate by receiving user interactions directed to one or more voice devices such as smartphones, voice-based assistants, voice-operated household appliances. Specifically, the voice device may be a computer, smart home appliance, voice-based assistant, or other relevant voice device in the user's environment (e.g., home, office, school). RVM may be configured to process user actions and commands through a machine learning model or other relevant artificial intelligence and execute respective corrective actions. [[ID=I]]

[0018] Specifically, RVM may be configured to determine a user's routine or usage pattern, such as by generating an activity model that takes into account various usage patterns of the user in the environment, not only the voice devices that are part of the usage pattern, based on machine learning. The routine or pattern may be based on previous usage of the voice device. Further, RVM can execute a corrective action based on detecting an input to the voice device. Specifically, RVM may be configured to detect an input to a voice device in the environment, and the input may be part of one or more user interactions received by the connected device. Based on a pre-generated activity model, in response to the detected input, RVM can determine a potential second input (the "second input"). The second input may be a future, upcoming, or predicted input corresponding to a routine or usage pattern in the environment. RVM may further monitor for additional input from the user when the user is performing an interaction with the voice device. RVM may monitor for deviations and, based on the monitoring, identify an anomaly (the "anomaly") in the activities in the environment.

[0019] The anomaly may be a lack of input from the user (e.g., the user does not speak the corresponding follow-up instruction, the user does not perform a specific follow-up action). The anomaly may be inaccurate input from the user (e.g., the user speaks an incorrect command, the user performs an incorrect physical action). The anomaly may also be input that does not match a pattern, routine, or sequence of inputs that is part of the activity model. The corrective action may include generating a response to the user that includes details of the missing or anomalous input. The generated response may be based not only on the anomaly in the activity, but also on patterns, routines, and a second input of the determined user.

[0020] The RVM may operate to overcome one or more problems of existing voice control devices. First, the RVM can be configured to operate by a connected device. The connected device may be a cloud-connected device such as a server for processing user interactions. In some embodiments, the connected device may be a dedicated device such as a computer configured to execute processing for the RVM. In some embodiments, the connected device may be one of the voice devices also configured to receive voice input from the user and respond with a voice output. By operating locally, the RVM can enable local processing of voice data without connecting to an external server.

[0021] Another advantage is that the RVM may operate without a fixed set of commands and responses. Specifically, the RVM can execute one or more artificial intelligence techniques (e.g., machine learning) to determine routines, activities, or patterns. Due to machine learning, the RVM may be able to detect specific anomalies regarding voice user interactions, even if a particular interaction was not previously part of the anomaly list. Further, the RVM may be able to generate a language response to the user, even if words or phrases were not initially part of the vocabulary stored in the device.

[0022] Furthermore, RVM is also effective when users continue to use voice devices in more complex ways. Specifically, as a user uses a voice device in a given environment, they may begin to use the device in more elaborate ways or in conjunction with other voice devices. For example, if a user starts using a voice-enabled coffee maker in the morning, the RVM may receive training data including the coffee maker's time and settings. As the user consistently wakes up at the same time and performs the same actions, the RVM can update its activity model with this data. Subsequently (e.g., days, weeks, months), the user may also install a smart lamp in the same room as the voice-enabled coffee maker and begin using voice commands or physical input. The usage of the smart lamp may also be input to the RVM, and the activity model may include usage information for both devices. Subsequently, when the user performs a user interaction directed at either the smart lamp or the voice-enabled coffee maker, each device may be monitored for a potential second input. Monitoring for potential inputs would involve monitoring deviations from patterns created in the activity models corresponding to both voice devices. As a result, RVM can compare user interaction inputs to a trained activity model, and RVM can identify anomalies not only in one of the voice devices, but also in either a smart lamp, a voice-activated coffee maker, or both.

[0023] Figure 1 shows typical key components of an exemplary computer system 100 (alternatively, a computer) that may be used in some embodiments of the present disclosure. It is understood that the individual components may differ in complexity, number, type, configuration, or combination thereof. The specific examples disclosed are illustrative and not necessarily limited to such variations. The computer system 100 may include a processor 110, memory 120, input / output interfaces (here I / O or I / O interfaces) 130, and a main bus 140. The main bus 140 can provide a communication path to other components of the computer system 100. In some embodiments, the main bus 140 may be connected to other components, such as a dedicated digital signal processor (not shown).

[0024] The processor 110 of the computer system 100 may consist of one or more cores 112A, 112B, 112C, 112D (collectively referred to as 112). The processor 110 may further include one or more memory buffers or caches (not shown) that provide temporary storage of instructions and data for the cores 112. The cores 112 can execute instructions for inputs provided from the cache or memory 120 and output the results to the cache or memory. The cores 112 may consist of one or more circuits configured to perform one or more methods consistent with embodiments of the present disclosure. In some embodiments, the computer system 100 may include multiple processors 110. In some embodiments, the computer system 100 may be a single processor 110 having a single core 112.

[0025] The memory 120 of the computer system 100 may include a memory controller 122. In some embodiments, the memory 120 may include random-access semiconductor memory, storage devices, or storage media (either volatile or non-volatile) for storing data and programs. In some embodiments, the memory may be in the form of modules (e.g., dual in-line memory modules). The memory controller 122 communicates with the processor 110 to facilitate the storage and retrieval of information in the memory 120. The memory controller 122 also communicates with the I / O interface 130 to facilitate the storage and retrieval of inputs or outputs in the memory 120.

[0026] The I / O interface 130 may include an I / O bus 150, a terminal interface 152, a storage interface 154, an I / O device interface 156, and a network interface 158. The I / O interface 130 may connect the main bus 140 to the I / O bus 150. The I / O interface 130 can direct instructions and data from the processor 110 and memory 120 to various interfaces of the I / O bus 150. Alternatively, the I / O interface 130 may direct instructions and data from various interfaces of the I / O bus 150 to the processor 110 and memory 120. These various interfaces may include the terminal interface 152, the storage interface 154, the I / O device interface 156, and the network interface 158. In some embodiments, the various interfaces may include a subset of the aforementioned interfaces (for example, an embedded computer system for an industrial application may not include the terminal interface 152 and the storage interface 154).

[0027] The logic modules of the entire computer system 100 (including, but not limited to, memory 120, processor 110, and I / O interface 130) can communicate failures and changes to one or more components to a hypervisor or operating system (not shown). The hypervisor or operating system can allocate various resources available in the computer system 100 and track the location of data in memory 120 and processes assigned to various cores 112. In embodiments where elements are combined or rearranged, the aspects and capabilities of the logic modules may be combined or redistributed. These variations will be apparent to those skilled in the art.

[0028] This disclosure includes a detailed description of cloud computing, but the implementations of the teachings described herein are not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented with any other type of computer environment that is currently known or may be developed in the future. Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or minimal interaction with service providers. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.

[0029] The characteristics are as follows:

[0030] On-demand self-service: Cloud consumers can unilaterally prepare computing power, such as server time and network storage, automatically as needed, without requiring human interaction with service providers.

[0031] Broad network access: Computing power is available over the network and accessible through standard mechanisms. This facilitates utilization by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, PDAs).

[0032] Resource pooling: A provider's computing resources are pooled and delivered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated as needed. Generally, consumers have a sense of location independence because they do not manage or know the exact location of the resources provided. However, consumers may be able to identify the location at a higher level of abstraction (e.g., country, state, data center).

[0033] Rapid Elasticity: Computing power can be prepared quickly and flexibly, allowing it to scale out automatically and immediately, and to be quickly released and scale in immediately. To consumers, the computing power available for preparation often appears unlimited and can be purchased in any quantity at any time.

[0034] Measured Services: Cloud systems leverage metric capabilities at a certain level of abstraction, appropriate for the type of service (e.g., storage, processing, bandwidth, active user accounts), to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.

[0035] The service model is as follows:

[0036] Software as a Service (SaaS): The functionality offered to consumers is the ability to use the provider's applications running on a cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, except for configuring a limited number of user-specific applications.

[0037] Platform as a Service (PaaS): The functionality offered to consumers is the ability to deploy applications they have created or acquired to cloud infrastructure using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, and storage, but they can control the deployed applications and, in some cases, the configuration of their hosting environment.

[0038] Infrastructure as a Service (IaaS): The functionality provided to consumers is the provision of processors, storage, networking, and other basic computing resources that enable consumers to deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they can control the operating system, storage, and deployed applications, and in some cases, partially control certain network components (e.g., host firewalls).

[0039] The deployment model is as follows:

[0040] Private Cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by that organization or a third party and can reside on-premises or off-premises.

[0041] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common interests (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can reside on-premises or off-premises.

[0042] Public Cloud: This cloud infrastructure is provided to a large number of people or large industry groups and is owned by organizations that sell cloud services.

[0043] Hybrid Cloud: This cloud infrastructure combines two or more cloud models (private, community, or public). While maintaining the unique entities of each model, they are bound together by standards or individual technologies to achieve data and application portability (e.g., cloud bursting for load balancing across clouds).

[0044] Cloud computing environments are service-oriented environments that emphasize statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is the infrastructure, which includes a network of interconnected nodes.

[0045] figure 2The diagram shows an exemplary cloud computing environment 50. As shown in the diagram, the cloud computing environment 50 includes one or more cloud computing nodes 10. Local computer devices used by cloud consumers (e.g., personal digital assistants (PDAs) or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or a combination thereof) can communicate with these nodes. The nodes 10 can communicate with each other. The nodes 10 can be grouped physically or virtually (not shown) in one or more networks, such as the private, community, public, or hybrid clouds or a combination thereof. This allows the cloud computing environment 50 to provide infrastructure, platforms, or software as a service, or a combination thereof, without requiring cloud consumers to maintain resources on their local computer devices. 2 The types of computer devices 54A to N shown are merely examples, and it should be understood that the computing node 10 and the cloud computing environment 50 can communicate with any type of electronic device via any type of network or network addressable connection (e.g., using a web browser) or both.

[0046] figure 3 Referring to the cloud computing environment 50 (Figure) 2 The set of functional abstraction layers provided by ) is shown. Note that the figure 3 The components, layers, and functions shown are illustrative only, and embodiments of the present invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided.

[0047] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include a mainframe 61, a reduced instruction set computer (RISC) architecture-based server 62, server 63, blade server 64, storage 65, and a network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68. The virtualization layer 70 provides an abstraction layer. From this layer, for example, the following virtual entities can be provided: a virtual server 71, virtual storage 72, a virtual network 73 including a virtual private network, a virtual application and operating system 74, and a virtual client 75.

[0048] As an example, the management layer 80 can provide the following functions: Resource preparation 81 enables the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 82 enables cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only protection of data and other resources but also identification and verification of cloud consumers and tasks. The user portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 enables the allocation and management of cloud computing resources to ensure that requested service levels are met. Service Level Agreement (SLA) planning and execution 85 enables the pre-arrangement and procurement of cloud computing resources that are expected to be needed in the future in accordance with the SLA.

[0049] Workload Layer 90 provides examples of the capabilities available in a cloud computing environment. Examples of workloads and capabilities available from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analytics processing 94, transaction processing 95, and RVM 96.

[0050] Figure 4 shows an example of a representative neural network (or "Network") 400 of one or more artificial neural networks capable of running RVM on one or more speech devices in an environment consistent with embodiments of the present disclosure. The exemplary neural network 400 consists of multiple layers. Network 400 includes an input layer 410, a hidden section 420, and an output layer 450. Network 400 depicts a feedforward neural network, but other neural network layouts, such as a recurrent neural network layout (not shown), may also be contemplated. In some embodiments, Network 400 may be a design-and-run neural network, and the depicted layout may be created by a computer programmer. In some embodiments, Network 400 may be a design-by-run neural network, and the depicted layout may be generated by a process of inputting data and analyzing that data according to one or more defined heuristics. Network 400 may operate in a forward-propagating manner by receiving inputs and outputting the results of those inputs. Network 400 may adjust the values ​​of various components of the neural network by backpropagating (reverse propagation).

[0051] The input layer 410 includes a series of input neurons 412-1, 412-2, 412-n (collectively, 412), and a series of input connections 414-1, 414-2, 414-3, 414-4, etc. (collectively, 414). The input layer 410 represents input from data that the neural network is to analyze (e.g., a series of speech inputs, physical actions, associated timestamps, expected routines, patterns, etc.). Each input neuron 412 can represent a subset of the input data. For example, the neural network 400 is provided with various numerical representations as input, with speech device information and user input being represented by numerical representations.

[0052] In another example, input neuron 412-1 may be the first pixel of the picture, input neuron 412-2 may be the second pixel of the picture, and so on. The number of input neurons 412 may correspond to the size of the input. For example, if neural network 400 is designed to analyze a 256x256 pixel image, the neural network layout may contain a set of 65,536 input neurons. The number of input neurons 412 may correspond to the type of input. For example, if the input is a 256x256 pixel color image, the neural network layout may contain a set of 196,608 input neurons (65,536 for each of the red, green, and blue values ​​of each pixel). 2The input neurons may include a 412 input neuron. The type of input neuron can correspond to the type of input. In the first example, the neural network can analyze a grayscale image, and each input neuron can be a decimal value between 0.00001 and 1 representing the grayscale level of a pixel (where 0.00001 represents a completely white pixel and 1 represents a completely black pixel). In the second example, the neural network can analyze a color image, and each input neuron can be represented as a three-dimensional vector of color values ​​for a given pixel in the input image. The first component of the vector may be a red integer between 0 and 255, the second component may be a green integer between 0 and 255, and the third component of the vector may be a blue integer between 0 and 255.

[0053] Input connections 414 represent the output of input neurons 412 to the hidden section 420. Each input connection 414 varies based on multiple weights (not shown) depending on the value of each input neuron 412. For example, the first input connection 414-1 has a value provided to the hidden section 420 based on input neuron 412-1 and a first weight. Continuing the example, the second input connection 414-2 has a value provided to the hidden section 420 based on input neuron 412-1 and a second weight. Continuing the example further, the third input connection 414-3 is based on input neuron 412-2 and a third weight, and so on. Alternatively, input connections 414-1 and 414-2 may share the same output component of input neuron 412-1, and input connections 414-3 and 414-4 may share the same output component of input neuron 412-2, and all four input connections 414-1, 414-2, 414-3, and 414-4 may have output components with four different weights. The network neural 400 may have different weightings for each connection 414, although some embodiments may intend similar weightings. In some embodiments, the values ​​for each input neuron 412 and connection 414 may be stored in memory.

[0054] Hidden section 420 includes one or more layers that receive input and produce output. Hidden section 420 includes a first hidden layer of computational neurons 422-1, 422-2, 422-3, 422-4, up to 422-n (collectively 422), a second hidden layer of computational neurons 426-1, 426-2, 426-3, 426-4, 426-5, up to 426-n (collectively 426), and a set of hidden connections 424 that connect the first and second hidden layers. It should be understood that neural network 400 is merely one of many neural networks that can monitor input from user interaction with an audio device consistent with some embodiments of this disclosure. As a result, in some embodiments of this disclosure, the exemplary hidden section may include more or fewer hidden layers than the two hidden layers described with respect to hidden section 420 (e.g., one hidden layer, seven hidden layers, twelve hidden layers, etc.).

[0055] The first hidden layer 422 contains computation neurons up to 422-1, 422-2, 422-3, 422-4, and 422-n. Each computation neuron in the first hidden layer 422 can receive one or more connections 414 as input. For example, computation neuron 422-1 receives input connections 414-1 and 414-2. Each computation neuron in the first hidden layer 422 also provides an output. The output is represented by a dotted line of hidden connections 424 flowing out of the first hidden layer 422. Each of the computation neurons 422 executes an activation function during forward propagation. In some embodiments, the activation function may be a process (e.g., a perceptron) that receives multiple binary inputs and computes a single binary output. In some embodiments, the activation function may be a process that receives several non-binary inputs (e.g., a number between 0 and 1, 0.671, etc.) and computes a single non-binary output (e.g., a number between 0 and 1, a number between -0.5 and 0.5, etc.). Various functions may be performed to compute the activation function (e.g., sigmoid neurons or other logistic functions, tanh neurons, softplus functions, softmax functions, rectified linear units, etc.). In some embodiments, each of the computational neurons 422 also includes a bias (not shown). The bias may be used to determine the likelihood or evaluation of a given activation function. In some embodiments, each of the bias values ​​for each of the computational neurons must necessarily be stored in memory.

[0056] The neural network 400 may include using a sigmoid neuron as the activation function for the computational neuron 422-1. Equation (Equation 1 below) can express the activation function of computational neuron 422-1 as f(neuron). The logic of computational neuron 422-1 may be the sum of each input connection that inputs to computational neuron 422-1 (i.e., input connection 414-1 and input connection 414-3), and is represented as j in Equation 1. For each j, the weight w is multiplied by the value x of a given connected input neuron 412. The bias of computational neuron 422-1 is represented as b. When each input connection j is summed, the bias b is subtracted. In this example, the output of computation neuron 422-1 is determined as follows: given a large positive number resulting from the sum of activations f(neurons) and bias, the output of computation neuron 422-1 approaches approximately 1; given a large negative number resulting from the sum of activations f(neurons) and bias, the output of computation neuron 422-1 approaches approximately 0; given a number intermediate between the large positive and large negative numbers resulting from the sum of activations f(neurons) and bias, the output changes slightly because the weights and bias change slightly.

[0057]

number

[0058] The second hidden layer 426 includes computation neurons 426-1, 426-2, 426-3, 426-4, 426-5, and 426-n. In some embodiments, the computation neurons of the second hidden layer 426 may operate similarly to the computation neurons of the first hidden layer 422. For example, computation neurons 426-1 to 426-n may each operate with the same activation function as computation neurons 422-1 to 422-n. In some embodiments, the computation neurons of the second hidden layer 426 may operate differently from the computation neurons of the first hidden layer 422. For example, computation neurons 426-1 to 426-n may have a first activation function, and computation neurons 422-1 to 422-n may have a second activation function.

[0059] Similarly, the connectivity of the hidden section 420 to, from, and between layers can also vary. For example, input connection 414 may be fully connected to the first hidden layer 422, and hidden connection 424 may be fully connected from the first hidden layer to the second hidden layer 426. In some embodiments, full connectivity may mean that each neuron in a given layer can be connected to all neurons in the previous layer. In some embodiments, full connectivity may mean that each neuron in a given layer functions completely independently and does not share connections. In the second example, input connection 414 may not be fully connected to the first hidden layer 422, and hidden connection 424 may not be fully connected from the first hidden layer to the second hidden layer 426.

[0060] Furthermore, the parameters to, from, and between layers of the hidden section 420 can vary. In some embodiments, the parameters may include weights and biases. In some embodiments, the parameters may be more or less than the weights and biases. In some embodiments of this disclosure, the network 400 may be a convolutional neural network or a convolutional network. A convolutional neural network may include a set of heterogeneous layers (e.g., an input layer 410, a convolutional layer 422, a pooling layer 426, and an output layer 450). In such a network, the input layer 410 may hold raw audio data of an audio input with start time, sample length, and two-dimensional volume of the microphone source. The convolutional layers of such a network may output from connections that are only locally present in the input layer to identify features of small sections of an image (e.g., the start of a verbal command from a user, the length of time it takes for the user to provide the verbal command, etc.). In this example, the convolutional layers may include weights and biases, as well as additional parameters (e.g., depth, stride, and padding). The pooling layer 426 of such a network takes the output of the convolutional layer 422 as input, but may perform a fixed-function operation (e.g., an operation that does not consider weights or biases). Also, considering this example, the pooling layer may not include any convolution parameters, nor may it include any weights or biases (e.g., it performs a downsampling operation).

[0061] The output layer 450 includes a series of output neurons 450-1, 450-2, 450-3, and up to 450-n (collectively 450). The output layer 450 holds the analysis results of the neural network 400. In some embodiments, the output layer 450 may be a classification layer used to identify features of inputs to the network 400. For example, the network 400 may be a classification network trained to identify Arabic numerals. In such an example, the network 400 may include an output layer 450 with 10 output neurons corresponding to which Arabic numerals the network identified (for example, output neuron 450-2 having a higher activation value than output neuron 450 may indicate that the neural network determined that the image contains the digit "1"). In some embodiments, the output layer 450 may also be a real-valued target (e.g., attempting to predict an outcome when the input is a set of previous outcomes) and may have specific output neurons (not shown). The output layer 450 is supplied from output connection 452. The output connection 452 provides activation from the hidden section 420. In some embodiments, the output connection 452 may include weights, and the neurons in the output layer 450 may include biases.

[0062] Training a neural network, as depicted by neural network 400, may include performing backpropagation. Backpropagation is different from forward propagation. Forward propagation may include supplying data to input neurons 410, performing computations on connections 414, 424, and 452, and performing computations on computation neurons 422 and 426. The term forward propagation may also be used to describe the layout of a given neural network (e.g., recursion, number of layers, number of neurons in one or more layers, whether layers are fully connected to other layers or not, etc.).

[0063] In contrast, backpropagation can determine the error of parameters (e.g., weights and biases) within the network 400 by backpropagating the error starting from output neuron 450 and passing through various connections 452, 424, 414 and layers 426, 422, respectively.

[0064] Backpropagation can involve running one or more algorithms on one or more training datasets to reduce the difference between what a given neural network decides from its inputs and what a given neural network should decide from its inputs. The difference between the network's decision and the correct decision is sometimes called the objective function (or alternatively, the cost function). When a given neural network is first created, data is provided, and it is computed by forward propagation, the result or decision may be an incorrect decision.

[0065] For example, the neural network 400 may be a classification network. Furthermore, the network 400 may receive oral input, such as a command directed to the first of five speech devices. Furthermore, the network 400 may determine that the oral command is relatively most likely to be directed to the third of the five speech devices. Furthermore, the network 400 may determine that the next most likely speech device is the fifth speech device, and the next most likely speech device is the first speech device (and so on for other inputs). Continuing this example, backpropagation may change the weight values ​​of connections 414, 424, and 452, and may change the bias values ​​of the first layer of computation neuron 422, the second layer of computation neuron 426, and output neuron 450. Continuing the example further, backpropagation may result in a future outcome that is a relatively accurate classification of the same speech input containing the same oral command (e.g., ranking a linguistic command directed to the first of five speech devices as a relatively likely speech device).

[0066] Equation 2 provides an example of an objective function ("exemplary function") in the form of a quadratic cost function (e.g., mean squared error), other functions may be chosen, and mean squared error is chosen for illustrative purposes. In Equation 2, all weights may be represented by w, and the bias may be represented by b of the neural network 400. The network 400 is given a given number of training inputs n in a subset (or whole) of training data having input values ​​x. The network 400 can obtain output a from x and should obtain a desired output y(x) from x. Backpropagation or training of the network 400 should be the reduction or minimization of the objective function "O(w,b)" by changing the set of weights and biases. Successful training of the network 400 should include not only the reduction of the difference between the answer a for an input value x and the correct answer y(x), but also the case when new input values ​​(e.g., from additional training data, from validation data, etc.) are given.

[0067]

number

[0068] Backpropagation algorithms can utilize many options in both the objective function (e.g., mean squared error, cross-entropy cost function, precision function, confusion matrix, precision-reproduction curve, mean absolute error, etc.) and the reduction of the objective function (e.g., gradient descent, batch-based stochastic gradient descent, Hessian optimization, momentum-based gradient descent, etc.). Backpropagation may include using a gradient descent algorithm (e.g., calculating partial derivatives of the objective function in relation to weights and biases for all training data). Backpropagation may include deciding on stochastic gradient descent (e.g., calculating partial derivatives for a subset of the training data or a subset of training inputs within a batch). Also, various backpropagation algorithms may include additional parameters (e.g., learning rate for gradient descent). Large changes to weights and biases due to backpropagation can lead to erroneous training (e.g., overfitting to the training data, reduction to the local minimum, excessive reduction beyond the global minimum, etc.). Therefore, erroneous training can be prevented by modifying the objective function to have more parameters (e.g., using an objective function that incorporates regularization to prevent overfitting). Furthermore, as a result, the changes in the neural network 400 may be small in any given iteration. The backpropagation algorithm can perform many more iterations to achieve more accurate learning as a result of the relative smallness of any given iteration.

[0069] For example, neural network 400 may have untrained weights and biases, and backpropagation may involve stochastic gradient descent to train the network on a subset of the training inputs (e.g., a batch of 10 training inputs from the entire training inputs). Continuing this example, network 400 could continue training using a second subset of the training inputs (e.g., a second batch of 10 training inputs from the whole, excluding the first batch), and this could be repeated until all the training inputs have been used to compute the gradient descent (e.g., one epoch of training data). In other words, if there are 10,000 training images in total and a batch size of 100 training inputs is used in one iteration, it would take 1,000 iterations to complete one epoch of training data. Many epochs can be performed to continue training the neural network. Many factors may be involved in determining the selection of additional parameters (e.g., a large batch size may result in improper training, a small batch size may result in too many training iterations, a large batch size may not fit in memory, a small batch size may not efficiently utilize discrete GPU hardware, too few training epochs may result in an untrained network, too many training epochs may result in overfitting of the trained network, etc.). Furthermore, Network 400 can be evaluated to quantify its ability to evaluate the dataset, for example, by using evaluation metrics (e.g., mean squared error, cross-entropy cost function, precision function, confusion matrix, precision-reproduction curve, mean absolute error, etc.).

[0070] Figure 5 shows an exemplary system 500 for voice device management consistent with some embodiments of the present disclosure. System 500 may be configured to run an RVM, which includes monitoring user activity during interaction with a voice-enabled computing device or a voice-controlled computing device. System 500 may include at least a network 510, one or more network-connected devices 520-1, 520-2, 520-3, and 520-4, and an RVM 530.

[0071] Network 510 can be implemented using any number of suitable physical, logical, or both communication topologies. Network 510 may include one or more private or public computing networks. For example, network 510 may include a private network associated with the workload (e.g., a network with a firewall that blocks unauthorized external access). Alternatively, or additionally, network 510 may include a public network such as the Internet. Thus, network 510 can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet, or a combination thereof. Network 510 may include one or more servers, networks, or databases, and can transfer data between elements of system 500 using one or more communication protocols. Furthermore, although illustrated as a single entity in Figure 5, in other examples network 510 may include multiple networks, such as public networks, private networks, or a combination thereof. Communication network 510 may include various types of physical communication channels or "links." Links can be wired, wireless, optical, or any other suitable medium, or a combination thereof. In addition, the communication network 510 may include various network hardware and software for performing routing, switching, and other functions, such as routers, switches, base stations, bridges, or any other devices that may be useful in facilitating data communications.

[0072] The connected device 520 may be a computing device configured to perform one or more operations of the RVM 530. The connected device 520 may be a computer system, such as computer 100. The connected device 520 may be a cloud computing system. For example, connected device 520-4 may be a representative of one or more servers that make up the cloud computing environment 50. The connected device 520 may be a voice device. For example, connected device 520-1 may be a voice-operated smartphone configured to receive and respond to verbal commands. The connected device 520 may include a household appliance configured to perform routine tasks in response to user input. For example, connected device 520-2 may be a washing machine, and connected device 520-3 may be a clothes dryer. The user can perform actions on the washing machine 520-2, the clothes dryer 520-3, or both, such as physically placing dirty clothes into the washing machine. Furthermore, the user can execute voice commands to the washing machine 520-2, the clothes dryer 520-3, or both. For example, they might say, "Please set the mode to 3" to refer to washing clothes under the third of five different washing settings.

[0073] RVM530 may include one or more operations of a configured set of software or hardware or both for performing voice device management. RVM530 may also take the form of multiple artificial intelligence components, such as machine learning models ("ML models"). Each ML model may be an instance of a neural network, such as neural network 400. An ML model may include a detection model 540, a missing input model 550, a misinput model 560, and a corrective action model 570. RVM530 may also include an activity model 532 and a natural language processor 534 ("NLP").

[0074] In some embodiments, the detection model 540, the missing input model 550, the erroneous input model 560, and the corrective action model 570 can perform machine learning on the activity model 532 using one or more of the following exemplary techniques: K-Nearest Neighbors (KNN), Learned Vector Quantization (LVQ), Self-Organizing Maps (SOM), Logistic Regression, Ordinary Least Squares Regression (OLSR), Linear Regression, Stepwise Regression, Multivariate Adaptive Regression Spline (MARS), Ridge Regression, Least Absolute Shrinkage Selection Operator. LASSO), elastic networks, least-angle regression (LARS), stochastic classifiers, naive Bayes classifiers, binary classifiers, linear classifiers, hierarchical classifiers, canonical correlation analysis (CCA), factor analysis, independent component analysis (ICA), linear discriminant analysis (LDA), multidimensional scaling (MDS), non-negative metric decomposition (NMF), partial least-squares regression (PLSR), principal component analysis (PCA), principal component regression (PCR), summon mapping, t-distribution stochastic neighbor embedding (t-SNE), bootstrap aggregation, ensemble mean, gradient boosting decision tree (G BRT), Gradient Boosting Machine (GBM), Inductive Bias Algorithms, Q-Learning, State-Action-Reward-State-Action (SARSA), Time-Lag (TD) Learning, A priori Algorithms, Equivalent Class Transformation (ECLAT) Algorithm, Gaussian Process Regression, Gene Expression Programming, Grouping Method for Data Processing (GMDH), Inductive Logic Programming, Instance-Based Learning, Logistic Model Trees, Information Fuzzy Networks (IFN), Hidden Markov Models, Gaussian Naive Bayes, Polynomial Naive Bayes, Meaning-Based 1-Dependency Estimator (AODE), Bayesian Networks (BN), Classification Regression Trees (CART), Chi-Square Automatic Interaction Detection (CHAID), Expectation Maximization Algorithms, Feedforward Neural Networks, Logic Learning Machines, Self-Organizing Maps, Single-Connected Clustering, Fuzzy Clustering, Hierarchical Clustering, Boltzmann Machines, Convolutional Neural Networks, Recurrent Neural Networks, Hierarchical Time Memory (HTM), or other machine learning techniques, or combinations thereof.

[0075] The ML model of RVM530 may be configured to process language using artificial intelligence techniques such as natural language processing, using a natural language processor 534. In some embodiments, the natural language processor 534 may include various components (not shown) that operate in hardware, software, or some combination. For example, the natural language processor 534 may include one or more data sources, a search application, and a report analyzer. The natural language processor 534 may also be a computer module that analyzes incoming content and other information. The natural language processor 534 may perform various methods and techniques for analyzing text information (e.g., syntactic analysis, semantic analysis, etc.). The natural language processor 534 may be configured to recognize and analyze any number of natural languages. In some embodiments, the natural language processor 534 may analyze a document or a passage of content from speech received from a user. Various components (not shown) of the natural language processor 534 include, but are not limited to, a tokenizer, a part-of-speech (POS) tagger, a semantic relation identifier, and a syntactic relation identifier. The natural language processor 534 may include a support vector machine (SVM) generator for processing the content of topics found in the corpus and for classifying those topics.

[0076] In some embodiments, the tokenizer may be a computer module that performs lexical analysis. The tokenizer can convert a sequence of characters into a sequence of tokens. Tokens may be strings of characters contained in an electronic document and classified as meaningful symbols. Furthermore, in some embodiments, the tokenizer can identify word boundaries within an electronic document and divide any text passage in the document into constituent text elements such as words, multi-word tokens, numbers, and punctuation marks. In some embodiments, the tokenizer can receive a string, identify the vocabulary within the string, and classify them into tokens.

[0077] In accordance with various embodiments, a POS tagger may be a computer module that marks up words in a passage to correspond to specific parts of speech. A POS tagger can read a text written in natural language or other text and assign a part of speech to each word or other token. Based on the definition of the word and the context of the word, a POS tagger can determine the part of speech to which a word (or other text element) corresponds. The context of a word may be based on its relationship to adjacent or related words within a phrase, sentence, or paragraph.

[0078] In some embodiments, the context of a word may depend on one or more previously analyzed electronic documents (e.g., previous voice invocation routines, previous usage patterns including voice utterances, information stored in activity model 532). Examples of parts of speech that can be assigned to a word include, but are not limited to, nouns, verbs, adjectives, and adverbs. Examples of other parts of speech categories that the POS tagger may assign include, but are not limited to, comparative adverbs, superordinate adverbs, WH-adverbs, conjunctions, determiners, negative particles, possessive markers, prepositions, and WH-pronouns. In some embodiments, the POS tagger may tag or annotate tokens in a passage with part-of-speech categories. In some embodiments, the POS tagger may tag tokens or words in a passage that are parsed by a natural language processing system.

[0079] In some embodiments, the semantic relation identifier may be a computer module configured to identify semantic relationships between recognized text elements (e.g., words, phrases) in a document. In some embodiments, the semantic relation identifier may determine functional dependencies and other semantic relationships between entities.

[0080] In accordance with various embodiments, a syntactic relation identifier may be a computer module that can be configured to identify syntactic relations within a passage composed of tokens. The syntactic relation identifier may determine the grammatical structure of a sentence, for example, which groups of words are associated as a phrase and which words are the subject or object of a verb. The syntactic relation identifier may conform to a formal grammar.

[0081] In some embodiments, the natural language processor 534 may be a computer module capable of parsing a document and generating corresponding data structures for one or more parts of the document. For example, in response to receiving speech from a user in a natural language processing system, the natural language processor 534 may output parsed text elements from the data. In some embodiments, the parsed text elements may be represented in the form of a parse tree or other graph structure. To generate parsed text elements, the natural language processor 534 may trigger a computer module including a tokenizer, a part-of-speech (POS) tagger, an SVM generator, semantic relation identifiers, and syntactic relation identifiers.

[0082] In some embodiments, a natural language processing system can perform machine learning (ML) text operations by leveraging one or more exemplary machine learning techniques. Specifically, the RVM530 may operate to perform machine learning text classification or machine learning text comparison or both. Machine learning text classification may include ML text operations that convert characters, text, words, and phrases into numerical values. The numerical values ​​are then input into a neural network that can determine various features, characteristics, and other information of the words in relation to a document or in relation to other words (for example, classifying numerical values ​​associated with a word can enable word classification). Machine learning text comparison may involve using the numerical values ​​of converted characters, text, words, and phrases to perform comparisons. The comparison may be between the numerical values ​​of a first word or other text and the numerical values ​​of a second word or other text. The determination of a machine learning text comparison may involve determining scoring, correlation, or association (for example, the relationship between a first numerical value of a first word and a second numerical value of a second word). The comparison is used to determine whether two words are similar or different based on one or more criteria. Numerical operations for machine learning-based text classification / comparison may be functions of mathematical operations performed via a neural network, such as linear regression, addition, or other related mathematical operations on numerical values ​​representing words or other text.

[0083] ML text manipulation can include word encoding, such as one-hot encoding of words from tokenizers, POS taggers, semantic relation identifiers, syntactic relation identifiers, etc. ML text manipulation can also include the use of text vectorization, such as vectorizing words from tokenizers, POS taggers, semantic relation identifiers, syntactic relation identifiers, etc. For example, a paragraph of text might contain the phrase "Oranges are fruit that grows on trees." Vectorizing the word "orange" might involve setting up input neurons in a neural network for different words in the phrase containing the word "orange." The output value might be an array of values ​​(e.g., 48 digits, thousands of digits). The output value might tend towards "1" for related words and towards "0" for unrelated words. Related words might be associated based on one or more of the following: similar parts of speech, syntactic meaning, locality within a sentence or paragraph, or other related "closeness" between the input and other parts of natural language (e.g., other parts of the phrase "Oranges are fruit that grows on trees," other parts of the paragraph containing the phrase, other parts of the language).

[0084] In some embodiments, each of the ML models (e.g., detection model 540, missing input model 550, incorrect input model 560, and corrective action model 570) may be configured to perform a different action.

[0085] The detection model 540 may be a fully connected neural network and may be configured to receive the current time and user interactions from the user (e.g., physical activity, voice commands) as inputs. The detection model 540 may include various outputs, such as the name of a particular connected device 520 and whether a user interaction was expected but not received. The detection model 540 may also perform calculations to determine conditions, such as whether an expected second user interaction or follow-up user interaction was expected (e.g., a softmax calculation). For example, when any voice device is activated at 7 a.m. by switching it on or by placing any object, such as a coffee cup, under the spout of a coffee maker, the detection model 540 may determine whether those actions, devices, or times, or combinations thereof, are part of a regular routine, pattern, or usage scenario of the user.

[0086] The missing input model 550 may be a recurrent neural network configured to act as an encoder / decoder. The missing input model 550 may receive a previous user interaction (e.g., an interaction received by the detection model 540) as input. The missing input model 550 may also receive a previous timestamp (e.g., a timestamp of a previous user interaction received by the detection model 540) and a current timestamp (e.g., time after the previous timestamp). The missing input model 550 may output a predicted time for a follow-up or second instruction in order to output the time for the predicted follow-up or second instruction. This may be considered a determination of a potential second input. The missing input model 550 may perform a classifier operation. The classifier operation may include monitoring for deviations from a potential second input and identifying anomalies in activity. For example, the classifier operation may identify or predict anomalies in activity, including the disappearance or omission of a predicted command. The missing input model 550 can output the name of a specific connected device 520 and an identified value (e.g., "yes" or "no") indicating that the expected command is missing. For example, if the first instruction was "grill at 300 degrees for 10 minutes", the missing input model 550 can begin determining the next most likely instruction and time. The next most likely instruction and time might be to lower the temperature to 200 degrees and grill for 10 minutes. Furthermore, the missing input model 550 may determine that this next most likely instruction is expected within 20 minutes of receiving the first instruction.

[0087] The erroneous input model 560 may be a recurrent neural network configured to function as an encoder / decoder for long-term short-term memory, etc. The erroneous input model 560 may also include dense layers, fully connected layers, or both. The recurrent portion of the erroneous input model 560 can pass data to the fully connected layers. The erroneous input model 560 may also include an encoder that creates a latent space and a decoder (e.g., a variational autoencoder ("VAE")) that uses that latent space to generate an output. The latent space created by the encoder may be in the form of two distributions: a mean distribution and a covariance distribution. As a result, the erroneous input model 560 may be configured to analyze all inputs and generate as an output a Gaussian distribution of the mean of all inputs and a Gaussian distribution of the standard deviation of all inputs. The erroneous input model 560 may be configured to score the output and determine whether the score exceeds a predetermined threshold. This could be considered the determination of a potential second input, monitoring of deviations from a potential second input, or both, and the identification of anomalies in activity. If the score exceeds a predetermined threshold, an anomaly input may be in the user interaction. For example, the erroneous input model 560 may determine that the deviation level of a command related to grilling on a stove exceeds a predetermined threshold. The expected command might be "grill for another 10 minutes," but the actually received command might be "grill for another 15 minutes," in which case the erroneous input model 560 may generate a score below a predetermined threshold associated with a specific voice device, a routine containing a specific voice device, or both. If the actually received command is "grill for another 40 minutes," the erroneous input model 560 might generate a score above a predetermined threshold associated with a specific voice device, a routine, or both.

[0088] The corrective action model 570 may be a recurrent neural network configured to perform encoding, decoding, or both. The corrective action model 570 can receive input from the detection model 540, the missing input model 550, and the erroneous input model 560. The corrective action model 570 may also output a corrective action to the user.

[0089] The ML model of RVM530 may collaborate to receive user interactions (e.g., physical activities and voice commands) with various voice devices in the environment to identify anomalous activity (e.g., inputs different from expected inputs). Specifically, 580 can receive user interactions from connected devices 520. User interactions may include voice commands such as "Turn on the coffee maker" or "Do the laundry for 10 minutes." The corrective action model 570 can generate corrective actions in 590 and provide these corrective actions to the user. Corrective actions may take the form of questions such as "Were you going to turn on the washing machine?" or "The washing machine has finished, would you like to start the drying cycle?" Corrective actions may also take the form of verbal statements, such as "You usually set the microwave to 3 minutes" in a voice file. Corrective actions may also take the form of visual statements. For example, connected device 520-1 may display a message such as "The microwave is scheduled for another 10 minutes of 200-degree grilling cycles," along with touchscreen buttons for "OK" or "Cancel."

[0090] Figures 6A, 6B, 7, 8, 9A, and 9B show examples of training data for one or more parts of RVM530 used when performing artificial intelligence techniques to identify abnormal inputs to a speech device and perform corrective actions accordingly. Specifically, the training data may include data inputs and data outputs for updating and training a machine learning model. The model, its trained data, or both may be stored in and / or updated in the activity model 532. The model may be further trained as additional real-world inputs are provided through the use of system 500.

[0091] Figure 6A shows a first portion of training data 600 for a first machine learning model of system 500 for identifying abnormal inputs, consistent with some embodiments of the present disclosure. Figure 6B shows a second portion of training data 600 for a first machine learning model of system 500, consistent with some embodiments of the present disclosure. Specifically, Figure 6A shows several rows 610 of the training data 600 provided to detection model 540. Furthermore, Figure 6B shows a continuation of rows 610 of the training data 600. Each row 610 represents a specific routine, pattern, or activity that can be processed by a neural network configured to detect abnormal behavior of an audio device. Row 610-1 may represent a header row containing descriptive information for the training data 600. Row 612 may represent an additional data row that may be included in the training data 600 but is not shown.

[0092] Column 620 can represent each element of the training data 600 that can be used to train and update the weights or biases, or both, of the detection model 540. Specifically, training the detection model 540 may include only a subset of elements 620 provided as input, such as elements 620-1, 620-2, 620-3, and 620-4. Additional elements, such as elements 620-5 and 620-6, may be provided as part of the expected output. The expected output may be used for comparison and for the purpose of training the neural network of the detection model 540. For example, if the detection of the training data 600 does not result in an accurate identification of the presence of an abnormal body movement or voice command, the expected output may be used to update the weights and biases of the detection model 540.

[0093] Figure 7 shows a portion of training data 700 for a second machine learning model of System 500 for identifying anomalous inputs, consistent with some embodiments of the present disclosure. Specifically, Figure 7 depicts several rows 710 of the training data 700 provided to a missing input model 550. Each row 710 represents a specific routine, pattern, or activity that can be processed by a neural network configured to detect anomalous operation of a speech device. Row 710-1 may represent a header row containing descriptive information for the training data 700. Row 712 may represent an additional data row, not shown, that may be included in the training data 700.

[0094] Column 720 can represent each element of the training data 700 that may be used to train and update the weights or biases, or both, of the missing input model 550. Specifically, training the missing input model 550 may include only a subset of the elements 720 provided as input, such as elements 720-1, 720-2, 720-3, and 720-4. Additional elements, such as element 720-5, may be provided as part of the expected output. The expected output can be used for comparison and for the purpose of training the neural network of the missing input model 550. For example, if the detection of the training data 700 does not result in accurate identification of anomalous inputs, including missing follow-up voice commands to a voice device, the expected output may be used to update the weights and biases of the missing input model 550.

[0095] Figure 8 shows a portion of training data 800 for a third machine learning model of system 500 for identifying anomalous inputs, consistent with several embodiments of the present disclosure. Specifically, Figure 7 shows several rows 810 of training data 800 provided to anomalous input model 560. Each row 810 represents a specific routine, pattern, or activity that can be processed by a neural network configured to detect anomalous operation of a speech device. Row 810-1 may represent a header row containing descriptive information for the training data 800. Row 812 may represent an additional data row, not shown, that may be included in the training data 800.

[0096] Column 820 may represent each element of the training data 800 that can be used to train and update the weights or biases, or both, of the erroneous input model 560. Specifically, training the erroneous input model 560 may include only a subset of elements 820 provided as input, such as elements 820-1, 820-2, 820-3, and 820-4. Additional elements, such as element 820-5, may be provided as part of the expected output. The expected output may be used for comparison and for the purpose of training the neural network of the erroneous input model 560. For example, if the detection of the training data 800 does not result in an accurate identification of erroneous inputs, including unexpected physical behavior or unexpected voice commands to a voice device, the expected output may be used to update the weights and biases of the erroneous input model 560.

[0097] Figure 9A shows a first portion of training data 900 for a fourth machine learning model of system 500 to provide precise corrective behaviors, consistent with some embodiments of the present disclosure. Figure 9B shows a second portion of training data 900 for a fourth machine learning model of system 500, consistent with some embodiments of the present disclosure. Specifically, Figure 9A shows several rows 910 of the training data 900 provided to the corrective behavior model 570. Furthermore, Figure 9B shows a continuation of rows 910 of the training data 900. Each row 910 represents a specific routine, pattern, or activity that can be processed by a neural network configured to format and generate corrective behaviors. Row 910-1 may represent a header row containing descriptive information for the training data 900. Row 912 may represent an additional row of data that is not depicted but may be included in the training data 900.

[0098] Column 920 can represent each element of the training data 900 that can be used to train and update the weights or biases, or both, of the modified operating model 570. Specifically, training the modified operating model 570 may include only a subset of elements 920 provided as input, such as elements 920-1, 920-2, 920-3, and 920-4. Additional elements, such as elements 920-5, 920-6, and 920-7, may be provided as part of the expected output. The expected output may be used for comparison and for the purpose of training the neural network of the modified operating model 570. For example, if the detection of the training data 900 does not yield accurate modified behavior, the expected output may be used to update the weights and biases of the modified operating model 570.

[0099] Figure 10 shows a method 1000 for performing a modification operation of an audio device, consistent with some embodiments of the present disclosure. Method 1000 may generally be implemented with fixed functional hardware, configurable logic, logic instructions, or any combination thereof. For example, logic instructions may include assembler instructions, ISA instructions, machine instructions, machine-dependent instructions, microcode, state setting data, configuration data for integrated circuits, state information for personalizing electronic circuits, or other structural components native to hardware (e.g., host processor, central processing unit / CPU, microcontroller, etc.), or a combination thereof.

[0100] Method 1000 can be initiated at 1005 by receiving one or more user interactions at 1010. User interactions may be directed to a set of one or more voice devices. Voice devices may be connected devices in the environment, such as connected device 520 of system 500. User interactions may be received by connected devices (e.g., connected device 520-1), such as a desktop computer or other associated computer system. User interactions may include physical actions, such as turning on a coffee maker, opening a washing machine door, or other actions or placement of items by the user. User interactions may be received at 1010 continuously, repeatedly (e.g., every second, every tenth of a second), or at other relevant intervals.

[0101] Based on user interaction, detection of a first input may be initiated at 1020. Detection may include processing by a connected device using location, motion, sound, or other relevant sensors. Specifically, user interaction received at 1010 may be transmitted to a connected device via a network, such as a local area network. Detection may be performed by a voice device or by a connected device. Each input of a user interaction, including the first input, may be provided to an activity model for processing. Specifically, the first input may act to trigger a machine learning model that performs one or more artificial intelligence operations to determine a particular type of activity, process, pattern, or routine associated with the first input. For example, when a coffee cup is placed in a coffee maker, a process is initiated to process that input and determine whether there is a routine associated with its placement.

[0102] If a first input is detected, i.e., Y at 1030, method 1000 can continue at 1040 by determining whether there is a second input. The determination may include predicting the second input using an activity model. In the coffee cup placement example, the determination may include performing machine learning to determine the potential additional input directed towards the coffee cups.

[0103] In 1050, method 1000 can be continued by monitoring deviations from a second input. The deviation may be the difference between a potential second input determined in 1040 and any actual input received from a user interaction in 1010. Monitoring deviations may include comparing the output of the activity model with a stream of inputs received from the connected device. A deviation from a potential second input may be a missing command. A missing command may be a command that is not received within a specific predetermined threshold. For example, based on the activity model, after laundry is placed in a voice device which is a smart washing machine, the predetermined threshold may be determined to be receiving a command within 45 seconds. A deviation from a potential second input may be an unexpected physical input or verbal command. An unexpected command may be a command that contradicts a predetermined input pattern or routine. For example, based on the activity model, after a voice device which is a smart oven has been heated to 400 degrees Fahrenheit for 20 minutes, the next input may be determined to be a command to heat to 150 degrees Fahrenheit for 30 minutes.

[0104] If an abnormality in activity is identified, i.e., Y in 1060, method 1000 may continue by performing a corrective action in 1070. The corrective action 1070 may be in the form of a command, such as a command instructing the user to perform an abnormal input. The corrective action 1070 may be in the form of a request, such as a question or prompt, provided to the user. The corrective action 1070 may be provided by one of the voice devices, such as a smart home appliance, which speaks the corrective action to the user. The corrective action 1070 may be provided by one of the connected devices, such as a smartwatch worn by the user. After the corrective action 1070 has been provided in 1070, or if no abnormality was identified, i.e., N in 1060, or if no first input was detected, i.e., N in 1030, method 1000 may terminate in 1095.

[0105] The present invention may be a system, method, or computer program product or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0106] A computer-readable storage medium can be a tangible device capable of holding and storing instructions used by an instruction execution device. A computer-readable storage medium may, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), SRAM, CD-ROM, DVD, memory stick, floppy disk, punch cards, or grooved raised structures, and mechanically encoded devices on which instructions are recorded, and suitable combinations thereof. The computer-readable storage medium as used herein should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted through wires.

[0107] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). The network consists of copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. The network adapter card or network interface of each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on the computer-readable storage medium within each computing / processing device.

[0108] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++ and procedural programming languages ​​such as the C programming language or similar programming languages. The computer-readable program instructions are executable as a standalone software package, either entirely on the user's computer or partially on the user's computer. Alternatively, they may be executable partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by personalizing them using state information of computer-readable program instructions in order to perform aspects of the present invention.

[0109] Aspects of the present invention are described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and any combination of blocks in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.

[0110] These computer-readable program instructions can be provided to a computer processor or other programmable data processing device to generate a machine, such that instructions executed via the processor of the computer or other programmable data processing device generate means for implementing functions / operations specified in one or more blocks of a flowchart or block diagram or both. These computer-readable program instructions can also be stored in a computer-readable storage medium that can be connected to a computer, a programmable data processing device, or other device or combination of devices that function in a particular way, such that the computer-readable storage medium on which the instructions are stored constitutes one of the outputs containing instructions that implement the modes of functions / operations specified in one or more blocks of a flowchart or block diagram or both.

[0111] Computer-readable program instructions, like instructions that perform a function / action specified in one or more blocks of a flowchart or block diagram or both on a computer, other programmable device, or other device, can also be loaded into a computer, other programmable data processing device, or other device and perform a series of operational steps on the computer, other programmable device, or other device to produce a computer-implemented process.

[0112] The flowcharts and block diagrams in the figures illustrate the configuration, function, and operation of executable implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or part of an instruction, which constitutes one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions shown in the blocks may differ from the order shown in the figures. For example, two blocks shown consecutively may actually be achieved as a single step, executed simultaneously, substantially simultaneously, partially or entirely in overlapping time, or the blocks may be executed in reverse order depending on the functions involved. It should also be noted that each block in a block diagram or flowchart diagram, or both, and any combination of blocks in a block diagram or flowchart diagram, or both, can be implemented by a special-purpose hardware-based system that performs a specified function or operation, or a combination of special-purpose hardware and computer instructions.

[0113] The descriptions of the various embodiments of this disclosure are presented for illustrative purposes only and are not intended to be exhaustive or to limit the embodiments disclosed. It will be apparent to those skilled in the art that many modifications and changes are possible without departing from the scope of the embodiments described. The terminology used herein has been chosen to best describe the principles of the embodiments, their practical application to market-based technologies or technical improvements, or to enable those skilled in the art to understand the embodiments disclosed herein.

[0114] The descriptions of the various embodiments of this disclosure are presented for illustrative purposes only and are not intended to be exhaustive or to limit the embodiments disclosed. It will be apparent to those skilled in the art that many modifications and changes are possible without departing from the scope of the embodiments described. The terminology used herein has been chosen to best describe the principles of the embodiments, their practical application to market-based technologies or technical improvements, or to enable those skilled in the art to understand the embodiments disclosed herein.

Claims

1. The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, Includes, The aforementioned potential second input includes an expected interaction, and the deviation is a method in which the expected interaction does not exist.

2. The first connected device receives one or more user interactions directed to one or more sets of voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, Includes, The method wherein the potential second input includes a first interaction, and the deviation is the second interaction.

3. The first connected device receives one or more user interactions directed to one or more sets of voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, Includes, A method wherein the aforementioned potential second input includes a predetermined range of acceptable values, and the deviation is outside the range.

4. The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, Includes, The aforementioned potential second input includes a predetermined range of acceptable values, including a lower limit and an upper limit. The deviation is close to either the lower limit or the upper limit of the range. moreover, Updating the activity model based on the aforementioned deviations. Methods that include...

5. The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, Includes, The method wherein the potential second input includes expected interactions within a predetermined threshold time period, and the deviation is an interaction outside the predetermined threshold time period.

6. A system, wherein the system is A memory, wherein the memory includes one or more instructions, A processor, wherein the processor is communicatively coupled to the memory, and the processor responds to reading one or more instructions. The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, The aforementioned potential second input includes an expected interaction, and the deviation is a system in which the expected interaction does not exist.

7. A system, wherein the system is A memory, wherein the memory includes one or more instructions, A processor, wherein the processor is communicatively coupled to the memory, and the processor responds to reading one or more instructions. The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, The system wherein the potential second input includes a first interaction, and the deviation is the second interaction.

8. A system, wherein the system is A memory, wherein the memory includes one or more instructions, A processor, wherein the processor is communicatively coupled to the memory, and the processor responds to reading one or more instructions. The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, The system wherein the aforementioned potential second input includes a predetermined range of acceptable values, and the deviation is outside the range.

9. A system, wherein the system is A memory, wherein the memory includes one or more instructions, A processor, wherein the processor is communicatively coupled to the memory, and the processor responds to reading one or more instructions. The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, The aforementioned potential second input includes a predetermined range of acceptable values, including a lower limit and an upper limit. The deviation is close to either the lower limit or the upper limit of the range. The aforementioned processor further, Updating the activity model based on the aforementioned deviations. A system configured to perform the following actions.

10. A system, wherein the system is A memory, wherein the memory includes one or more instructions, A processor, wherein the processor is communicatively coupled to the memory, and the processor responds to reading one or more instructions. The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, The system wherein the potential second input includes expected interactions within a predetermined threshold time period, and the deviation is an interaction outside the predetermined threshold time period.

11. A computer program, wherein the computer program is Includes program instructions, and said program instructions, The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, The aforementioned potential second input includes an expected interaction, and the deviation is a computer program in which the expected interaction does not exist.

12. A computer program, wherein the computer program is Includes program instructions, and said program instructions, The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, A computer program in which the potential second input includes a first interaction, and the deviation is the second interaction.

13. A computer program, wherein the computer program is Includes program instructions, and said program instructions, The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, A computer program in which the aforementioned potential second input includes a predetermined range of acceptable values, and the deviation is outside the range.

14. A computer program, wherein the computer program is Includes program instructions, and said program instructions, The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, The aforementioned potential second input includes a predetermined range of acceptable values, including a lower limit and an upper limit. The deviation is close to either the lower limit or the upper limit of the range. The aforementioned program instruction further, Updating the activity model based on the aforementioned deviations. A computer program configured to execute [something].

15. A computer program, wherein the computer program is Includes program instructions, and said program instructions, The first connected device receives one or more user interactions directed to one or more voice control devices in the environment, Based on the user interaction, a first input to the first voice control device of the set of voice control devices is detected, In response to the first input, a potential second input to the set of voice control devices is determined based on the activity model, In response to the first input, the system monitors the user interaction for deviations from the potential second input, Based on the aforementioned monitoring, anomalies in the activity in the environment are identified, In response to an abnormality in the aforementioned activity, corrective actions are performed, It is configured to perform, A computer program in which the potential second input includes expected interactions within a predetermined threshold time period, and the deviation is an interaction outside the predetermined threshold time period.

Citation Information

Patent Citations

  • Speech interaction method and apparatus

    JP2003208196A

  • Anomaly detection for voice controlled devices

    US10706848B1

  • Smart Home Automation Systems and Methods

    US20140108019A1

  • Voice control device and voice control method

    US20140149122A1

  • Follow-up voice query prediction

    US20180012594A1