Reactive Voice Device Management
Patent Information
- Application Number
- JP2024517095
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-07
- Filing Date
- 2022-09-29
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2042-09-29
AI Technical Summary
Voice-based devices often have limited functionality, requiring strict commands and failing to respond to user inputs that deviate from predefined commands, leading to operational difficulties and inefficiencies, especially in dynamic environments where user interactions are unpredictable.
Reactive Voice Device Management (RVM) employs machine learning models to detect user inputs, determine potential subsequent inputs, monitor for deviations, and identify anomalies, enabling corrective actions through natural language processing and artificial intelligence techniques.
RVM enhances the usability of voice-based devices by allowing them to adapt to user interactions beyond predefined commands, improving responsiveness and accuracy in dynamic environments without constant network connectivity.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to voice-based devices, and more particularly, to managing voice-based devices using artificial intelligence techniques. [Background technology]
[0002] A voice-based device may operate based on an audio interface. A voice-based device may receive commands from a user. A voice-based device may operate based on strict commands or only on specific commands. Summary of the Invention
[0003] According to embodiments, a method, a system, and a computer program product are disclosed.
[0004] One or more user interactions directed to a set of one or more voice controlled devices in the environment are received by a first connected device. A first input to a first voice controlled device of the set of voice controlled devices is detected based on the user interaction. A potential second input to the set of voice controlled devices is determined based on the activity model in response to the first input. Deviations from the potential second input are monitored from the user interactions in response to the first input. Anomalies in activity in the environment are identified based on the monitoring. Corrective action is performed in response to the anomalies in activity.
[0005] The above summary is not intended to describe each illustrated embodiment or every embodiment of the present disclosure.
[0006] Preferred embodiments of the invention will now be described, by way of example only, with reference to the following drawings: [Brief description of the drawings]
[0007] [Figure 1]1 illustrates representative major components of an exemplary computer system that may be used in accordance with some embodiments of the present disclosure. [Diagram 2] 1 illustrates a cloud computing environment according to an embodiment of the present invention. [Diagram 3] 1 illustrates an abstraction model layer according to one embodiment of the present invention. [Figure 4] 1 illustrates a representative neural network example of one or more artificial neural networks capable of performing reactive voice device management ("RVM") on one or more voice devices in an environment consistent with embodiments of the present disclosure. [Diagram 5] 1 illustrates an example system for voice device management consistent with certain embodiments of the present disclosure. [Figure 6A] 1 illustrates a first portion of training data for a first machine learning model of a system for identifying anomalous inputs, consistent with certain embodiments of the present disclosure. [Figure 6B] 1 illustrates a second portion of training data for a first machine learning model of a system, consistent with certain embodiments of the present disclosure. [Figure 7] 1 illustrates a portion of training data for a second machine learning model of the system for identifying anomalous inputs, consistent with certain embodiments of the present disclosure. [Figure 8] 13 illustrates a portion of training data for a third machine learning model of the system for identifying anomalous inputs, consistent with certain embodiments of the present disclosure. [Figure 9A] 13 illustrates a first portion of training data for a fourth machine learning model of the system for providing accurate corrective action, consistent with certain embodiments of the present disclosure. [Figure 9B] 13 shows a second portion of training data for a fourth machine learning model of the system, consistent with certain embodiments of the present disclosure. [Figure 10] 1 illustrates a method 1000 for performing corrective action on an audio device, consistent with certain embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] While the invention is amenable to various modifications and alternative forms, specific aspects thereof have been shown by way of example in the drawings and will be described in detail. It should be understood, however, that it is not intended to limit the invention to the particular embodiments described. On the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the scope of the invention.
[0009] Aspects of the present disclosure relate to voice-based devices, and more particularly to managing voice-based devices using artificial intelligence techniques. The present disclosure is not necessarily limited to such applications, but various aspects of the present disclosure may be understood through the discussion of various examples using this context.
[0010] A voice-controlled client device (alternatively, a voice-controlled device, or a voice-based device) ("voice device") may be a computing device that operates based on audio input from a user, such as speech. Voice devices may be increasingly popular due to one or more factors. One factor includes that voice devices may facilitate a user's usage that may be perceived as intuitive. Users may be accustomed to speaking commands or questions to a computer because the voice device may be programmed to respond to similar commands (e.g., "What's the weather like today?"). Another factor is that voice devices may enable multitasking in real-world scenarios (e.g., a user may be able to operate a voice device while their hands are occupied with other tasks). For example, a user may be able to instruct a voice device to turn on the lights while walking their dog.
[0011] Another factor may be that the cost of computer components (memory, processors, voice transceivers, etc.) has decreased to such a level that voice devices are readily available or incorporated into all kinds of devices. These voice devices include computing devices (laptops, desktops, etc.), portable electronic devices (smartphones, tablets, etc.), or wearable client devices (augmented reality glasses, etc.), or a combination thereof. Voice devices may also be home appliances (e.g., smart refrigerators, voice-controlled washing machines, beverage machines with voice-based interfaces, etc.). In some cases, voice devices may include devices that operate solely based on receiving voice or verbal input from a user (e.g., voice-based assistants that have no touch screen, physical buttons, or other ways to receive input other than voice commands).
[0012] Voice devices may have limited functionality in certain scenarios, making it difficult or impossible to operate the voice device. Specifically, a voice device may only operate with rigid or fixed commands. For example, a voice device installed in a home office may respond to "What's the temperature?" but not to "What's the weather right now?" or to "Is it raining here?" Another issue may be that the voice device has limited knowledge about how to respond. Limited knowledge may include the voice device responding to the user by playing a limited range of voice samples or overly general voice samples (e.g., an audio file saying "Invalid command", a sound wave saying "Something's wrong", or a visible message on the screen saying "I can't answer").
[0013] Existing solutions may not be adequate to address all scenarios. One existing solution may be to connect the device to a network such as the Internet and perform additional processing (e.g., by a computer or technical support user). Such a solution may violate the user's privacy, such as monitoring audio input to an audio device in a residential environment. Furthermore, constant monitoring may not be possible at certain times of the day. For example, a user may travel to a remote location with a smartphone acting as an audio device. The smartphone may not have network connectivity and network-based audio listening and processing may not work at all.
[0014] Another existing solution may be to generate a very large selection of commands and responses to try to cover all possible scenarios that may occur with a voice device. Specifically, a voice device may contain tens to hundreds of commands that a user may potentially provide. This existing solution also has drawbacks. For example, the memory or storage subsystem of the voice device may need to be large to provide audio recordings of all potential audible questions and potential answers. Furthermore, because language is not fixed and is constantly evolving, a relatively large data set may not be able to cover future ways of communicating.
[0015] Another drawback is that users are not perfect and may not always be able to operate a voice device in a predictable or understandable manner. Specifically, a given user may forget to send a voice command even though the voice device expects the voice command for correct operation. The user may be busy with other activities and forget to speak the voice command aloud. For example, a user may put tea in a smart microwave. The smart microwave may be a voice device configured to heat the item after receiving a voice command. After placing a tea cup in the smart microwave appliance, the user may typically need to speak a specific command to start heating the tea. The user may also be actively caring for a child and may inadvertently forget to speak a specific verbal command.
[0016] In a second example, a first user may be putting tea in a microwave oven at a residence. The first user may inadvertently utter an unexpected, unwanted, or unusual command. For example, the user may be talking on the phone with another person while trying to heat the tea. The other person on the phone may ask how long the first user plans to take to get ready to go out, and the first user may reply, "I'll be ready to go out in 45 minutes." The microwave may misinterpret the time to heat the tea as 45 minutes.
[0017] Reactive voice device management ("RVM") may operate to perform detection of user inputs and responsively determine potential additional inputs for beneficial processing and operation (e.g., management) of voice devices. The RVM may operate by receiving user interactions directed to one or more voice devices, such as smartphones, voice-based assistants, voice-operated home appliances, etc. In particular, the voice devices may be computers, smart appliances, voice-based assistants, or other relevant voice devices in the user's environment (e.g., home, office, school). The RVM may be configured to process user actions and commands through machine learning models or other relevant artificial intelligence and perform respective corrective actions.
[0018] In particular, the RVM may be configured to determine a user's routine or usage pattern, such as by generating an activity model based on machine learning that takes into account the user's various usage patterns in the environment as well as the voice device that is part of the usage pattern. The routine or pattern may be based on previous usage of the voice device. Furthermore, the RVM may perform corrective actions based on detecting an input to the voice device. Specifically, the RVM may be configured to detect an input to the voice device in the environment, where the input may be part of one or more user interactions received by the connected device. Based on the pre-generated activity model and in response to the detected input, the RVM may determine a potential second input ("second input"). The second input may be an upcoming, future, or predicted input that corresponds to the routine or usage pattern in the environment. The RVM may further monitor the user's additional input as the user performs interactions with the voice device. The RVM may monitor deviations and, based on the monitoring, may identify anomalies in activity in the environment ("anomalies").
[0019] An anomaly may be a missing input from the user (e.g., the user does not speak a corresponding follow-up instruction, the user does not perform a particular follow-up action). An anomaly may be an incorrect input from the user (e.g., the user speaks an incorrect command, the user performs an incorrect physical action). An anomaly may be an input that does not match a pattern, routine, or sequence of inputs that are part of the activity model. Corrective action may include generating a response to the user that includes details of the missing or anomalous input. The generated response may be based on the activity anomaly as well as the pattern, routine, and determined user second input.
[0020] The RVM may operate to overcome one or more of the problems of existing voice controlled devices. First, the RVM may be configured to operate with a connected device. The connected device may be a cloud connected device, such as a server, for processing user interactions. In some embodiments, the connected device may be a dedicated device, such as a computer, configured to perform processing for the RVM. In some embodiments, the connected device may be one of the voice devices that is also configured to receive voice input from a user and respond with voice output. By operating locally, the RVM may enable local processing of voice data without connecting to an external server.
[0021] Another advantage is that the RVM may operate without a fixed set of commands and responses. Specifically, the RVM may implement one or more artificial intelligence techniques (e.g., machine learning) to determine routines, activities, or patterns. Because of machine learning, the RVM may be able to detect certain anomalies regarding a user's interaction with voice, even if the particular interaction was not previously part of the anomaly list. Additionally, the RVM may be able to generate verbal responses to the user, even if the words or phrases were not originally part of the device's stored vocabulary.
[0022] Additionally, the RVM is useful as the user continues to use the voice device in more sophisticated ways. In particular, as the user uses the voice device in an environment, the user may begin to use the device in more elaborate ways or in concert with other voice devices. For example, as the user begins to use the voice-enabled coffee maker in the morning, the RVM may receive training data including the time and settings of the coffee maker. As the user consistently wakes up at the same time and performs the same actions, the RVM may update the activity model to include this data. Later (e.g., days, weeks, months), the user may also place a smart lamp in the same room as the voice-enabled coffee maker and begin to use voice commands or physical inputs. Usage of the smart lamp may also be input into the RVM, and the activity model may include usage information for both devices. As the user subsequently performs user interactions directed toward either the smart lamp or the voice-enabled coffee maker, each device may be monitored for potential second inputs. Monitoring for potential inputs would monitor for deviations from the patterns created in the activity models corresponding to both voice devices. As a result, the RVM can compare user interaction inputs to the trained activity model, allowing the RVM to identify anomalies not just to one of the voice devices, but to either the smart lamp, the voice-enabled coffee maker, or both.
[0023] FIG. 1 illustrates representative major components of an exemplary computer system 100 (alternatively, a computer) that may be used in accordance with some embodiments of the present disclosure. It is understood that the individual components may vary in complexity, number, type, or configuration, or combinations thereof. The specific example disclosed is for illustrative purposes and not necessarily the only such variation. The computer system 100 may include a processor 110, a memory 120, an input / output interface (here I / O or I / O interface) 130, and a main bus 140. The main bus 140 may provide a communication path to other components of the computer system 100. In some embodiments, the main bus 140 may be connected to other components, such as a dedicated digital signal processor (not shown).
[0024] The processor 110 of the computer system 100 may be comprised of one or more cores 112A, 112B, 112C, 112D (collectively 112). The processor 110 may further include one or more memory buffers or caches (not shown) that provide temporary storage of instructions and data for the cores 112. The cores 112 may execute instructions on inputs provided from the cache or memory 120 and output results to the cache or memory. The cores 112 may be comprised of one or more circuits configured to perform one or more methods consistent with embodiments of the present disclosure. In some embodiments, the computer system 100 may include multiple processors 110. In some embodiments, the computer system 100 may be a single processor 110 with a single core 112.
[0025] The memory 120 of the computer system 100 may include a memory controller 122. In some embodiments, the memory 120 may include a random access semiconductor memory, storage device, or storage medium (either volatile or non-volatile) for storing data and programs. In some embodiments, the memory may be in the form of a module (e.g., a dual in-line memory module). The memory controller 122 may communicate with the processor 110 and facilitate the storage and retrieval of information in the memory 120. The memory controller 122 may communicate with the I / O interface 130 and facilitate the storage and retrieval of inputs or outputs in the memory 120.
[0026] I / O interface 130 may include I / O bus 150, terminal interface 152, storage interface 154, I / O device interface 156, and network interface 158. I / O interface 130 may connect main bus 140 to I / O bus 150. I / O interface 130 may direct instructions and data from processor 110 and memory 120 to various interfaces of I / O bus 150. I / O interface 130 may also direct instructions and data from various interfaces of I / O bus 150 to processor 110 and memory 120. The various interfaces may include terminal interface 152, storage interface 154, I / O device interface 156, and network interface 158. In some embodiments, the various interfaces may include a subset of the aforementioned interfaces (e.g., an embedded computer system for industrial applications may not include terminal interface 152 and storage interface 154).
[0027] Logic modules throughout computer system 100 (including, but not limited to, memory 120, processor 110, and I / O interface 130) may communicate faults and changes to one or more components to a hypervisor or operating system (not shown). The hypervisor or operating system may allocate the various resources available to computer system 100 and track the location of data in memory 120 and processes assigned to the various cores 112. In embodiments that combine or rearrange elements, aspects and capabilities of the logic modules may be combined or redistributed. These variations will be apparent to those skilled in the art.
[0028] Although the present disclosure includes detailed descriptions of cloud computing, implementation of the teachings described herein is not limited to a cloud computing environment. Rather, embodiments of the present invention may be practiced with any other type of computing environment now known or developed in the future. Cloud computing is a model of service delivery to enable convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.
[0029] The characteristics are as follows:
[0030] On-Demand Self-Service: Cloud consumers can unilaterally provision computing capacity, such as server time or network storage, automatically as needed, without the need for human interaction with the service provider.
[0031] Broad network access: Computing power is available over the network and can be accessed through standard mechanisms, facilitating use by heterogeneous thin- and thick-client platforms (e.g., cell phones, laptops, PDAs).
[0032] Resource Pooling: Computing resources from a provider are pooled and offered to multiple consumers using a multi-tenant model. Different physical and virtual resources are dynamically allocated and reallocated depending on demand. Consumers generally have no control or knowledge of the exact location of the resources provided to them, so there is a sense of location independence, although consumers may be able to determine location at a higher level of abstraction (e.g. country, state, data center).
[0033] Rapid elasticity: Computing capacity can be provisioned quickly and elastically, in some cases automatically, to scale out immediately and to be released quickly to scale in immediately. To the consumer, the computing capacity available for provisioning often appears unlimited and can be purchased at any time and in any quantity.
[0034] Metered Services: Cloud systems leverage metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported to provide transparency to both providers and consumers of the services utilized.
[0035] The service model is as follows:
[0036] Software as a Service (SaaS): The functionality offered to the consumer is the availability of a provider's applications running on a cloud infrastructure that can be accessed from a variety of client devices through a thin-client interface such as a web browser (e.g., webmail). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even the individual application functions, except for limited user-specific application configuration settings.
[0037] Platform as a Service (PaaS): The capability offered to the consumer is to deploy applications that the consumer creates or acquires using programming languages and tools supported by the provider onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of the hosting environment.
[0038] Infrastructure as a Service (IaaS): The functionality offered to the consumer is the provision of processors, storage, networking, and other basic computing resources onto which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating systems, storage, and deployed applications, and may have partial control over some network components (e.g., host firewalls).
[0039] The deployment model is as follows:
[0040] Private Cloud: The cloud infrastructure is dedicated to a specific organization. It can be managed by that organization or a third party and can exist on-premise or off-premise.
[0041] Community Cloud: The cloud infrastructure is shared by multiple organizations to support a specific community with common concerns (e.g., mission, security requirements, policies, and compliance). The cloud infrastructure can be managed by the organizations or a third party and can exist on-premise or off-premise.
[0042] Public cloud: The cloud infrastructure is available to the general public or large industry organizations and is owned by an organization that sells cloud services.
[0043] Hybrid Cloud: This cloud infrastructure combines two or more cloud models (private, community or public), each of which retains its inherent semantics but is bound together by standards or specific technologies that enable data and application portability (e.g. cloud bursting for load balancing between clouds).
[0044] A cloud computing environment is a service-oriented environment with an emphasis on statelessness, low coupling, modularity and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0045] Referring to FIG. 1, an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10, to which local computing devices used by cloud consumers (e.g., personal digital assistants (PDAs) or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automobile computer systems 54N, or combinations thereof) can communicate. The nodes 10 can communicate with each other. The nodes 10 can be physically or virtually grouped (not shown) in one or more networks, such as, for example, a private, community, public, or hybrid cloud, or combinations thereof, as described above. This allows the cloud computing environment 50 to provide infrastructure, platform, or software, or combinations thereof, as a service, for which the cloud consumer does not need to maintain resources on the local computing device. It should be understood that the types of computing devices 54A-N shown in FIG. 1 are merely exemplary, and that the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic device over any type of network or network addressable connection (e.g., using a web browser), or both.
[0046] Referring to Figure 2, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 1) is shown. It should be understood in advance that the components, layers and functions shown in Figure 2 are merely exemplary and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0047] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframes 61, reduced instruction set computer (RISC) architecture based servers 62, servers 63, blade servers 64, storage devices 65, and networks and network components 66. In some embodiments, the software components include network application server software 67 and database software 68. The virtualization layer 70 provides an abstraction layer from which the following virtual entities can be provided, for example: virtual servers 71, virtual storage 72, virtual networks 73, including virtual private networks, virtual applications and operating systems 74, and virtual clients 75.
[0048] By way of example, the management layer 80 may provide the following functions: Resource provisioning 81 allows dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 allows cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. By way of example, these resources may include application software licenses. Security allows for identification and verification of cloud consumers and tasks as well as protection for data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 allows for allocation and management of cloud computing resources such that requested service levels are met. Service level agreement (SLA) planning and fulfillment 85 allows for pre-arranging and procurement of anticipated future needs for cloud computing resources according to SLAs.
[0049] The workload layer 90 provides examples of functionality available to a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and RVM 96.
[0050] FIG. 4 illustrates an example of a representative neural network (or "network") 400 of one or more artificial neural networks capable of executing RVM on one or more audio devices in an environment consistent with embodiments of the present disclosure. The exemplary neural network 400 is composed of multiple layers. The network 400 includes an input layer 410, a hidden section 420, and an output layer 450. Although the network 400 illustrates a feed-forward neural network, other neural network layouts may be contemplated, such as a recurrent neural network layout (not shown). In some embodiments, the network 400 may be a design-and-run neural network, where the illustrated layout may be created by a computer programmer. In some embodiments, the network 400 may be a design-by-run neural network, where the illustrated layout may be generated by a process of inputting data and analyzing that data according to one or more defined heuristics. The network 400 may operate in a forward propagation manner by receiving inputs and outputting the results of the inputs. The network 400 may adjust the values of various components of the neural network by backward propagation (backpropagation).
[0051] The input layer 410 includes a set of input neurons 412-1, 412-2, through 412-n (collectively, 412) and a set of input connections 414-1, 414-2, 414-3, 414-4, etc. (collectively, 414). The input layer 410 represents inputs from the data that the neural network is to analyze (e.g., a set of voice inputs, physical actions, associated timestamps, expected routines, patterns, etc.). Each input neuron 412 may represent a subset of the input data. For example, the neural network 400 is provided with various numerical representations as inputs, where voice device information and input from a user are represented by numerical representations.
[0052] In another example, input neuron 412-1 may be the first pixel of a picture, input neuron 412-2 may be the second pixel of the picture, etc. The number of input neurons 412 may correspond to the size of the input. For example, if neural network 400 is designed to analyze a 256 pixel by 256 pixel image, the neural network layout may include a set of 65,536 input neurons. The number of input neurons 412 may correspond to the type of input. For example, if the input is a 256 pixel by 256 pixel color image, the neural network layout may include a set of 196,608 input neurons (65,536 [256 2] input neurons). The type of input neurons 412 may correspond to the type of input. In a first example, the neural network may analyze a black and white image, and each of the input neurons may be a decimal value between 0.00001 and 1 that represents the grayscale shade of the pixel (where 0.00001 represents a completely white pixel and 1 represents a completely black pixel). In a second example, the neural network may analyze a color image, and each of the input neurons may be represented as a three-dimensional vector of the color value of a given pixel in the input image. A first component of the vector may be an integer value of red between 0 and 255, a second component may be an integer value of green between 0 and 255, and a third component of the vector may be an integer value of blue between 0 and 255.
[0053] The input connections 414 represent the outputs of the input neurons 412 to the hidden section 420. Each of the input connections 414 varies based on a number of weights (not shown) depending on the value of each input neuron 412. For example, a first input connection 414-1 has a value provided to the hidden section 420 based on input neuron 412-1 and a first weight. Continuing with the example, a second input connection 414-2 has a value provided to the hidden section 420 based on input neuron 412-1 and a second weight. Continuing with the example, a third input connection 414-3 is based on input neuron 412-2 and a third weight, and so on. Alternatively, input connections 414-1 and 414-2 may share the same output component of input neuron 412-1, input connections 414-3 and 414-4 may share the same output component of input neuron 412-2, and all four input connections 414-1, 414-2, 414-3, and 414-4 may have output components with four different weights. Network neural 400 may have different weightings for each connection 414, although some embodiments may contemplate similar weights. In some embodiments, each value of input neurons 412 and connections 414 may be stored in memory.
[0054] The hidden section 420 includes one or more layers that receive inputs and generate outputs. The hidden section 420 includes a first hidden layer of computational neurons 422-1, 422-2, 422-3, 422-4, up to 422-n (collectively 422), a second hidden layer of computational neurons 426-1, 426-2, 426-3, 426-4, 426-5, up to 426-n (collectively 426), and a set of hidden connections 424 that connect the first and second hidden layers. It should be understood that the neural network 400 illustrates only one of many neural networks that can monitor inputs from user interactions with a voice device consistent with some embodiments of the present disclosure. As a result, in some embodiments of the present disclosure, an exemplary hidden section may include more or fewer hidden layers than the two hidden layers described with respect to the hidden section 420 (e.g., one hidden layer, seven hidden layers, twelve hidden layers, etc.).
[0055] The first hidden layer 422 includes computational neurons 422-1, 422-2, 422-3, 422-4, through 422-n. Each computational neuron in the first hidden layer 422 can receive one or more of the connections 414 as inputs. For example, computational neuron 422-1 receives input connection 414-1 and input connection 414-2. Each computational neuron in the first hidden layer 422 also provides an output. The output is represented by the dotted hidden connections 424 flowing out of the first hidden layer 422. Each of the computational neurons 422 performs an activation function during forward propagation. In some embodiments, the activation function may be a process (e.g., a perceptron) that receives multiple binary inputs and computes a single binary output. In some embodiments, the activation function may be a process that receives several non-binary inputs (e.g., a number between 0 and 1, 0.671, etc.) and computes a single non-binary output (e.g., a number between 0 and 1, a number between -0.5 and 0.5, etc.). Various functions may be implemented to compute the activation functions (e.g., sigmoid neurons or other logistic functions, tanh neurons, softplus functions, softmax functions, rectified linear units, etc.). In some embodiments, each of the computational neurons 422 also includes a bias (not shown). The bias may be used to determine the likelihood or evaluation of a given activation function. In some embodiments, each of the values of the bias for each of the computational neurons must necessarily be stored in memory.
[0056] The neural network 400 may include using a sigmoid neuron for the activation function of the computational neuron 422-1. An equation (Equation 1 below) may represent the activation function of the computational neuron 422-1 as f(neuron). The logic of the computational neuron 422-1 may be the sum of each of the input connections (i.e., input connections 414-1 and input connections 414-3) that feed into the computational neuron 422-1, represented as j in Equation 1. For each j, a weight w is multiplied by the value x of a given connected input neuron 412. The bias of the computational neuron 422-1 is represented as b. As each input connection j is summed, a bias b is subtracted. In this example, the output of computational neuron 422-1 is determined according to the following: given a large positive number resulting from the sum of activations f (neuron) and the bias, the output of computational neuron 422-1 will be close to 1; given a large negative number resulting from the sum of activations f (neuron) and the bias, the output of computational neuron 422-1 will be close to 0; given a number intermediate between a large positive number and a large negative number resulting from the sum of activations f (neuron) and the bias, the output will change slightly as the weights and biases change slightly.
[0057]
number
[0058] The second hidden layer 426 includes computational neurons 426-1, 426-2, 426-3, 426-4, 426-5, through 426-n. In some embodiments, the computational neurons of the second hidden layer 426 may operate similarly to the computational neurons of the first hidden layer 422. For example, the computational neurons 426-1 through 426-n may operate with similar activation functions as the computational neurons 422-1 through 422-n, respectively. In some embodiments, the computational neurons of the second hidden layer 426 may operate differently than the computational neurons of the first hidden layer 422. For example, the computational neurons 426-1 through 426-n may have a first activation function, and the computational neurons 422-1 through 422-n may have a second activation function.
[0059] Similarly, the connectivity to, from, and between the various layers of the hidden section 420 may vary. For example, the input connections 414 may be fully connected to the first hidden layer 422, and the hidden connections 424 may be fully connected from the first hidden layer to the second hidden layer 426. In some embodiments, fully connected may mean that each neuron in a given layer may be connected to every neuron in the previous layer. In some embodiments, fully connected may mean that each neuron in a given layer functions completely independently and does not share connections. In the second example, the input connections 414 may not be fully connected to the first hidden layer 422, and the hidden connections 424 may not be fully connected from the first hidden layer to the second hidden layer 426.
[0060] Additionally, parameters to, from, and between various layers of the hidden section 420 may vary. In some embodiments, the parameters may include weights and biases. In some embodiments, the parameters may be more or less than weights and biases. In some embodiments of the present disclosure, the network 400 may be a convolutional neural network or a convolutional network. A convolutional neural network may include a series of heterogeneous layers (e.g., an input layer 410, a convolutional layer 422, a pooling layer 426, and an output layer 450). In such a network, the input layer 410 may hold raw audio data of the audio input with onset time, sample length, and a two-dimensional volume of microphone sources. The convolutional layers of such a network may output from connections that exist only locally in the input layer to identify features of a small section of an image (e.g., the onset of a verbal command from a user, the length of time it takes the user to provide a verbal command, etc.). Considering this example, the convolutional layers may include weights and biases as well as additional parameters (e.g., depth, stride, and padding). The pooling layer 426 of such a network takes as input the output of the convolutional layer 422, but may perform a fixed function operation (e.g., an operation that does not consider weights or biases). Also, considering this example, the pooling layer may not include any convolution parameters, and may not include any weights or biases (e.g., perform a downsampling operation).
[0061] The output layer 450 includes a set of output neurons 450-1, 450-2, 450-3, through 450-n (collectively 450). The output layer 450 holds the analysis results of the neural network 400. In some embodiments, the output layer 450 may be a classification layer used to identify features of the inputs to the network 400. For example, the network 400 may be a classification network trained to identify Arabic numerals. In such an example, the network 400 may include an output layer 450 of 10 output neurons corresponding to which Arabic numerals the network has identified (e.g., output neuron 450-2 having a higher activation value than output neuron 450 may indicate that the neural network has determined that the image contains the numeral "1"). In some embodiments, the output layer 450 may be a real-valued target (e.g., attempting to predict an outcome when the input is a set of previous outcomes) or may have a specific output neuron (not shown). The output layer 450 is fed by output connections 452. The output connections 452 provide activations from the hidden section 420. In some embodiments, the output connections 452 may include weights and the neurons of the output layer 450 may include biases.
[0062] Training a neural network depicted by neural network 400 may include performing backpropagation. Backpropagation is distinct from forward propagation. Forward propagation may include feeding data to input neurons 410, performing computations on connections 414, 424, 452, and performing computations on computational neurons 422, 426. The term forward propagation may also be used to describe the layout of a given neural network (e.g., recursion, number of layers, number of neurons in one or more layers, whether layers are fully connected or unconnected to other layers, etc.).
[0063] In contrast, backpropagation can determine errors in the parameters (e.g., weights and biases) in the network 400 by starting from the output neuron 450 and backpropagating the error through the various connections 452, 424, 414 and layers 426, 422, respectively.
[0064] Backpropagation may involve running one or more algorithms based on one or more training data to reduce the difference between what a given neural network decides from the inputs and what the given neural network should decide from the inputs. The difference between the network's decision and the correct decision may be called the objective function (alternatively, the cost function). When a given neural network is first created, provided with data, and calculated by forward propagation, the result or decision may be an incorrect decision.
[0065] For example, the neural network 400 may be a classification network. Further, the network 400 may receive a verbal input, such as a command directed to a first voice device of the five voice devices. Further, the network 400 may determine that the verbal command is relatively most likely directed to a third voice device of the five voice devices. Further, the network 400 may determine that the next most likely voice device is the fifth voice device, and the next most likely voice device is the first voice device (and similarly for other inputs). Continuing with the example, backpropagation may change the values of the weights of the connections 414, 424, and 452, and may change the values of the biases of the first layer of computational neurons 422, the second layer of computational neurons 426, and the output neuron 450. Continuing with the example, backpropagation may result in future results that are relatively accurate classifications of the same voice inputs that include the same verbal command (e.g., ranking a verbal command directed to a first voice device of the five voice devices as a relatively more likely voice device).
[0066] Equation 2 provides an example of an objective function ("example function") in the form of a quadratic cost function (e.g., mean squared error), although other functions may be selected, and mean squared error is selected for illustrative purposes. In Equation 2, the weights may all be represented by w, and the biases may be represented by b for neural network 400. Network 400 is provided with a predetermined number of training inputs n in a subset (or the entirety) of training data having input values x. Network 400 should be able to obtain an output a from x, and obtain a desired output y(x) from x. Backpropagation or training of network 400 should be a reduction or minimization of the objective function "O(w,b)" via modification of a set of weights and biases. Successful training of network 400 should include not only a reduction in the difference between the answer a for an input value x and the correct answer y(x), but also when new input values (e.g., from additional training data, from validation data, etc.) are presented.
[0067]
number
[0068] Backpropagation algorithms may utilize many options in both objective function (e.g., mean squared error, cross-entropy cost function, accuracy function, confusion matrix, precision-recall curve, mean absolute error, etc.) and reduction of objective function (e.g., gradient descent, batch-based stochastic gradient descent, Hessian optimization, momentum-based gradient descent, etc.). Backpropagation may include using a gradient descent algorithm (e.g., computing partial derivatives of the objective function with respect to weights and biases with respect to all training data). Backpropagation may include determining stochastic gradient descent (e.g., computing partial derivatives with respect to a subset of training data or a subset of training inputs in a batch). Various backpropagation algorithms may also include additional parameters (e.g., learning rate for gradient descent). Large changes to weights and biases through backpropagation may lead to erroneous training (e.g., overfitting to training data, reducing to a local minimum, reducing too far past a global minimum, etc.). Therefore, modifications to the objective function with more parameters can prevent erroneous training (e.g., utilizing an objective function that incorporates regularization to prevent overfitting). Also, as a result, the changes to neural network 400 may be small at any given iteration, and the backpropagation algorithm may run many iterations to perform more accurate learning as a result of the relative smallness at any given iteration.
[0069] For example, the neural network 400 may have untrained weights and biases, and the backpropagation may include stochastic gradient descent to train the network on a subset of the training inputs (e.g., a batch of 10 training inputs from the entire set of training inputs). Continuing with the example, the network 400 may continue training using a second subset of the training inputs (e.g., a second batch of 10 training inputs from the entire set outside the first batch), which may be repeated until all of the training inputs have been used to compute the gradient descent (e.g., one epoch of training data). Stated differently, if there are a total of 10,000 training images and one iterative training uses a batch size of 100 training inputs, then 1,000 iterative trainings will be required to complete one epoch of training data. Many epochs may be performed to continue training the neural network. There may be many factors that determine the selection of additional parameters (e.g., a large batch size may result in inadequate training, a small batch size may result in too many training iterations, a large batch size may not fit in memory, a small batch size may not efficiently utilize discrete GPU hardware, too few training epochs may not result in a fully trained network, too many training epochs may result in overfitting the trained network, etc.). Additionally, network 400 may be evaluated to quantify its performance on a dataset, such as by using evaluation metrics (e.g., mean squared error, cross-entropy cost function, accuracy function, confusion matrix, precision-recall curve, mean absolute error, etc.).
[0070] 5 illustrates an example system 500 for voice device management consistent with certain embodiments of the present disclosure. System 500 may be configured to perform RVM, which includes monitoring user activity during interaction with a voice-enabled or voice-controlled computing device. System 500 may include at least a network 510, one or more network-connected devices 520-1, 520-2, 520-3, and 520-4, and an RVM 530.
[0071] Network 510 may be implemented using any number of suitable physical and / or logical communication topologies. Network 510 may include one or more private or public computing networks. For example, network 510 may include a private network associated with the workload (e.g., a network having a firewall that blocks unauthorized external access). Alternatively, or in addition, network 510 may include a public network, such as the Internet. Thus, network 510 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet, or a combination thereof. Network 510 may include one or more servers, networks, or databases, and may transfer data between elements of system 500 using one or more communication protocols. Additionally, while illustrated as a single entity in FIG. 5, in other examples network 510 may include multiple networks, such as a combination of public and / or private networks. Communication network 510 may include various types of physical communication channels or "links." Links may be wired, wireless, optical, or any other suitable medium, or combinations thereof. Additionally, the communications network 510 may include various network hardware and software for performing routing, switching, and other functions, such as routers, switches, base stations, bridges, or any other equipment that may be useful in facilitating data communications.
[0072] The connected device 520 may be a computing device configured to perform one or more operations of the RVM 530. The connected device 520 may be a computer system, such as computer 100. The connected device 520 may be a cloud computing system. For example, the connected device 520-4 may be representative of one or more servers that make up the cloud computing environment 50. The connected device 520 may be a voice device. For example, the connected device 520-1 may be a voice-operated smartphone configured to receive and respond to verbal commands. The connected device 520 may include a home appliance configured to perform routine tasks in response to input from a user. For example, the connected device 520-2 may be a clothes washer and the connected device 520-3 may be a clothes dryer. A user may perform an operation on the clothes washer 520-2 or the clothes dryer 520-3 or both, such as physically placing dirty clothes into the clothes washer. The user can also issue voice commands to the clothes washer 520-2 or the clothes dryer 520-3 or both, for example, saying "set mode to 3" to refer to washing clothes under the third of five different clothes wash settings.
[0073] RVM 530 may include one or more operations of a configured set of software and / or hardware for performing voice device management. RVM 530 may be in the form of multiple artificial intelligence components, such as machine learning models ("ML models"). Each of the ML models may be an instance of a neural network, such as neural network 400. The ML models may include a detection model 540, a missing input model 550, an erroneous input model 560, and a corrective action model 570. RVM 530 may also include an activity model 532, and a natural language processor 534 ("NLP").
[0074] In some embodiments, the detection model 540, the missing input model 550, the erroneous input model 560, and the corrective action model 570 may perform machine learning on the activity model 532 using one or more of the following example techniques: K-nearest neighbors (KNN), learning vector quantization (LVQ), self-organizing maps (SOM), logistic regression, ordinary least squares regression (OLSR), linear regression, stepwise regression, multivariate adaptive regression splines (MARS), ridge regression, least absolute shrinkage selection operator (LSSR), or a combination of both. LASSO), Elastic Net, Least Angle Regression (LARS), Probabilistic Classifier, Naive Bayes Classifier, Binary Classifier, Linear Classifier, Hierarchical Classifier, Canonical Correlation Analysis (CCA), Factor Analysis, Independent Component Analysis (ICA), Linear Discriminant Analysis (LDA), Multidimensional Scaling (MDS), Non-negative Metric Decomposition (NMF), Partial Least Squares Regression (PLSR), Principal Component Analysis (PCA), Principal Component Regression (PCR), Sammon Mapping, t-SNE, Bootstrap Tabulation, Ensemble Average, Gradient Boosting Decision Tree (G BRT), Gradient Boosting Machine (GBM), Inductive Bias Algorithm, Q-Learning, State-Action-Reward-State-Action (SARSA), Temporal Difference (TD) Learning, A Priori Algorithm, Equivalence Class Transformation (ECLAT) Algorithm, Gaussian Process Regression, Gene Expression Programming, Group Method of Data Processing (GMDH), Inductive Logic Programming, Instance-Based Learning, Logistic Model Trees, Information Fuzzy Networks (IFN), Hidden Markov Models, Gaussian Naive Bayes, Polynomial Naive Bayes, Averaged One-Dependence Estimator (AODE), Bayesian Networks (BN), Classification and Regression Trees (CART), Chi-Square Automatic Interaction Detection (CHAID), Expectation Maximization Algorithm, Feedforward Neural Networks, Logic Learning Machines, Self-Organizing Maps, Single Linkage Clustering, Fuzzy Clustering, Hierarchical Clustering, Boltzmann Machines, Convolutional Neural Networks, Recurrent Neural Networks, Hierarchical Temporal Memory (HTM), or other machine learning techniques, or a combination thereof.
[0075] The ML model of the RVM 530 may be configured to process language using artificial intelligence techniques such as natural language processing using a natural language processor 534. In some embodiments, the natural language processor 534 may include various components (not shown) operating in hardware, software, or some combination. For example, the natural language processor 534 may include one or more data sources, a search application, and a report analyzer. The natural language processor 534 may be a computer module that analyzes received content and other information. The natural language processor 534 may perform various methods and techniques for analyzing text information (e.g., syntactic analysis, semantic analysis, etc.). The natural language processor 534 may be configured to recognize and analyze any number of natural languages. In some embodiments, the natural language processor 534 may parse a passage of a document or content from speech received from a user. The various components (not shown) of the natural language processor 534 include, but are not limited to, a tokenizer, a part of speech (POS) tagger, a semantic relation identifier, and a syntactic relation identifier. The natural language processor 534 may include a support vector machine (SVM) generator to allow the processor 534 to process topical content found in the corpus and classify topics.
[0076] In some embodiments, the tokenizer may be a computer module that performs lexical analysis. The tokenizer may convert a sequence of characters into a sequence of tokens. A token may be a string of characters that is included in an electronic document and classified as a meaningful symbol. Additionally, in some embodiments, the tokenizer may identify word boundaries within an electronic document and divide any text passage within the document into constituent text elements, such as words, multi-word tokens, numbers, and punctuation marks. In some embodiments, the tokenizer may receive a string of characters, identify vocabulary within the string, and classify them into tokens.
[0077] Consistent with various embodiments, a POS tagger may be a computer module that marks up words in a passage to correspond to a particular part of speech. A POS tagger can read a sentence or other text written in a natural language and assign a part of speech to each word or other token. A POS tagger can determine the part of speech to which a word (or other text element) corresponds based on the word's definition and the word's context. A word's context may be based on its relationship to adjacent or related words in a phrase, sentence, or paragraph.
[0078] In some embodiments, the context of a word may depend on one or more previously analyzed electronic documents (e.g., previous voice-activated routines, previous usage patterns including voice utterances, information stored in the activity model 532). Examples of parts of speech that may be assigned to a word include, but are not limited to, noun, verb, adjective, adverb, etc. Examples of other part-of-speech categories that the POS tagger may assign include, but are not limited to, comparative adverbs, superordinate adverbs, WH adverbs, conjunctions, determiners, negation particles, possessive markers, prepositions, WH pronouns, etc. In some embodiments, the POS tagger may tag or annotate tokens of a passage with part-of-speech categories. In some embodiments, the POS tagger may set tags on tokens or words of a passage that are analyzed by a natural language processing system.
[0079] In some embodiments, the semantic relationship identifier may be a computer module that may be configured to identify semantic relationships of recognized text elements (e.g., words, phrases) in a document. In some embodiments, the semantic relationship identifier may determine functional dependencies and other semantic relationships between entities.
[0080] Consistent with various embodiments, the syntactic relation identifier may be a computer module that may be configured to identify syntactic relations within a passage composed of tokens. The syntactic relation identifier may determine the grammatical structure of a sentence, such as, for example, which groups of words are related as phrases and which words are subjects or objects of verbs. The syntactic relation identifier may conform to a formal grammar.
[0081] In some embodiments, the natural language processor 534 may be a computer module that may parse a document and generate a corresponding data structure for one or more portions of the document. For example, in response to receiving speech received from a user in a natural language processing system, the natural language processor 534 may output parsed text elements from the data. In some embodiments, the parsed text elements may be represented in the form of a parse tree or other graph structure. To generate the parsed text elements, the natural language processor 534 may trigger computer modules including a tokenizer, a part-of-speech (POS) tagger, an SVM generator, a semantic relation identifier, and a syntactic relation identifier.
[0082] In some embodiments, the natural language processing system may leverage one or more of the exemplary machine learning techniques to perform machine learning (ML) text operations. In particular, the RVM 530 may operate to perform machine learning text classification or machine learning text comparison or both. Machine learning text classification may include ML text operations that convert characters, text, words, and phrases into numerical values. The numerical values may then be input into a neural network to determine various features, characteristics, and other information of the words with respect to the document or in relation to other words (e.g., classifying the numerical values associated with the words may enable classification of the words). Machine learning text comparison may include using the numerical values of the converted characters, text, words, and phrases to perform the comparison. The comparison may be a comparison of the numerical value of a first word or other text with the numerical value of a second word or other text. The determination of the machine learning text comparison may be a scoring, correlation, or determining an association relationship (e.g., a relationship between a first numerical value of a first word and a second numerical value of a second word). The comparison is used to determine whether two words are similar or different based on one or more criteria. The numerical operations for machine learning text classification / comparison may be a function of mathematical operations performed through neural networks, such as performing linear regression, addition, or other related mathematical operations on numbers representative of words or other text.
[0083] ML text manipulation can include encoding of words, such as one-hot encoding of words from a tokenizer, a POS tagger, a semantic relation identifier, a syntactic relation identifier, etc. ML text manipulation can include use of vectorization of text, such as vectorization of words from a tokenizer, a POS tagger, a semantic relation identifier, a syntactic relation identifier, etc. For example, a paragraph of text may include the phrase "oranges are fruits that grow on trees." Vectorization of the word "orange" may include setting input neurons of a neural network to various words of the phrase that includes the word "orange." The output value may be an array of values (e.g., 48 digits, thousands of digits). The output value may tend toward "1" for related words and toward "0" for unrelated words. Related words may be associated based on one or more of similar parts of speech, syntactic meaning, locality within a sentence or paragraph, or other associated "closeness" between the input and other parts of the natural language (e.g., other parts of the phrase "oranges are fruits that grow on trees," other parts of the paragraph that includes the phrase, other parts of the language).
[0084] In some embodiments, each of the ML models (e.g., detection model 540, missing input model 550, erroneous input model 560, and corrective action model 570) may be configured to perform a different action.
[0085] The detection model 540 may be a fully connected neural network and may be configured to receive as inputs the current time and a user interaction from the user (e.g., physical activity, voice command). The detection model 540 may include various outputs, such as the name of the particular connected device 520 and whether a user interaction was expected but not received. The detection model 540 may also perform operations to determine conditions such as whether an expected second user interaction or a follow-up user interaction was expected (e.g., a softmax operation). For example, at 7:00 a.m., when any voice device is activated by switching on any voice device or by placing any object such as a coffee cup under the spout of a coffee maker, the detection model 540 determines whether those actions, devices, or times, or combinations thereof, are part of the user's regular routine, pattern, or usage scenario.
[0086] The missing input model 550 may be a recurrent neural network configured to operate as an encoder / decoder. The missing input model 550 may receive as input a previous user interaction (e.g., an interaction received by the detection model 540). The missing input model 550 may also receive a previous timestamp (e.g., a timestamp of a previous user interaction received by the detection model 540), and a current timestamp (e.g., a time after the previous timestamp). The missing input model 550 may output a predicted time of a follow-up or second command to output a time of a predicted follow-up or second command. This may be considered as a determination of a potential second input. The missing input model 550 may perform a classifier operation. The classifier operation may include monitoring for deviations from a potential second input and identifying activity anomalies. For example, the classifier operation may identify or predict activity anomalies including loss or missing of a predicted command. The missing input model 550 may output the name of the particular connected device 520 and an identified value (e.g., "yes," "no") that the predicted command is missing. For example, if the first command was "grill at 300 degrees for 10 minutes," the missing input model 550 may begin determining the next most likely command and time. The next most likely command and length of time might be to reduce the temperature to 200 degrees and grill for 10 minutes. Further, the missing input model 550 may determine that this next most likely command is expected within 20 minutes of receiving the first command.
[0087] The false input model 560 may be a recurrent neural network configured to function as an encoder / decoder, such as a long short-term memory. The false input model 560 may also include a dense layer or a fully connected layer or both. The recurrent portion of the false input model 560 may pass data to the fully connected layer. The false input model 560 may also include an encoder that creates a latent space and a decoder (e.g., a variational autoencoder ("VAE")) that uses the latent space to generate an output. The latent space created by the encoder may be in the form of two distributions: a mean distribution and a covariance distribution. As a result, the false input model 560 may be configured to analyze all inputs and generate as output a Gaussian distribution of the mean of all inputs and a Gaussian distribution of the standard deviation of all inputs. The false input model 560 may be configured to score the output and determine whether the score exceeds a predefined threshold. This may be considered as determining a potential second input and / or monitoring deviations from a potential second input and identifying activity anomalies. If the score exceeds a predefined threshold, an anomalous input may be present in the user interaction. For example, the error model 560 may determine that a deviation level of a command related to grilling on a stove exceeds a predetermined threshold. The expected command may be "grill for 10 more minutes," but the actual command received is "grill for 15 more minutes," and the error model 560 may generate a score below the predetermined threshold associated with the particular voice device and / or routine that includes the particular voice device. If the actual command received is "grill for 40 more minutes," the error model 560 may generate a score above the predetermined threshold associated with the particular voice device and / or routine.
[0088] The corrective action model 570 may be a recurrent neural network configured to perform encoding and / or decoding. The corrective action model 570 may receive inputs from the detection model 540, the missing input model 550, and the erroneous input model 560. The corrective action model 570 may also output a corrective action to a user.
[0089] The ML models of the RVM 530 may cooperate to receive user interactions (e.g., physical activities and voice commands) with various audio devices in the environment to identify anomalous activity (e.g., input that differs from expected input). Specifically, user interactions may be received at 580 from the connected device 520. The user interactions may include voice commands such as "turn on the coffee maker" via voice or "run the laundry for 10 minutes" via speech. The corrective action model 570 may generate corrective actions at 590 and provide these corrective actions to the user. The corrective actions may be in the form of questions such as "Did you mean to turn on the washing machine?" or "Now that the washing machine is done, do you want to start the drying cycle?" The corrective actions may be in the form of verbal utterances, such as explaining in an audio file, "You normally set your microwave for 3 minutes." The corrective actions may be in the form of visual statements. For example, connected device 520-1 may display the message "Microwave oven will schedule a 200 degree grill cycle for another 10 minutes," along with touchscreen buttons for "OK" or "Cancel."
[0090] 6A, 6B, 7, 8, 9A, and 9B show examples of training data for one or more portions of the RVM 530 for use in implementing artificial intelligence techniques that identify anomalous inputs to an audio device and take corrective action in response. In particular, the training data may include data inputs and data outputs for updating and training a machine learning model. The model and / or its trained data may be stored and / or updated in the activity model 532. The model may be further trained as additional real-world inputs are provided through use of the system 500.
[0091] FIG. 6A illustrates a first portion of training data 600 for a first machine learning model of a system 500 for identifying anomalous inputs, consistent with some embodiments of the present disclosure. FIG. 6B illustrates a second portion of training data 600 for a first machine learning model of a system 500, consistent with some embodiments of the present disclosure. Specifically, FIG. 6A illustrates multiple rows 610 of training data 600 provided to a detection model 540. Further, FIG. 6B illustrates a continuation of the rows 610 of training data 600. Each row 610 represents a particular routine, pattern, or activity that may be processed by a neural network configured to detect anomalous operation of a voice device. Row 610-1 may represent a header row that includes descriptive information for training data 600. Row 612 may represent additional rows of data that are not illustrated but may be included in training data 600.
[0092] Column 620 may represent each element of training data 600 that may be used to train and update weights and / or biases of detection model 540. Specifically, training of detection model 540 may include only a subset of elements 620 provided as inputs, such as elements 620-1, 620-2, 620-3, and 620-4. Additional elements, such as elements 620-5 and 620-6, may be provided as part of expected output. The expected output may be used as a comparison and for training the neural network of detection model 540. For example, if detection of training data 600 does not result in accurate identification of the presence of an abnormal body movement or voice command, the expected output may be used to update the weights and biases of detection model 540.
[0093] 7 illustrates a portion of training data 700 for a second machine learning model of a system 500 for identifying anomalous inputs, consistent with certain embodiments of the present disclosure. Specifically, FIG. 7 illustrates multiple rows 710 of training data 700 provided to missing input model 550. Each row 710 represents a particular routine, pattern, or activity that may be processed by a neural network configured to detect anomalous operation of a voice device. Row 710-1 may represent a header row that includes descriptive information for training data 700. Row 712 may represent additional rows of data that may be included in training data 700, not shown.
[0094] Column 720 may represent each element of the training data 700 that may be used to train and update the weights and / or biases of the missing input model 550. Specifically, the training of the missing input model 550 may include only a subset of the elements 720 provided as inputs, such as elements 720-1, 720-2, 720-3, and 720-4. Additional elements, such as element 720-5, may be provided as part of the expected output. The expected output may be used as a comparison and for purposes of training the neural network of the missing input model 550. For example, if the detection of the training data 700 does not result in an accurate identification of an anomalous input, including a missing follow-up voice command to a voice device, the expected output may be used to update the weights and biases of the missing input model 550.
[0095] 8 illustrates a portion of training data 800 for a third machine learning model of the system 500 for identifying anomalous inputs, consistent with certain embodiments of the present disclosure. Specifically, FIG. 7 illustrates multiple rows 810 of training data 800 provided to the erroneous input model 560. Each row 810 represents a particular routine, pattern, or activity that may be processed by a neural network configured to detect anomalous operation of a voice device. Row 810-1 may represent a header row that includes descriptive information for the training data 800. Row 812 may represent additional rows of data that may be included in the training data 800, although not shown.
[0096] Column 820 may represent each element of training data 800 that may be used to train and update weights and / or biases of erroneous input model 560. Specifically, training of erroneous input model 560 may include only a subset of elements 820 provided as inputs, such as elements 820-1, 820-2, 820-3, and 820-4. Additional elements, such as element 820-5, may be provided as part of the expected output. The expected output may be used as a comparison and for training the neural network of erroneous input model 560. For example, if detection of training data 800 does not result in accurate identification of anomalous inputs, including unexpected physical movements or unexpected voice commands to a voice device, the expected output may be used to update the weights and biases of erroneous input model 560.
[0097] FIG. 9A illustrates a first portion of training data 900 for a fourth machine learning model of the system 500 for providing accurate corrective actions, consistent with some embodiments of the present disclosure. FIG. 9B illustrates a second portion of training data 900 for a fourth machine learning model of the system 500, consistent with some embodiments of the present disclosure. Specifically, FIG. 9A illustrates multiple rows 910 of the training data 900 provided to the corrective action model 570. Further, FIG. 9B illustrates a continuation of the rows 910 of the training data 900. Each row 910 represents a particular routine, pattern, or activity that may be processed by a neural network configured to format and generate corrective actions. Row 910-1 may represent a header row that includes descriptive information for the training data 900. Row 912 may represent additional rows of data that are not depicted but may be included in the training data 900.
[0098] Column 920 may represent each element of training data 900 that may be used to train and update weights and / or biases of corrective motion model 570. Specifically, training of corrective motion model 570 may include only a subset of elements 920 provided as inputs, such as elements 920-1, 920-2, 920-3, and 920-4. Additional elements, such as elements 920-5, 920-6, and 920-7, may be provided as part of the expected output. The expected output may be used as a comparison and for training the neural network of corrective motion model 570. For example, if detection of training data 900 does not result in an accurate corrective motion, the expected output may be used to update the weights and biases of corrective motion model 570.
[0099] 10 is a method 1000 of performing a corrective action on an audio device consistent with some embodiments of the present disclosure. Method 1000 may be implemented generally with fixed functionality hardware, configurable logic, logic instructions, etc., or any combination thereof. For example, logic instructions may include assembler instructions, ISA instructions, machine instructions, machine-dependent instructions, microcode, state setting data, configuration data for integrated circuits, state information for personalizing electronic circuits, or other structural components native to hardware (e.g., a host processor, a central processing unit / CPU, a microcontroller, etc.), or combinations thereof.
[0100] Method 1000 may begin at 1005 by receiving one or more user interactions at 1010. The user interactions may be directed to a set of one or more audio devices. The audio devices may be connected devices in the environment, such as connected devices 520 of system 500. The user interactions may be received by a connected device (e.g., connected device 520-1), such as a desktop computer or other associated computer system. The user interactions may include physical actions, such as turning on a coffee maker, opening a washing machine door, or other user actions or placement of items. The user interactions may be received at 1010 continuously, repeatedly (e.g., every second, every tenth of a second), or at other relevant intervals.
[0101] Based on the user interaction, detection of a first input may be initiated at 1020. Detection may include processing by a connected device using position, motion, voice, or other relevant sensors. Specifically, the user interaction received at 1010 may be transmitted to the connected device over a network, such as a local area network. Detection may be performed by a voice device or may be performed by the connected device. Each input of the user interaction, including the first input, may be provided to an activity model for processing. Specifically, the first input may operate to trigger a machine learning model that performs one or more artificial intelligence operations to determine a particular type of activity, process, pattern, or routine associated with the first input. For example, when a coffee cup is placed in a coffee maker, a process is initiated to process the input to determine whether there is a routine associated with the placement.
[0102] If the first input is detected, i.e., Y at 1030, method 1000 may continue by determining whether there is a second input, at 1040. The determining may include using an activity model to make a prediction of the second input. In the coffee cup placement example, the determining may include performing machine learning to determine potential additional inputs directed at the coffee cup.
[0103] At 1050, the method 1000 may continue by monitoring for deviations from the second input. The deviations may be a difference between the potential second input determined at 1040 and any actual input received from user interaction at 1010. Monitoring for deviations may include comparing the output of the activity model to a stream of inputs received from connected devices. The deviations from the potential second input may be a missed command. A missed command may be a command that is not received within a certain predefined threshold. For example, based on the activity model, it may be determined that after laundry is placed on a voice device that is a smart washing machine, the predefined threshold is to receive a command within 45 seconds. The deviations from the potential second input may be an unexpected physical input or verbal command. An unexpected command may be a command that is inconsistent with a predefined input pattern or routine. For example, based on the activity model, it may be determined that after a voice device that is a smart oven has been heated to 400 degrees Fahrenheit for 20 minutes, the next input should be a command to heat to 150 degrees Fahrenheit for 30 minutes.
[0104] If an anomaly in the activity is identified, i.e., Y at 1060, the method 1000 may continue by performing a corrective action at 1070. The corrective action 1070 may be in the form of a command, such as a command to prompt the user for the abnormal input. The corrective action 1070 may be in the form of a request, such as a question or prompt, provided to the user. The corrective action 1070 may be provided by one of the voice devices, such as a smart appliance, that speaks the corrective action to the user. The corrective action 1070 may be provided by one of the connected devices, such as a smart watch worn by the user. After the corrective action 1070 is provided at 1070, or if there was no anomaly identified, i.e., N at 1060, or if there was no first input detected, i.e., N at 1030, the method 1000 may end at 1095.
[0105] The present invention may be a system, method or computer program product, or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer readable storage medium having stored thereon computer readable program instructions for causing a processor to carry out aspects of the present invention.
[0106] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium may be, by way of example and not by way of limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), SRAM, CD-ROM, DVD, memory sticks, floppy disks, punch cards or ridges in grooves or other mechanically encoded devices that have instructions recorded thereon, and suitable combinations thereof. As used herein, a computer-readable storage medium should not be construed as a transitory signal per se, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a wave guide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electric signal transmitted through a wire.
[0107] The computer readable program instructions described herein can be downloaded from the computer readable storage medium to each computing / processing device or to an external computer or storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). The network may be comprised of copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface of each computing / processing device receives the computer readable program instructions from the network and transfers the computer readable program instructions for storage in the computer readable storage medium within the respective computing / processing device.
[0108] The computer readable program instructions for carrying out the operations of the present invention may be either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language and similar programming languages. The computer readable program instructions may be executed entirely on the user's computer, as a stand-alone software package, or partially on the user's computer. Alternatively, they may be executed partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the computer readable program instructions to perform aspects of the invention.
[0109] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0110] These computer readable program instructions can be provided to a processor of a computer or other programmable data processing apparatus to generate a machine, such that the instructions, executed via a processor of the computer or other programmable data processing apparatus, generate means for implementing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer readable program instructions can also be stored in a computer readable storage medium connectable to a computer, programmable data processing apparatus, or other device or combination that functions in a particular way, such that the computer readable storage medium having the instructions stored thereon constitutes one of a product including instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0111] Computer readable program instructions, such as instructions to perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams on a computer, other programmable apparatus, or other device, may also be loaded into a computer, other programmable data processing apparatus, or other device to perform a series of operational steps on the computer, other programmable apparatus, or other device to generate a computer-implemented process.
[0112] The flowcharts and block diagrams in the figures illustrate the configuration, functionality, and operation of executable implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or part of an instruction, which constitutes one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions shown in the blocks may differ from the order shown in the figures. For example, two blocks shown in succession may in fact be accomplished as one step and executed simultaneously, substantially simultaneously, partially or fully in a time-overlapping manner, or the blocks may be executed in reverse order depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special purpose hardware-based system that executes the specified functions or operations or executes a combination of special purpose hardware and computer instructions.
[0113] The description of various embodiments of the present disclosure is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. It will be apparent to those skilled in the art that many modifications and variations are possible without departing from the scope of the described embodiments. The terms used in this specification are selected to best explain the principles of the embodiments, practical applications or technical improvements to the technology found in the market, or to allow those skilled in the art to understand the embodiments disclosed herein.
[0114] The description of various embodiments of the present disclosure is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. It will be apparent to those skilled in the art that many modifications and variations are possible without departing from the scope of the described embodiments. The terms used in this specification are selected to best explain the principles of the embodiments, practical applications or technical improvements to the technology found in the market, or to allow those skilled in the art to understand the embodiments disclosed herein.
Claims
1. receiving, by a first connected device, one or more user interactions directed to a set of one or more voice-controlled devices in the environment; Detecting a first input to a first voice control device of the set of voice control devices based on the user interaction; determining potential second inputs to the set of voice control devices based on an activity model in response to the first input; monitoring the user interaction for deviation from the potential second input in response to the first input; identifying anomalies in activity in the environment based on the monitoring; and performing corrective action in response to an abnormality in the activity; A method comprising:
2. The method of claim 1 , wherein the first connected device is a voice-controlled device of the set of voice-controlled devices.
3. The method of claim 1 , wherein the first input comprises at least one voice command by a user.
4. The method of claim 1 , wherein the first input comprises at least one physical action by a user.
5. The method of claim 1 , wherein the activity model comprises a machine learning model.
6. the machine learning model relates only to the first voice-controlled device; the activity model includes a second machine learning model associated with only a second voice control device of the set of voice control devices; The method of claim 5.
7. The method further comprises: training the machine learning model with a set of training data; generating one or more voice control device routines based on the machine learning model, the voice control device routines including a plurality of potential inputs to one or more voice control devices of the set of voice control devices including the potential second inputs; The method of claim 5 , comprising:
8. The method of claim 7 , wherein the training data includes previously issued voice commands.
9. The method of claim 7 , wherein the training data includes previously issued physical actions.
10. The method of claim 7 , wherein the training data also includes temporal information describing one or more temporal relationships.
11. 11. The method of claim 10, wherein a first one of the temporal relationships comprises a predetermined threshold amount of time between a potential first input to a first voice control device and the potential second input to the first voice control device.
12. The corrective action is generating a response associated with the potential second input based on the deviation; sending the response to a user's client device; causing the client device of the user to provide the response; and The method of claim 1 , comprising:
13. The method of claim 12 , wherein the response is generated based on a machine learning model.
14. The method of claim 1 , wherein the potential second input comprises an expected interaction and the deviation is the absence of the expected interaction.
15. The method of claim 1 , wherein the potential second input comprises a first interaction and the deviation is a second interaction.
16. The method of claim 1 , wherein the potential second input comprises a predetermined range of acceptable values, and the deviation is outside of the range.
17. the potential second input comprises a predetermined range of acceptable values including a lower limit and an upper limit; the deviation is closer to one of the lower and upper limits of the range, The method further comprises: updating the activity model based on the deviations; The method of claim 1 , comprising:
18. The method of claim 1 , wherein the potential second input comprises an expected interaction within a predetermined threshold amount of time, and the deviation is an interaction outside the predetermined threshold amount of time.
19. 1. A system comprising: a memory, the memory containing one or more instructions; a processor communicatively coupled to the memory, the processor, in response to reading the one or more instructions, receiving, by a first connected device, one or more user interactions directed to a set of one or more voice-controlled devices in the environment; Detecting a first input to a first voice control device of the set of voice control devices based on the user interaction; determining potential second inputs to the set of voice control devices based on an activity model in response to the first input; monitoring the user interaction for deviation from the potential second input in response to the first input; identifying anomalies in activity in the environment based on the monitoring; and performing corrective action in response to an abnormality in the activity; A system configured to run
20. A computer program, the computer program comprising: and program instructions, the program instructions comprising: receiving, by a first connected device, one or more user interactions directed to a set of one or more voice-controlled devices in the environment; Detecting a first input to a first voice control device of the set of voice control devices based on the user interaction; determining potential second inputs to the set of voice control devices based on an activity model in response to the first input; monitoring the user interaction for deviation from the potential second input in response to the first input; identifying anomalies in activity in the environment based on the monitoring; and performing corrective action in response to an abnormality in the activity; A computer program configured to execute the