Artificial Intelligence Voice Response System for Users with Speech Disorders
The system addresses communication barriers for speech-impaired individuals by using IoT sensors and AI models to predict user intent and provide customized voice menus, improving interaction with AI systems while ensuring data privacy and user control.
Patent Information
- Application Number
- JP2023512417
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-11
- Filing Date
- 2021-09-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-09-06
AI Technical Summary
Individuals with speech disorders or fatigue are unable to effectively communicate voice commands to artificial intelligence (AI) systems due to speech impediments or physical conditions, hindering interaction with voice response systems.
A system that collects user data from connected devices, trains a voice response system, identifies a wake-up signal, and engages with the user through IoT biometric sensors and augmentative and alternative communication devices to facilitate communication, using LSTM-RNN models for pattern recognition and deep reinforcement learning to predict user intent and provide customized menus.
Enables individuals with speech impairments to interact with AI systems by predicting user intent and providing tailored voice responses, enhancing communication capabilities and ensuring data privacy and user control over data usage.
Smart Images

Figure 0007710813000001 
Figure 0007710813000002 
Figure 0007710813000003
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of computing, and more particularly to virtual assistants.
Background Art
[0002] Due to speech impediment or other speech articulation disorders or both, a speech disorder, a person may not be able to form a speech command understandable by an artificial intelligence (AI) speech response system by constructing language, using appropriate words, or both. Fatigue or other physical conditions, or both, may also cause a person to be unable to submit a speech command, speak a detailed request to an AI speech response system, or both.
Summary of the Invention
[0003] Embodiments of the present invention disclose a method, a computer system, and a computer program product for voice response. The present invention can include collecting user data from at least one connected device. The present invention can include training a voice response system based on the collected user data. The present invention can include identifying a wake-up signal based on the trained voice response system. The present invention can include determining that user engagement is intended based on identifying the wake-up signal. The present invention can include engaging with the user through at least one connected device.
Brief Description of the Drawings
[0004] These and other objects, features, and advantages of the present invention will become apparent from the following detailed description of its exemplary embodiments, read in conjunction with the accompanying drawings. The various features of the drawings are not to scale as they are provided for the purpose of clarifying the understanding of the present invention by those skilled in the art when read in conjunction with the detailed description.
[0005]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
DETAILED DESCRIPTION OF THE INVENTION
[0006] Although detailed embodiments of the claimed structures and methods are disclosed herein, it is to be understood that the disclosed embodiments are merely exemplary of the claimed structures and methods that may be embodied in various forms. However, the present invention can be embodied in many different forms and should not be construed as limited to the exemplary embodiments described herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope of the invention to those skilled in the art. In the description, well-known features and techniques may be omitted so as not to unnecessarily obscure the presented embodiments.
[0007] The present invention can integrate a system, a method, a computer program product, or a combination thereof at any possible technical detail level. The computer program product can include one or more computer-readable storage media having computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0008] The computer-readable storage media can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage media can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage media includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a punch card, or a mechanically encoded device such as a raised structure in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, the computer-readable storage media is not construed to be a transitory signal per se, such as a radio wave, or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0009] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as, for example, the Internet, a local area network, a wide area network, or a wireless network, or combinations thereof. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers or edge servers, or combinations thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.
[0010] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code described in any combination of one or more programming languages including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, may be executed partly on the user's computer and partly as a stand-alone software package, may be executed partly on the user's computer and partly on a remote computer, or may be executed entirely on the remote computer or server. In the last scenario, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may utilize the state information of the computer-readable program instructions to execute the computer-readable program instructions to customize the electronic circuit in order to implement aspects of the present invention.
[0011] Aspects of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0012] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, thereby creating means for implementing the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both by instructions executed by the processor of the computer or other programmable data processing apparatus. These computer program instructions can also be stored in a computer-readable medium so that the instructions stored in the computer-readable medium include a product that includes instructions for implementing the mode of functions / operations specified in one or more blocks of a flowchart, a block diagram, or both. These computer program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby providing a process for implementing the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both by instructions executed on the computer or other programmable apparatus.
[0013] These computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby providing a process for implementing the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both by instructions executed on the computer or other programmable apparatus.
[0014] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart may represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in accordance with the functionality involved, actually be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams or flowchart diagrams, or combinations of blocks in the block diagrams or flowchart diagrams or both, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or a combination of dedicated hardware and computer instructions.
[0015] The exemplary embodiments described below provide a system, method, and program product for voice response. Accordingly, the present embodiments have the ability to improve the technical field of voice response systems by enabling a speech impaired user to communicate with a voice response system using one or more connected devices including an augmentative and alternative communication device. More specifically, the present invention can include collecting user data from at least one connected device. The present invention can include training a voice response system based on the collected user data. The present invention can include identifying a wakeup signal based on the trained voice response system. The present invention can include determining that user engagement is intended based on identifying the wakeup signal. The present invention can include engaging with the user via at least one connected device.
[0016] As described above, due to speech function disorders or other speech articulation disorders or both, which are speech disorders, a person may not be able to construct language, or utilize appropriate words, or both, in order to form voice commands that can be understood by an artificial intelligence (AI) voice response system. Fatigue or other physical conditions, or both, due to illness may also cause an individual to be unable to submit voice commands, or speak elaborate requests to an AI voice response system, or both.
[0017] Accordingly, it may be advantageous to provide means, particularly for an artificial intelligence (AI) system, which may include, but is not limited to, observing human conversations, including ambient conversations, learning menu options using motion signals or biometric signals or both, and generating a customized voice menu that can assist a user with a speech disorder when executing an intended voice response or voice command.
[0018] According to at least one embodiment, an artificial intelligence (AI) system can predict when a user submits a voice command and whether the user desires to submit it, or that there is a possibility that the user cannot submit a voice command, or both.
[0019] According to at least one embodiment, when predicting when a user submits a voice command and whether the user desires to submit it, or that there is a possibility that the user cannot submit a voice command, or both, the user's rules of thumb or health condition or both can be taken into account. Using the user's rules of thumb or health condition or both, a spoken menu can be predicted for the topic of voice commands or voice requests or both, and optionally, at least one appropriate voice command can be selected therefrom and provided to the user.
[0020] According to at least one embodiment, the voice response program can ensure that the user's voice response data and / or integrated data sources, or both, cannot be used by any other system without the user's sufficient knowledge and approval. Through system integration, the user of the voice response program can integrate tools such as IoT biometric sensors, augmentative and alternative communication devices (AAC devices), and / or video streams to provide enhanced functionality and give options for further training the user's own instance of the voice response program. The integration process with the voice response program can be explicitly opt-in, and any collected data cannot be shared outside the user's own personal instance of the voice response program.
[0021] Referring to FIG. 1, an exemplary networked computer environment 100 according to one embodiment is shown. The networked computer environment 100 can include a computer 102 having a processor 104 and a data storage device 106 capable of executing a software program 108 and a voice response program 110a. The networked computer environment 100 can also include a server 112 capable of executing a voice response program 110b that can interact with a database 114 and a communication network 116. The networked computer environment 100 can include a plurality of computers 102 and servers 112, only one of which is shown. The communication network 116 can include various types of communication networks such as a wide area network (WAN), a local area network (LAN), a telecommunications network, a wireless network, a public switched network, or a satellite network, or combinations thereof. The connected device 118 is shown as a separate entity in itself, but can be integrated into another part of the computer network environment. It should be understood that FIG. 1 provides only an illustration of one implementation and does not imply any limitation with respect to environments in which different embodiments can be implemented. Many modifications to the shown environment can be made based on design and implementation requirements.
[0022] Client computer 102 can communicate with server computer 112 via communication network 116. Communication network 116 can include connections such as wired, wireless communication links, or optical fiber cables. As described with reference to FIG. 3, server computer 112 can include internal component 902a and external component 904a respectively, and client computer 102 can include internal component 902b and external component 904b respectively. Server computer 112 can also operate in a cloud computing service model such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS). Server 112 can also be deployed in a cloud computing deployment model such as a private cloud, community cloud, public cloud, or hybrid cloud. Client computer 102 can be, for example, a mobile device, a phone, a portable terminal, a netbook, a laptop computer, a tablet computer, a desktop computer, or any type of computing device capable of executing programs, accessing the network, and accessing database 114. According to various implementations of this embodiment, voice response programs 110a, 110b can interact with database 114, which can be incorporated into various storage devices such as, but not limited to, computer / mobile device 102, networked server 112, or cloud storage service.
[0023] According to this embodiment, a user using the client computer 102 or the server computer 112 can use the voice response programs 110a and 110b (respectively) to enable a user with a speech disorder to communicate with a voice response system using one or more connected devices (e.g., the connected device 118) including an augmentative and alternative communication device. The voice response method will be described in more detail below in connection with FIG. 2.
[0024] Referring now to FIG. 2, an operational flowchart is shown that depicts an exemplary voice response process 200 used by the voice response programs 110a and 110b, according to at least one embodiment.
[0025] At 202, the voice response programs 110a, 110b collect user data. The data collection module of the voice response programs 110a, 110b can collect data including, but not limited to, past behavioral data and / or conversation data, and new data supplied by and collected in real time by the user's connected devices (e.g., the connected device 118).
[0026] The data collection module can capture behavioral pattern data, biometric pattern data, or movement pattern data, or a combination thereof, from a user with a speech disorder or any other user, or both, and store the captured (i.e., collected) data in a knowledge corpus (e.g., the database 114).
[0027] In particular, wearable devices including Internet of Things (IoT)-connected rings, glasses, clothing (e.g., having a heart sensor and / or a respiration sensor, or both), watches, shoes, and / or fitness trackers can supply data to a data collection module, which can include camera-supplied data and / or device data from any other IoT biometric sensor, or both.
[0028] Data can also be collected from various augmentative and alternative communication (AAC) devices (e.g., combinations thereof). An AAC device can be a device that enables and / or facilitates communication for impairment patterns and / or disability patterns presented by individuals with expressive communication disorders. An augmentative communication device can be used by individuals who can speak somewhat but cannot understand or have limited speaking ability. An alternative communication device can be used by individuals who cannot speak and rely on another communication method to express their thoughts (e.g., desires and requests in particular).
[0029] Data can be collected from a video device or an audio streaming device, or both. Once the raw video stream of data is collected, it can pass through an image processing system or a video processing system, or both, to classify engagement indicators for model input (e.g., identify raising a hand, eye blinking, etc.). The image processing system or the video processing system, or both, can be an IBM Watson (trademark) (Watson and all Watson-based trademarks are trademarks or registered trademarks of International Business Machines Corporation in the United States or other countries or both) visual recognition solution among other solutions. The Watson (trademark) visual recognition solution can use deep learning algorithms to analyze images for faces (e.g., face recognition), scenes, objects, or any other content, or combinations thereof, tag the analyzed visual content, and classify and search it.
[0030] The raw audio stream data collected from an audio streaming device can be passed to an audio-to-text processor such as Watson (trademark) speech-to-text so that the content can be analyzed by natural language processing (NLP) algorithms. NLP algorithms such as the Watson (trademark) tone analyzer (e.g., dynamically determine user satisfaction or dissatisfaction) and sentiment analysis (in particular, determine whether the user is nervous, angry, disappointed, sad, happy, etc.) application programming interfaces (APIs), and the Watson (trademark) natural language classifier (e.g., collect audio content and keyword indicator data) can be used.
[0031] For example, in a medical facility, if at least one user of the voice response programs 110a, 110b has a speech disorder and is unable to express voice commands, the voice response programs 110a, 110b can be utilized and trained. In this example, the data collected by the voice response programs 110a, 110b can include both commands spoken by the user with a speech disorder and / or commands spoken by any part of the medical support team, and changes resulting from action parameters and / or biometric parameters identified by the connected device or wearable device or both.
[0032] In 204, the voice response system is trained based on the collected data. Using a long short-term memory (LSTM) recurrent neural network (RNN) for time series sequencing (e.g., for connected sequencing patterns such as speech), the intended topic (i.e., topic, user topic) of the voice requests of the user with a speech disorder among other users can be predicted.
[0033] As previously described with respect to step 202 above, the data collected by the data collection module can be interpreted to identify the action pattern data, biometric pattern data, or movement pattern data or combinations thereof of the user (e.g., the user with a speech disorder or any other user of the voice response programs 110a, 110b, or both), and the intended topic or request or both of the user can be predicted. This can be further done using an LSTM-RNN model, which is described in more detail with respect to step 208 below.
[0034] At 206, a wake-up signal is identified. Once the knowledge corpus (e.g., database 114) is complete (e.g., sufficient data has been collected to make predictions about future outcomes in the knowledge base), any data collected by connected devices (e.g., in particular, connected wearable devices, IoT sensors, cameras) that track changes in the behavioral parameters or biometric parameters or both of a user with speech impairments can wake up the artificial intelligence (AI) device and trigger device engagement with the user.
[0035] Connected IoT devices can passively listen to the user's conversation until a wake-up signal is identified and can only start storing data once the wake-up signal has been identified. However, users of voice response programs 110a, 110b can trigger the connected IoT device to turn off the listening function and start listening only when a command is issued.
[0036] At 208, voice response programs 110a, 110b determine that the user wishes to engage with the connected device. When the artificial intelligence (AI) device wakes up, all data collected by the connected device can be passed to a random forest algorithm to perform binary classification (e.g., classify the data to interpret whether the user wishes to engage with the system or not based on classification rules). For example, voice response programs 110a, 110b can obtain all inputs from the data collection module, execute the inputs through a random forest model, and use binary classification (e.g., 0 represents data that is not needed and the user does not wish to engage, 1 represents data that is needed and the user wishes to engage) to determine whether the input is needed (e.g., whether the user wishes to engage).
[0037] When the voice response programs 110a and 110b determine, based on the classification rules, that the user desires to engage with the system, they can pass the collected data to a deep reinforcement learning model (i.e., the LSTM-RNN model) to determine how to proceed with the engagement with the user.
[0038] The user's consent or refusal to engage with the voice response programs 110a and 110b can be fed back to the deep learning model to further adjust the model. Negative user feedback can act as a penalty, and positive user feedback can act as a reward. The deep reinforcement learning model can act as a feedback loop and classify the data as positive or negative to further adjust the model towards the desired result. This can assist the deep reinforcement learning model in adjusting the current state and determining future actions for engagement with the voice response programs 110a and 110b.
[0039] At 210, the voice response programs 110a and 110b engage with the user. To engage with a user with speech impairments (i.e., the user), the voice response programs 110a and 110b can provide the user with a customized menu related to the predicted topic. The voice response program can determine a voice request that can be executed considering the action signal or biometric signal or both collected by the data collection module as described above with respect to step 202. While navigating the voice menu, user feedback including consent feedback or dissent feedback, or both (e.g., positive or negative, or both biometric data and / or action data received as a result of a given question) can be analyzed. The voice menu can be navigated by the voice response programs 110a and 110b until a customized menu related to the predicted topic is determined and a voice command can be executed accordingly.
[0040] Continuing with the example from 202 above, a user with speech impairment in a medical facility may be asked "Are you hungry?" and "Are you thirsty?". After the question "Are you thirsty?", a visual signal can be identified (e.g., the facial expression made by the user), and the next set of questions can include "Do you want water?" and "Do you want tea?". Using this data (e.g., video data) observed by the connected devices of the voice response programs 110a, 110b or the wearable device or both, a knowledge corpus can be generated to identify the intended topic and the associated hierarchical voice menu.
[0041] Here, the LSTM-RNN model can be used to process the user's speech and determine how to proceed based on the user's speech. The LSTM-RNN model can be an artificial recurrent neural network architecture used in the field of deep learning, which, unlike a standard feedforward neural network, functions based on feedback connections. The LSTM-RNN model can process not only a single data point (e.g., an image of the user obtained by a connected device) but also an entire sequence of data (e.g., the utterances or video of the interaction between the user and the device). For example, the LSTM-RNN model can be applied to tasks such as non-segmented speech recognition, handwritten character recognition, and anomaly detection in network traffic or intrusion detection systems.
[0042] In this application, the LSTM-RNN model can be used to process a user's voice request by decomposing the observed part of the utterance into sequential dependent inputs and predicting the user's intended topic. This speech-to-text function can function such that the input voice is used as sequential dependent inputs and the predicted intended topic is used as the result output based on the LSTM-RNN model.
[0043] Here, the LSTM-RNN model can be used to improve a knowledge corpus (e.g., database 114) by correlating the collected behavioral inputs, body language, or biometric signals, or combinations thereof, with the intended topic or a hierarchical voice menu related to the intended topic.
[0044] To correlate data with a particular aspect of the voice menu, the voice menu can be defined, or identified, or both, in the knowledge corpus (e.g., database 114). The voice response programs 110a, 110b can identify an appropriate voice menu based on the collected data by, for example, identifying the most common commands in view of the type of data received (e.g., based in particular on a particular behavioral input or biometric signal, or both).
[0045] According to at least one embodiment, based on engagement with the user, the voice response programs 110a, 110b can dynamically create a voice menu over time and can start by using an existing voice menu associated with a particular domain on the connected IoT device. For example, if the user says "Set an Alexa timer", the IoT device can respond by having the user start descending through the associated existing hierarchical voice menu for "timer", by asking "What would you like to call the timer", and in particular "What time would you like it for". The voice response programs 110a, 110b can learn to interact with the existing voice menu based on further user commands such as in particular "Set the time", "Set a stop point", "Remind me", or "Make sure I don't forget". Based on receiving these related commands, as described above, the voice response programs 110a, 110b can learn that the user is entering the hierarchical voice menu for "timer".
[0046] Action data or biometric data or both, including spoken text or sound or both, can be interpreted as related to the user's activities (e.g., eating, drinking, watching TV, or listening to music, or combinations thereof, among many other things), and the menu can be customized accordingly. Patterns in the user's actions (i.e., action patterns) can, as previously described with respect to step 202 above, assist in identifying the intended topic, and the voice response programs 110a, 110b can create a hierarchical set of questions based on the observed interaction or set of interactions of the impaired user with the artificial intelligence (AI) device.
[0047] According to at least one embodiment, the voice response programs 110a, 110b may start from a set of existing (e.g., pre-configured on an IoT device) or learned patterns based on interactions and / or observed actions with the user (e.g., the user's current health status or experience or both), for non-routine events (e.g., observed body movements that are different from the user's normal body movements as determined by the voice response programs 110a, 110b and / or any connected devices, or events where there is no prior data related to the user's requests that can be used by the voice response programs 110a, 110b), and / or may process non-routine events by starting from such a set of patterns, and / or may initiate a phone call to a living person (e.g., a person configured in the user profile of the voice response programs 110a, 110b) who can assist the voice response programs 110a, 110b in understanding the non-routine event.
[0048] In 208, if the voice response programs 110a, 110b determine that the user did not wish to engage, the program ends.
[0049] It can be understood that FIG. 2 provides only an illustration of one embodiment and does not imply any limitation as to how different embodiments can be implemented. Many modifications can be made to the illustrated embodiment based on design and implementation requirements.
[0050] FIG. 3 is a block diagram 900 of the internal and external components of the computer shown in FIG. 1 according to an exemplary embodiment of the present invention. It should be understood that FIG. 3 provides only an illustration of one implementation and does not imply any limitation as to the environment in which different embodiments can be implemented. Many modifications can be made to the illustrated environment based on design and implementation requirements.
[0051] Data processing systems 902, 904 represent any electronic device capable of executing machine-readable program instructions. The data processing systems 902, 904 can represent smartphones, computer systems, PDAs, or other electronic devices. Examples of computing systems, environments, or configurations, or combinations thereof, that can be represented by the data processing systems 902, 904 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, network PCs, minicomputer systems, and distributed cloud computing environments that include any of the above systems or devices.
[0052] The user client computer 102 and the network server 112 can each include a respective set of internal components 902a, b and external components 904a, b shown in FIG. 3. Each set of internal components 902a, b includes one or more processors 906, one or more computer-readable RAMs 908, and one or more computer-readable ROMs 910 on one or more buses 912, one or more operating systems 914, and one or more computer-readable tangible storage devices 916. One or more operating systems 914, software programs 108, and voice response programs 110a within the client computer 102, and voice response programs 110b within the network server 112 can be stored on one or more computer-readable tangible storage devices 916 for execution by one or more processors 906 via one or more RAMs 908 (typically including cache memory). In the embodiment shown in FIG. 3, each of the computer-readable tangible storage devices 916 is a magnetic disk storage device of an internal hard drive. Alternatively, each of the computer-readable tangible storage devices 916 is a semiconductor storage device such as a ROM 910, an EPROM, a flash memory, or any other computer-readable tangible storage device capable of storing computer programs and digital information.
[0053] Each set of internal components 902a, b also includes an R / W drive or interface 918 for reading and writing with one or more portable computer-readable tangible storage devices 920, such as CD-ROMs, DVDs, memory sticks, magnetic tapes, magnetic disks, optical disks, or semiconductor storage devices. Software programs, such as software program 108 and voice response programs 110a and 110b, can be stored in one or more of the respective portable computer-readable tangible storage devices 920, read through the respective R / W drives or interfaces 918, and loaded onto the respective hard drives 916.
[0054] Each set of internal components 902a, b can also include a network adapter (or switch port card) or interface 922, such as a TCP / IP adapter card, a wireless wi-fi interface card, or a 3G or 4G wireless interface card, or other wired or wireless communication links. The software program 108 and voice response programs 110a within the client computer 102, and the voice response program 110b within the network server computer 112, can be downloaded from an external computer (such as a server) via a network (such as the Internet, a local area network, or other wide area network), and the respective network adapters or interfaces 922. The software program 108 and voice response programs 110a within the client computer 102, and the voice response program 110b within the network server computer 112 are loaded onto the respective hard drives 916 from the network adapter (or switch port adapter) or interface 922. The network can include copper wires, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or combinations thereof.
[0055] Each set of external components 904a, b can include a computer display monitor 924, a keyboard 926, and a computer mouse 928. The external components 904a, b can also include a touch screen, a virtual keyboard, a touch pad, a pointing device, and other human interface devices. Each set of internal components 902a, b also includes a device driver 930 for interfacing with the computer display monitor 924, the keyboard 926, and the computer mouse 928. The device driver 930, the R / W drive or interface 918, and the network adapter or interface 922 also include hardware and software (stored in the storage device 916 or ROM 910, or both).
[0056] This disclosure includes a detailed description of cloud computing, but it is to be understood in advance that the implementation of the teachings detailed herein is not limited to a cloud computing environment. Rather, embodiments of the disclosed technology can also be implemented in conjunction with any other type of computing environment now known or later developed.
[0057] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0058] The characteristics are as follows. On-demand self-service: Cloud consumers can, as needed, automatically and unilaterally provision computing capabilities such as server time and network storage without the need for a human to interact with the service provider. Broad network access: The capabilities are available over the network and accessed through standard mechanisms that promote use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs). Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically assigned and reassigned according to demand. Consumers generally have no control or knowledge over the exact location of the provided resources, but may be able to specify a higher level of abstraction (e.g., country, state, or data center). Rapid elasticity: Capabilities can be provisioned quickly and elastically, and in some cases automatically, to quickly scale out and rapidly released to quickly scale in. To the consumer, the capabilities available for provisioning often appear to be unlimited and can be purchased in any quantity at any time. Measured service: The cloud system automatically controls and optimizes resource use by using some form of metering capability appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts) at some level of abstraction. Resource usage can be monitored, controlled, and reported to provide transparency for both the provider and consumer of the utilized service.
[0059] The service model is as follows. Software as a Service (SaaS): The function provided to the consumer is to use the provider's applications that run on cloud infrastructure. These applications are accessible from various client devices through a client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, server, operating system, storage, or individual application capabilities, with the exception of limited user-specific application configuration settings assumed. Platform as a Service (PaaS): The function provided to the consumer is to deploy the applications created or acquired by the consumer, which are created using programming languages and tools supported by the provider, onto the cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure including the network, server, operating system, or storage, but controls the deployed applications and, in some cases, the environment configuration that hosts the applications. Infrastructure as a Service (IaaS): The function provided to the consumer is to provision processing, storage, network, and other basic computing resources on which the consumer can deploy and run any software that may include an operating system and applications. The consumer does not manage or control the underlying cloud infrastructure, but has limited control over the operating system, storage, control of the deployed applications, and, in some cases, selection of network components (e.g., host firewall).
[0060] The deployment model is as follows. Private Cloud: The cloud infrastructure is operated solely for an organization. This can be managed by the organization or a third party and can exist on-premises or off-premises. Community Cloud: The cloud infrastructure is shared by several organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). This can be managed by those organizations or a third party and can exist on-premises or off-premises. Public Cloud: The cloud infrastructure is available to the general public or a large industry group and is owned by an organization that sells cloud services. Hybrid Cloud: The cloud infrastructure remains a distinct entity but is a hybrid of two or more clouds (private, community, or public) tied together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application portability.
[0061] Cloud computing environments are service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0062] Referring now to FIG. 4, an exemplary cloud computing environment 1000 is shown. As illustrated, cloud computing environment 1000 includes one or more cloud computing nodes 100 that can communicate with local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or cellular telephone 1000A, desktop computer 1000B, laptop computer 1000C, or in-vehicle computer system 1000N, or a combination thereof. The nodes 100 can communicate with one another. The nodes 100 can be physically or virtually grouped in one or more networks, such as the private cloud, community cloud, public cloud, or hybrid cloud, described above, or a combination thereof (not shown). This enables cloud computing environment 1000 to provide Infrastructure as a Service, Platform as a Service, or Software as a Service, or a combination thereof, such that a cloud consumer need not maintain resources on a local computing device. The types of computing devices 1000A-N shown in FIG. 4 are intended to be exemplary only, and it should be understood that cloud computing nodes 100 and cloud computing environment 1000 can communicate with any type of computerized device via any type of network or network addressable connection, or both (e.g., using a web browser).
[0063] Referring now to FIG. 5, a set of functional abstraction layers 1100 provided by cloud computing environment 1000 is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 5 are intended to be exemplary only, and embodiments of the present invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided.
[0064] The hardware and software layer 1102 includes hardware components and software components. Examples of hardware components include mainframe 1104, RISC (Reduced Instruction Set Computer) architecture-based server 1106, server 1108, blade server 1110, storage device 1112, and network and networking components 1114. In some embodiments, the software components include network application server software 1116 and database software 1118.
[0065] The virtualization layer 1120 provides an abstraction layer that can provide the following examples of virtual entities: virtual server 1122, virtual storage 1124, virtual network 1126 including a virtual private network, virtual applications and operating systems 1128, and virtual client 1130.
[0066] In one example, the management layer 1132 can provide the functions described below. Resource provisioning 1134 provides for the dynamic procurement of computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 1136 provides for cost tracking when resources are utilized within a cloud computing environment and for billing or charging for the consumption of these resources. In one example, these resources can include application software licenses. Security provides for authentication of cloud consumers and tasks and for protection of data and other resources. The user portal 1138 provides access to the cloud computing environment for consumers and system administrators. Service level management 1140 provides for the allocation and management of cloud computing resources so that required service levels are met. Planning and fulfillment of service level agreements (SLAs) 1142 provides for the pre-placement and procurement of cloud computing resources whose future requirements are predicted in accordance with an SLA.
[0067] The workload layer 1144 provides examples of functions that can utilize a cloud computing environment. Examples of workloads and functions that can be provided from this layer include mapping and navigation 1146, software development and life cycle management 1148, virtual classroom education delivery 1150, data analysis processing 1152, transaction processing 1154, and voice response 1156. The voice response programs 110a, 110b provide a method for enabling a user with a speech impairment to communicate with a voice response system using one or more connected devices including an augmentative and alternative communication device.
[0068] The descriptions of the various embodiments of the present disclosure are presented for illustrative purposes, but these are not intended to be exhaustive or to limit the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, the practical application, or a technical improvement over the technologies found in the marketplace, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for voice response, executed by computer information processing, comprising: collecting user data from at least one connected device; training a voice response system based on the collected user data; identifying a wake-up signal, which is a change in biometric parameters recorded on a connected Internet of Things (IoT) device, based on the trained voice response system; determining that user engagement is intended based on identifying the wake-up signal; engaging with the user through the at least one connected device. A method as described above.
2. The method according to claim 1, wherein the at least one connected device is an assistive and alternative communication device.
3. Training the voice response system based on the collected user data further comprises: predicting the topic of a voice request using a long short-term memory recurrent neural network. The method according to claim 1.
4. Determining that user engagement is intended further comprises: performing a binary classification of the collected user data using a random forest algorithm. The method according to claim 1.
5. A method for voice response, executed by computer information processing, comprising: collecting user data from at least one connected device; training a voice response system based on the collected user data; identifying a wake-up signal based on the trained voice response system; determining that user engagement is intended based on identifying the wake-up signal; engaging with the user through the at least one connected device, wherein engaging with the user through the at least one connected device further comprises: providing the user with a customized menu based on the user data; analyzing user feedback; predicting user topics. A method as described above.
6. A method for voice response, executed by computer information processing, comprising: Collecting user data from at least one connected device; Training a voice response system based on the collected user data; Identifying a wake-up signal based on the trained voice response system; Determining that user engagement is intended based on identifying the wake-up signal; Engaging with the user through the at least one connected device and including, wherein the user data is stored in a database, and the database is updated based on engagement with the user to correlate the user data with user topics predicted by a long short-term memory recurrent neural network, a method. **Claim 7** A voice menu, the method according to claim 6, predetermined in the database. **Claim 8** A computer system for voice response, comprising one or more processors and one or more computer-readable memories, the computer system being configured to execute the method according to any one of claims 1 to 7, a computer system. **Claim 9** A computer program for voice response, executable by a processor, and causing the processor to execute the method according to any one of claims 1 to 7, a computer program. **Claim 10** A computer-readable storage medium storing the computer program according to claim 9. **Claim 11** A method for voice response, executed by computer information processing, comprising collecting user data from at least one connected device; training a voice response system based on the collected user data; identifying, based on the trained voice response system, a wake-up signal that is a change in biometric parameters recorded on a connected Internet of Things (IoT) device; determining that user engagement is intended based on identifying the wake-up signal; engaging with the user through the at least one connected device and including, wherein engaging with the user through the at least one connected device is Initiating a telephone call to a live person who can assist the voice response system in understanding non-routine events received from the user A method comprising.
Citation Information
Patent Citations
Voice interactive device and voice interactive method
JP2017211608A
Multimodal Gesture-Based Interactive System and Method Using One Single Sensing System
JP2018505455A
Voice-activated selective memory for a voice capture device - Patent Application 20070122997
JP2020533628A
Personalized gesture recognition for user interaction with assistant systems
WO2019204651A1