Access authentication in AI systems

The system allows secure sharing of selected knowledge corpus portions between AI voice response systems, addressing the challenge of user-specific command recognition and control in smart environments by enabling authorized access and interaction.

JP7729881B2Active Publication Date: 2025-08-26INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023520089
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-05
Filing Date
2021-10-19
Publication Date
2025-08-26
Estimated Expiration
2041-10-19

AI Technical Summary

Technical Problem

Existing AI voice response systems struggle to securely and efficiently share knowledge corpora between users, limiting the ability of one user's system to recognize and respond to the commands of another user without access to their personalized knowledge corpus, thereby restricting interoperability and functionality in smart environments.

Method used

A system and method that enables a first user's AI voice response system to electronically perceive the presence of a second user, request access authorization, instantiate a communication session, and retrieve a selected portion of the second user's knowledge corpus through a portable device, allowing the first user's system to respond to voice commands based on the retrieved data while maintaining security and privacy.

Benefits of technology

Facilitates secure and efficient sharing of selected knowledge corpus portions, enabling broader user interaction and control of IoT devices by allowing a first user's AI voice response system to recognize and execute commands of a second user, enhancing interoperability and functionality in smart environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007729881000001
    Figure 0007729881000001
  • Figure 0007729881000002
    Figure 0007729881000002
  • Figure 0007729881000003
    Figure 0007729881000003
Patent Text Reader

Abstract

Access authentication in an artificial intelligence system includes electronically recognizing the physical presence of a second user using an artificial intelligence voice response system (AIVRS) of a first user. A voice request is generated by the first user's AIVRS and transmitted to the second user, requesting access authorization to a knowledge corpus stored by the second user's AIVRS. Based on the second user's voice response, the first user's AIVRS instantiates an electronic communication session with the second user's AIVRS. The session is initiated via an electronic communication connection with the second user's portable device. A selected portion of the knowledge corpus is retrieved by the first user's AIVRS from the second user's AIVRS. The selected portion is based on the voice response. Based on the selected portion, an action is initiated by one or more IoT devices in response to a voice prompt interpreted by the first user's AIVRS.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to computer-based data exchange, and more particularly to secure data access and exchange between computer systems endowed with artificial intelligence. [Background technology]

[0002] It is estimated that more than one million homes in the United States are equipped with automation systems, and the number of smart devices installed in U.S. homes is expected to soon exceed 50 million. Smart devices are endowed with artificial intelligence (AI) and can operate according to a user's specific intents and preferences. User intents and preferences are part of the knowledge corpus that smart devices can acquire using machine learning. Smart devices increasingly communicate with each other over the Internet or other data communication networks and coordinate their integrated functions, forming the Internet of Things (IoT). Summary of the Invention

[0003] In one or more embodiments, a method includes using a first user's artificial intelligence (AI) voice response system to electronically perceive the physical presence of a second user. The method includes transmitting a voice request generated by the first user's AI voice response system requesting access authorization to a knowledge corpus electronically stored by the second user's AI voice response system and receiving a voice response from the second user. The method includes instantiating, by the first user's AI voice response system, an electronic communication session with the second user's AI voice response system based on the voice response, the electronic communication session being initiated by the first user's AI voice response system via an electronic communication connection with the second user's portable device. The method includes retrieving, by the first user's AI voice response system over a data communication network, a selected portion of the knowledge corpus from the second user's AI voice response system, the selected portion of the knowledge corpus being selected based on the second user's voice response to the voice request transmitted by the first user's AI voice response system. The method includes initiating an action by one or more IoT devices in response to a voice prompt interpreted by the first user's AI voice response system based on a selected portion of the knowledge corpus.

[0004] In one or more embodiments, a system includes a first user's artificial intelligence (AI) voice response system operably coupled to a processor configured to initiate operations. These operations include electronically perceiving the physical presence of a second user. These operations include transmitting a voice request to access an electronically stored knowledge corpus by the second user's AI voice response system and receiving a voice response from the second user. These operations include instantiating, by the first user's AI voice response system, an electronic communication session with the second user's AI voice response system based on the voice response, the electronic communication session being initiated via an electronic communication connection with the second user's portable device. These operations include retrieving a selected portion of the knowledge corpus from the second user's AI voice response system over a data communication network, the selected portion of the knowledge corpus being selected based on the second user's voice response to the request. These operations include initiating an action by one or more IoT devices in response to a voice prompt interpreted by the first user's AI voice response system based on the selected portion of the knowledge corpus.

[0005] In one or more embodiments, a computer program product includes one or more computer-readable storage media having instructions stored thereon. The instructions are executable by a processor and operably coupled to a first user's artificial intelligence (AI) voice response system to initiate operations. The operations include electronically perceiving the physical presence of a second user. The operations include transmitting a voice request to access an electronically stored knowledge corpus by the second user's artificial intelligence (AI) voice response system and receiving a voice response from the second user. The operations include instantiating, by the first user's voice response system, an electronic communication session with the second user's AI voice response system based on the voice response, the electronic communication session being initiated via an electronic communication connection with the second user's portable device. The operations include retrieving a selected portion of the knowledge corpus from the second user's AI voice response system over a data communication network, the selected portion of the knowledge corpus being selected based on the second user's voice response to the request. These operations include initiating actions by one or more IoT devices in response to voice prompts interpreted by the first user's AI voice response system based on selected portions of the knowledge corpus.

[0006] This "Summary" section is provided merely to introduce certain concepts and is not intended to identify key features or essential features of any of the claimed subject matter. Other features of the inventive structure will be apparent from the accompanying drawings and from the detailed description that follows.

[0007] Configurations of the present invention are illustrated by way of example in the accompanying drawings. However, these drawings should not be construed as limiting the configurations of the present invention to only the particular implementations shown. Various aspects and advantages will become apparent upon review of the following detailed description and upon reference to the drawings. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 illustrates an exemplary computing environment in which an AI voice response system is operable, according to an embodiment, where an intelligent virtual assistant is provided with the ability to securely access a remotely stored knowledge corpus. [Figure 2] 1 is a flowchart of a method for securely accessing a knowledge corpus generated by another AI voice response system using one AI voice response system, according to an embodiment. [Figure 3] FIG. 1 illustrates a cloud computing environment according to an embodiment. [Figure 4] FIG. 1 illustrates an abstract model layer according to an embodiment. [Figure 5] FIG. 1 illustrates a cloud computing node according to an embodiment. [Figure 6] FIG. 1 illustrates an exemplary portable device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] While the present disclosure concludes with claims defining novel features, it is believed that the various features described within this disclosure will be better understood by examining the description in conjunction with the drawings. The processes, machines, manufacture, and any variations thereof described herein are provided for illustrative purposes. The specific structural and functional details described within this disclosure should not be construed as limiting, but merely as a basis for the claims and as a representative basis for teaching those skilled in the art to employ the described features in various ways in substantially any appropriately detailed structure. Furthermore, the terms and phrases used within this disclosure are not intended to be limiting, but rather to provide an understandable description of the described features.

[0010] This disclosure relates to computer-based data exchange, and more particularly to secure data access and exchange between artificially intelligent computer systems.

[0011] Various kinds of automated systems and devices, ranging from individual devices to multiple devices integrated into the IoT, are increasingly being powered by AI. Intelligent virtual assistants and chatbots, for example, implement natural language understanding (NLU) to create interfaces between humans and computers, where language structure and meaning can be determined by a machine from human utterances.

[0012] AI voice response systems rely on a knowledge corpus to recognize and execute voice, text, or voice-to-text, or combination, prompts and commands to control the voice-enabled AI system and to respond to information queries. As defined herein, a “knowledge corpus” is a data structure containing a collection of words, phrases, and sentences with syntax and meaning that express an individual's intentions and preferences. Intentions and preferences correspond to linguistic (voice or text) commands for a device to perform an action and a preferred way for the device to perform the action, respectively. For example, a linguistic utterance such as “Prepare Bill a cup of coffee the way he likes it” includes a command for an automatic coffee maker to prepare a cup of coffee and a preference that the coffee be prepared according to Bill's preferences (e.g., strong, with cream, and no sugar). Having built a knowledge corpus from past usage by “training,” using machine learning to recognize speakers and understand the syntax and meaning of utterances, an AI voice response system can respond to commands by setting the automatic coffee maker's parameters according to Bill's preferences.

[0013] An "artificial intelligence voice response system (AIVRS)," as defined herein, is any system that provides a computer-human interface in which linguistic structure and meaning are determined by a machine based on human speech to select, create, or modify data, or a combination thereof, used to control an automated device or system. The system uses machine learning to learn the intent and preferences of a device or system user from processing the user's speech, monitoring instances of user activity, and creating a corresponding history or pattern of device or system use. An AIVRS may be implemented, for example, in software running on a cloud-based server. An AIVRS running on a cloud-based server can interact with a user, for example, by connecting to a smart speaker via a data communications network (e.g., the Internet). The smart speaker may include a speaker and a virtual assistant or chatbot to facilitate verbal interaction with the AIVRS user.

[0014] Different users have different voice characteristics, use different utterances, have different activity patterns, and therefore different knowledge corpora. A particular knowledge corpus typically corresponds to a specific individual. For privacy and / or data security reasons, access to a particular user's knowledge corpus may be restricted to that individual. While this protects privacy and provides security, such restrictions may limit a user's ability to use another AIVRS, because someone else's AIVRS is typically excluded from accessing the former user's knowledge corpus. Without such access, the latter's AIVRS would be blind to the user's intentions and preferences and would be unable to recognize the user's commands or prompts. For example, consider a situation in which an individual visits a friend's home or a colleague's office. The friend or colleague may invite the individual and use the friend's or colleague's AIVRS to control IoT devices in the smart home or smart office based on the individual's personal preferences regarding things like room temperature, lighting, background music, or the amount of sugar in coffee prepared by the automatic coffee machine. Without access to a personal knowledge corpus, an individual cannot utilize the AIVRS of a friend or colleague to control any of these automated functions.

[0015] Aspects of the systems, methods, and computer program products disclosed herein instantiate communication sessions between multiple AIVRSs to facilitate dynamically providing selected portions of one user's knowledge corpus to another user's AIVRS. The provision of the selected portions of the knowledge corpus enables the first user's AIVRS to capture intents and preferences learned by the second user's AIVRS and to respond to the second user's prompts and commands while maintaining security of the second user's knowledge corpus.

[0016] In one configuration, the AIVRS of a first user recognizes that a command originates from or is associated with a second user and that the AIVRS does not possess a knowledge corpus related to the second user to implement the command. Accordingly, the AIVRS of the first user attempts to obtain a knowledge corpus (or a relevant portion of a knowledge corpus) to implement the command by enlisting a nearby portable device (e.g., a smartphone, a smartwatch) carried or worn by the second user. If the AIVRS of the first user determines that authorization is required and obtains authorization from the second user, the AIVRS of the first user instantiates a communication session with the AIVRS of the second user and obtains access rights to the selected portion of the knowledge corpus generated by the AIVRS of the second user. The AIVRS of the first user can instantiate the communication session via a link on a data communications network.

[0017] The arrangements described herein are directed to and provide improvements over existing computer technology. The arrangements improve computer technology by enabling AIVRSs to function efficiently on behalf of a wider range of users. The computer technology is further improved by reducing communication overhead (e.g., required message exchanges) to facilitate acquisition of a knowledge corpus for the AIVRS to operate on behalf of a wider range of users. A further improvement is the exchange of relevant portions of the knowledge corpus while maintaining the privacy of the wider users and the security of the knowledge corpus corresponding to the wider users.

[0018] Further aspects of the embodiments described within this disclosure will now be described in more detail with reference to the figures. For purposes of simplicity and clarity of illustration, the elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Furthermore, where considered appropriate, reference numerals have been repeated among the figures to indicate corresponding, similar, or similar features.

[0019] FIG. 1 illustrates an exemplary computing environment 100. Computing environment 100 illustratively includes AIVRS 102 and AIVRS 104. AIVRS 102 and AIVRS 104 may each be implemented in separate computer systems and / or separate computing nodes (not shown), such as computing node 500 (FIG. 5). Thus, each computing node may comprise, for example, a cloud-based server. AIVRS 102 and AIVRS 104 are both illustratively implemented in software, which may be stored in memory and executed on one or more processors of one or more computer systems, such as computer system 512 of computing node 500. Illustratively, AIVRS 102 is operably coupled to speech interface 106. The voice interface 106, in some embodiments, is a smart speaker including a speaker and an integrated virtual assistant for handling human-computer interactions and for controlling IoT devices (not explicitly shown) communicatively coupled to the voice interface 106 via a wired or wireless connection (e.g., Wi-Fi, Bluetooth). Control of the IoT devices is implemented through various voice prompts and commands captured by the voice interface 106 and recognized by the AIVRS 102 based on a knowledge corpus 108 created by the AIVRS 102 to correspond to the user. The knowledge corpus 108 is illustratively stored in the memory of the same computer system on which the AIVRS 102 executes or a computer system operably coupled to the computer system on which the AIVRS 102 executes.

[0020] The AIVRS 102 can reside on a computer node (e.g., a cloud-based server) remote from the voice interface 106, and the IoT device is operably coupled to the voice interface 106 and to the AIVRS 102 via the voice interface 106. In other embodiments, the AIVRS 102, voice interface 106, and knowledge corpus 108 can be integrated into a single device with sufficient computational resources in terms of processing power, memory storage, etc. to support each. The voice interface 106 and connected IoT device can be located, for example, within a user's smart home or smart office corresponding to the knowledge corpus 108. In certain embodiments in which the AIVRS 102 and voice interface 106 are located remotely from one another, the AIVRS 102 can be operably coupled to the voice interface 106 via a data communications network 110. The data communications network 110 can typically be the Internet. The data communications network 110 may be or include a wide area network (WAN), a local area network (LAN), or a combination thereof, or other type of network, and may include wired, wireless, fiber optic, or other connections, or a combination thereof.

[0021] The AIVRS 102 provides natural language understanding (NLU) that allows linguistic structure and meaning to be determined by a machine from human speech received by the voice interface 106. The voice commands and prompts received by the voice interface 106 can be used to select, create, and modify data stored in electronic memory to control various automated systems and IoT devices (e.g., appliances, climate controllers, lighting systems, entertainment systems, energy monitors). The voice interface 106 can provide voice output in response to voice input, where the voice output is generated based on the linguistic structure and meaning of the voice input as determined by the NLU provided by the AIVRS 102.

[0022] The AIVRS 102 uses machine learning to iteratively build a knowledge corpus 108 corresponding to a user based on the user's usage history or patterns of various automated systems or IoT devices. The AIVRS 102 stores the knowledge corpus as data stored in electronic memory. The data stored in electronic memory corresponding to a user is identified by an account and / or other authentication information associated with the user. Different knowledge corpora can be created for different users based on the user's usage patterns of automated systems and IoT devices, and the knowledge corpora are stored as data in electronic memory on a cloud-based server or other server. The AIVRS determines the user's unique identity using speech-based recognition and authentication. Concurrent with or prior to building the knowledge corpus 108, the AIVRS 102 uses machine learning to train the knowledge corpus 108 to recognize the corresponding user based on the user's unique voice characteristics (e.g., acoustic signal patterns). In response to a command or prompt issued as a human utterance, the AIVRS 102 identifies the user who pronounced the command or prompt (as other AI voice response systems do).

[0023] The AIVRS 102 includes a knowledge corpus accessor module (KCAM) 112. The KCAM 112 is illustratively implemented in software executing on one or more processors of a computer system of a computing node on which the AIVRS 102 is implemented. In other embodiments, the KCAM 112 may be implemented in a computer system remote from, but operably coupled to, the AIVRS 102. The KCAM 112 may, in yet other embodiments, be implemented in dedicated circuitry or a combination of software and circuitry.

[0024] The KCAM 112 operatively interfaces with other features of the AIVRS 102 to enable a user to participate in a voice interaction with the AIVRS 102, even though the AIVRS has not constructed a knowledge corpus corresponding to the user. The AIVRS 102 does not contain a user-aware knowledge corpus that enables the AIVRS 102 to recognize and implement the user's voice commands, but a user-aware knowledge corpus 114 exists, which is constructed by and operable with the AIVRS 104, but is separate from the AIVRS 102. The AIVRS 104 may run on a different computer system, on a different computing node (e.g., a cloud-based server) remote from the AIVRS 102, but even when running on the same computer system (e.g., a cloud-based server), the AIVRS 104 is distinct from the AIVRS 102. Nevertheless, the KCAM 112 in conjunction with the AIVRS 102 allows the user to participate in a voice interaction with the AIVRS 102 by accessing selected portions of the knowledge corpus 114 via the AIVRS 104 and providing the selected portions to the AIVRS 102 while maintaining the security of the knowledge corpus 114.

[0025] The KCAM 112 detects the physical presence of a user, and the AIVRS 102 does not store a knowledge corpus for this user. The user's physical presence may be detected according to one or more of the procedures described in detail below. The KCAM 112 responds to detecting the user's physical presence by causing the AIVRS 102 to generate a voice request to access a separately stored knowledge corpus corresponding to the user. The AIVRS 102 communicates the voice request to the user via the voice interface 106. If the user responds affirmatively with permission via a voice prompt received by the AIVRS 102 via the voice interface 106, the KCAM 112 causes the AIVRS 102 to access the appropriate knowledge corpus by instantiating an electronic communication session with another AIVRS. For example, at the direction of the KCAM 112, the AIVRS 102 instantiates an electronic communication session with the AIVRS 104, which includes the knowledge corpus 114.

[0026] The AIVRS 102 can initiate an electronic communication session through an electronic communication connection with a user's portable device 116, and the AIVRS 102 does not store a knowledge corpus about this user. The portable device 116 can be any device carried (e.g., a smartphone) or worn (e.g., a smartwatch) by the user. The portable device 116 can include one or more of the components of the exemplary portable device 600 (FIG. 6). The electronic communication connection can include a wireless connection 118 between the voice interface 106 and the user's portable device 116, via which authorization spoken by the second user to the portable device for transmission to the AIVRS 104 is wirelessly transmitted to the portable device 116. The electronic communication connection need not be primarily peer-to-peer. The electronic communication connection may include, for example, an authorization generated by the AIVRS 104 and communicated to the AIVRS 102 via the data communication network 110 (e.g., the Internet) in response to an authorization spoken by the second user into the portable device 116. The authorization provides the AIVRS 102 with information (e.g., an IP address) for instantiating the electronic communication session.

[0027] Based on the user's voice response, received by the speech interface 106, granting permission to access the knowledge corpus 114, the AIVRS 102 retrieves a selected portion of the knowledge corpus 114 from the AIVRS 104 via the data communications network 110. The knowledge corpus 114 constructed by the AIVRS 104 corresponds to the user, and the selected portion received by the AIVRS 102 can be a portion selected based on the user's voice response to a request that the KCAM 112 causes the AIVRS 102 to generate and communicate via the speech interface 106.

[0028] By granting permission for selected portions of the knowledge corpus 114, the user can command automated systems and IoT devices controlled by the AIVRS 102 using utterances recognized by the AIVRS 102 based on the selected portions of the knowledge corpus 114 obtained from the AIVRS 104.

[0029] In a situation involving two individuals in the same environment (e.g., a smart home or a smart office) containing IoT devices that can be controlled by voice commands recognized by the AIVRS 102, for example, the knowledge corpus 108 may be machine-learned to recognize commands spoken by only one of the users. The AIVRS 102 may correspond to the first user and have built the knowledge corpus 108 by machine-learning to recognize the voice and prompts of the first user, but may not recognize the voice and prompts of the second user. The first user may be a normal (and legitimate) user, such as the owner (or family member) of the smart home or the occupant of the smart office, and the AIVRS 102 controls the IoT devices by recognizing voice commands from electronically stored data in the knowledge corpus 108.

[0030] A second user (e.g., Bill) may be a guest or colleague of a first user (e.g., Abel) visiting an environment (e.g., a smart home or smart office) where an IoT device (e.g., an automatic coffee maker) is voice-controlled by the AIVRS 102. If the first user (Abel) activates the AIVRS 102 using voice prompts and speaks the command, "Prepare Bill a cup of coffee according to Bill's preferences," the AIVRS 102 can respond to this command only if the AIVRS 102 has access to a knowledge corpus containing the second user's (Bill's) coffee preferences. If the AIVRS 102 cannot respond to this command, the KCAM 112 can instruct the AIVRS 102 to generate a speech output communicated via the voice interface 106 stating, "Sorry, I don't know Bill's coffee preferences," and to begin the process of accessing the appropriate knowledge corpus to be able to effectively respond to this command.

[0031] The KCAM 112 can operatively instruct the AIVRS 102 to recognize the presence of a second user, and the KCAM 112 can determine whether a knowledge corpus corresponding to the second user is electronically stored as data in an electronic memory accessible by the AIVRS 102. If so, the AIVRS 102 can respond to commands according to the second user's intentions and preferences based on the electronically stored data containing the knowledge corpus corresponding to the second user. If not, the KCAM 112 causes the AIVRS 102 to generate a request requesting access authorization to a knowledge corpus corresponding to the second user that was generated and stored by a different AIVRS.

[0032] In particular embodiments, using the access permissions for the selected portion of the knowledge corpus 114, and based on the context of the permissions granted via the second user's voice response, the AIVRS 102 accesses data (e.g., intents and preferences) to create an ontology graph tree for performing a specific action (e.g., automated coffee preparation). The ontology graph tree or other knowledge graph may be created by the AIVRS 102 by first obtaining context data for resource authentication and requesting access permissions from the AIVRS 104, as directed by the KCAM 112. The AIVRS 102 may retrieve the selected portion of the knowledge corpus 114 from the AIVRS 104 and electronically store (at least temporarily) the selected portion. The AIVRS 102 may create a usage pattern corresponding to the second user based on the accessed portion of the knowledge corpus 114. In particular embodiments, the AIVRS 102 implements a deep learning neural network to recognize usage patterns and intents and preferences. For example, the AIVRS 102 can implement a bidirectional LSTM (a combination of forward and backward recurrent neural networks that aids in statistical learning of dependencies between individual steps in a time series or data sequence (e.g., words in a sentence). The bidirectional LSTM can be used with reasoning hops to determine contextual data, including pattern history, about a second user.

[0033] The second user can speak a command including an access permission to grant permission to access the separately stored knowledge corpus 114. The AIVRS 104 that generated the knowledge corpus 114 identifies the second user based on speech recognition and analyzes the second user's voice command. Based on the voice command, the AIVRS 104 associated with the knowledge corpus 114 identifies a selected portion to which access permission is granted by performing natural language processing. Once the AIVRS 102 has obtained access permission to the selected portion, it can build a pattern history, intentions, preferences, etc. contained in the selected portion, thereby enabling the AIVRS 102 to respond to one or more voice commands of the second user (e.g., "Make me coffee the way I enjoy my coffee") and / or one or more voice commands of the first user that refer to the second user (e.g., "Make Bill coffee the way Bill enjoys his coffee").

[0034] In particular embodiments, KCAM 112 can sense the physical presence of a second user in various ways. For example, as previously described, a first user for whom AIVRS 102 has already created a knowledge corpus may utter a command referencing a second user for whom AIVRS 102 has not created a knowledge corpus. KCAM 112, using natural language processing, can determine that this reference is to a name (e.g., "Bill") that is not currently associated with a stored knowledge corpus and can initiate the knowledge corpus access process previously described.

[0035] In other embodiments, the KCAM 112 can distinguish, based on speech recognition, between the voice of a first user for whom the AIVRS 102 has already created a knowledge corpus and the voice of an individual (second user) for whom no knowledge corpus is available to the AIVRS 102. In yet other embodiments, the AIVRS 102 can be operatively coupled to one or more IoT cameras. Based on facial recognition applied to facial images captured by the one or more IoT cameras, the KCAM 112 can distinguish between a first user whose facial image corresponds to someone for whom the AIVRS 102 has already created a knowledge corpus and a second user whose facial image corresponds to someone with whom the AIVRS 102 has not previously interacted.

[0036] In any of a variety of circumstances, the KCAM 112 can initiate the previously described knowledge corpus access process in response to the AIVRS 102 detecting the presence of a user for whom the AIVRS 102 does not have access authorization to an existing knowledge corpus. For example, in an embodiment in which presence is detected based on facial recognition, the KCAM 112 can be configured to communicate with one or more IoT cameras using a machine-to-machine communication protocol via a representational state transfer (REST) ​​application program interface (API) to determine visual context data using a fast convolutional neural network (CNN). The KCAM 112 directs the AIVRS 102 to create a knowledge graph based initially on obtaining the context data via the fast CNN for resource authentication and generating a prompt to retrieve a selected portion of a separately stored knowledge corpus (e.g., knowledge corpus 114) and store (at least temporarily) the corresponding user's selected intents and preferences. The KCAM 112 can instruct the AIVRS 102 to create a pattern history based on the repeated application of a bidirectional LSTM model used with inference hops to include the corresponding user's pattern history as well as intent and preferences based thereon to respond to the voice command using the selected portion.

[0037] Optionally, after the AIVRS 102 is granted permission to access a remotely located knowledge corpus, the granting user can use a voice command to further authorize the AIVRS 102 to store selected portions (e.g., on a cloud-based server) for use in subsequent sessions. The granting user can specify, via voice command, which portions can be stored and for how long. For example, the user can specify that the AIVRS 102 will have permanent access to the user's coffee-making preferences, but that certain other intentions and preferences will be discarded within a specified time period (e.g., two hours). In default mode, the AIVRS 102 discards the retrieved selected portions of the knowledge corpus after a predetermined time interval (e.g., 24 hours) has elapsed without receiving a retention prompt (or other voice command) from the granting user's initial access authorization.

[0038] In certain embodiments, the KCAM 112 instructs the AIVRS 102 to initiate generation and transmission of a knowledge corpus access request only after first determining that the sensed user is not associated with an already stored knowledge corpus. Thus, for a user who has already granted the AIVRS 102 permission to store selected portions of the knowledge corpus, a new communication session need not be instantiated between the AIVRS 102 and another AIVRS to access the remotely stored knowledge corpus. In other embodiments, the user need not authorize the storage of selected portions of the knowledge corpus, but instead may authorize the AIVRS 102 to instantiate a communication session and store active permission to only selectively obtain access permissions to the knowledge corpus. Each session may be initiated in response to a voice command from the user requesting access to the remotely located knowledge corpus from the AIVRS 102. The voice commands may specify the selection of different portions of the knowledge corpus in each separate session.

[0039] For example, in either of the above situations involving a second user, the second user, having granted the AIVRS 102 access to the knowledge corpus 114, may further specify that the AIVRS 102 may again access the knowledge corpus 114 in subsequent sessions only in response to new voice prompts, and only with respect to selected portions specified in each new prompt. Thus, in a later session, the second user may issue a new prompt granting the AIVRS 102 access to the second user's coffee-making preferences for use in controlling an automatic coffee maker in response to a voice command. In an even later session, the second user may issue a permission for the AIVRS 102 to again access the knowledge corpus 114 only with respect to selected portions related to the second user's music preferences for use in controlling an entertainment system in response to a new voice command from the second user. In either case, however, to facilitate real-time or near real-time interaction, the AIVRS 102 does not have permission to store each preference (coffee or music preference) for more than a short period of time.

[0040] In some embodiments, the KCAM 112 can also incorporate certain security features for a first user, i.e., a user who regularly (and legitimately) interacts with the AIVRS 102 and for whom the AIVRS created the knowledge corpus 108. In one configuration, a first user can specify that the AIVRS 102 never respond to certain specified voice commands from a second user, even though the second user has given permission to access the second user's own knowledge corpus. For example, the AIVRS 102 can be excluded (e.g., via another voice command) from modifying or disabling the security system except in response to the first user's voice command. Additionally, the KCAM 112 can enable a first user to create a circle of trust. The circle of trust specifies identification information (e.g., based on voice recognition or facial recognition or both) and specific commands, allowing the AIVRS 102 to respond to the commands based on who is issuing the commands. The web of trust can optionally be hierarchical, whereby users at different tiers are permitted to issue different types of voice commands: for example, a visiting friend who provides access to his or her own corpus of knowledge may have greater command latitude (e.g., first tier) than a mere acquaintance (e.g., second tier).

[0041] In some embodiments, the KCAM 112 may search a publicly accessible knowledge corpus or portion of a knowledge corpus prior to perceiving the presence of a user for whom the AIVRS 102 does not have an existing knowledge corpus. The KCAM 112 may initiate a search in response to data indicating the future presence of an individual for whom the AIVRS 102 does not have an existing knowledge corpus. For example, the KCAM 112 may be configured to intermittently (e.g., daily) search an IoT calendar system. By searching an IoT calendar system or other data source associated with a first user, the KCAM 112 may obtain data (e.g., date and time) that a second user for whom the AIVRS 102 does not have an existing knowledge corpus is expected to be present within the operational vicinity of the first user's IoT device (e.g., within a smart home or smart office). In searching for the second user's publicly accessible intentions and preferences, KCAM 112 may, for example, search one or more databases of the first user (e.g., computer notes about previous encounters with the second user) or publicly accessible sites associated with the second user's name (e.g., social networking sites, websites of related professional or business organizations).

[0042] A calendar system or other data source associated with the first user may indicate the nature of the second user's future presence. For example, the calendar may specify that the first user will meet the second user socially or for business. Because the KCAM 112 reasonably anticipates that the second user's intentions and preferences regarding coffee and / or temperature will be present in voice commands that the second user may provide to the AIVRS 102 to the extent permitted by the KCAM 112, the KCAM 112 may initiate a search of a publicly accessible knowledge corpus to identify those intentions and preferences. The intentions and preferences may be stored by the AIVRS 102 for responding to the second user's voice commands representing preferences regarding coffee prepared by an automatic coffee maker and / or temperature settings for a central air conditioning system when the first user as a host prompts the second user with voice commands regarding coffee and / or temperature.

[0043] Optionally, in the absence of retrieval of information from a publicly available knowledge corpus, KCAM 112 may instruct AIVRS 102 to generate and communicate (e.g., via cellular or data communications network 116) a text or other electronic message, such as an email, to the second user requesting prior access to the knowledge corpus of the second user's AIVRS (e.g., AIVRS 104). Prior access enables AIVRS 102 to store the second user's intents and preferences within the operative vicinity of speech interface 106 for use during the next interaction between the first and second users.

[0044] 2 is a flowchart of a method 200 for securely accessing a knowledge corpus generated by one AIVRS by another AIVRS, according to an embodiment. Method 200 may be performed by a system including a KCAM operatively coupled to or integrated with a first user's AIVRS, as described with reference to FIG. 1. At block 202, the system uses the first user's AIVRS to electronically perceive the physical presence of a second user.

[0045] In block 204, the system transmits a voice request generated by the AIVRS of the first user requesting access authorization to the knowledge corpus electronically stored by the AIVRS of the second user. The system can receive a voice response from the second user. The voice response of the second user can grant access authorization to the knowledge corpus electronically stored by the AIVRS of the second user.

[0046] In block 206, based on the voice response, the system instantiates an electronic communication session with the second user's AIVRS using the first user's AIVRS. The electronic communication session may be initiated by the first user's AIVRS through an electronic communication connection with the second user's portable device. The portable device may be, for example, a smartphone, smartwatch, or other such device carried or worn by the second user. The electronic communication connection between the first user's AIVRS and the portable device may include a wireless connection between a voice interface of the first user's AIVRS and the portable device. Authorization by the second user spoken into the portable device for transmission to the second user's AIVRS may be wirelessly transmitted to the voice interface. The electronic communication connection may include authorization transmitted by the portable device to the second user's AIVRS and the first user's AIVRS via, for example, a data communication network (e.g., the Internet). This authorization provides the first user's AIVRS with the network information (eg, IP address) needed to instantiate an electronic communication session.

[0047] In block 208, the system retrieves, by the AIVRS of the first user, a selected portion of the knowledge corpus from the AIVRS of the second user over the data communications network, the selected portion of the knowledge corpus being selected based on the second user's vocal response to the vocal request communicated by the AIVRS of the first user.

[0048] At block 210, the system initiates one or more actions by one or more IoT devices in response to a voice command (or other voice prompt). The voice command may be spoken by a second user or by a first user referencing a second user. The voice command is interpreted by the first user's AIVRS based on a selected portion of the knowledge corpus.

[0049] In some embodiments, the system transmits a voice request in response to searching, by the first user's AIVRS, data stored in electronic memory and determining that the data stored in electronic memory lacks the selected portion of the knowledge corpus in addition to the electronically stored previous authorizations. The system retrieves the selected portion of the knowledge corpus over the data communications network in response to determining that the data stored in electronic memory includes the electronically stored previous authorizations and receiving a response to the request for access rights to the knowledge corpus. In response to determining that the data stored in electronic memory already includes the selected portion of the knowledge corpus, the system ceases retrieving the selected portion over the data communications network.

[0050] In some embodiments, the system perceives the second user by using the first user's AIVRS to recognize the first user's voice prompts and using the first user's AIVRS to identify a reference to a second user whose identity is not already known to the first user's AIVRS. The system additionally or alternatively, in certain embodiments, perceives the second user by performing visual recognition of the second user based on images captured by an IoT camera operably coupled to the first user's AIVRS.

[0051] In some embodiments, the system can use the first user's AIVRS to capture data from the first user's device communicatively coupled to the first user's AIVRS to obtain information about the second user. This data can correspond to the second user's intentions and preferences (e.g., coffee-making preferences) that the first user's AIVRS can use to respond to the second user's voice commands. For example, the first user's device can be a computer that stores email exchanges between the first user and the second user in which the second user describes what kind of coffee the second user likes. Or, for example, the first user's device can be a smartphone that stores text messages describing the type of music the second user likes. The first user's AIVRS can use this data to build its own knowledge corpus about the second user, which can be used by the first user's AIVRS to respond to a voice command from the second user requesting that an automatic coffee maker prepare coffee for the second user or that an entertainment system play music.

[0052] The system can discard a selected portion of the second user's knowledge corpus received by the first user's AIVRS. The system can automatically discard the selected portion if a predetermined time interval has elapsed without the system receiving a retention prompt from the second user. If the second user wants the first user's AIVRS to electronically store the selected portion, the second user can pronounce a command specifying a period of time for electronic storage of the selected portion.

[0053] Although this disclosure includes detailed descriptions of cloud computing, it is expressly noted that implementation of the subject matter presented herein is not limited to cloud computing environments. Rather, embodiments of the present invention may be implemented in conjunction with any other type of computing environment now known or later developed.

[0054] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computational resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) and for rapidly provisioning and releasing these resources with minimal administrative effort or interaction with a service provider. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0055] The features are as follows:

[0056] On-Demand Self-Service: Cloud consumers can automatically and unilaterally provision server time, network storage, and other computing capabilities as needed, without the need for human interaction with the service provider.

[0057] Broad network access: Functionality is available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin-client or thick-client platforms (e.g., mobile phones, laptops, and PDAs).

[0058] Resource Pooling: Providers' computing resources are pooled to serve multiple consumers using a multi-tenant model, with various physical and virtual resources dynamically allocated and reallocated according to demand. Consumers typically have no control or knowledge regarding the exact location of the resources provided, although at higher levels of abstraction there is a sense of location independence in that a location (e.g., country, state, or data center) may be specified.

[0059] Rapid Elasticity: Capabilities can be provisioned quickly and elastically, sometimes automatically, to scale out quickly, and released quickly to scale in quickly. These capabilities available for provisioning often appear to the consumer as unlimited, and any quantity can be purchased at any time.

[0060] Metered Services: Cloud systems automatically control and optimize resource usage by leveraging metering at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of utilized services.

[0061] The service model is as follows:

[0062] SaaS (Software as a Service): This is the ability offered to consumers to use a provider's applications running on a cloud infrastructure. These applications can be accessed from a variety of client devices through a thin-client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, except for the possibility of limited user-specific application configuration settings.

[0063] PaaS (Platform as a Service): This is the capability offered to consumers to deploy applications they create or acquire, written using programming languages ​​and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of the application hosting environment.

[0064] Infrastructure as a Service (IaaS): This functionality is provided to consumers to provision processing, storage, network, and other basic computing resources onto which they can deploy and run any software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do have control over the operating system, storage, deployed applications, and in some cases, limited control over selected network components (e.g., host firewalls).

[0065] The deployment model is as follows:

[0066] Private Cloud: This cloud infrastructure is operated solely for the organization. It can be managed by the organization or a third party and can reside on-premise or off-premise.

[0067] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with shared concerns (e.g., mission, security requirements, policy, and compliance considerations). This cloud infrastructure can be managed by these organizations or a third party and can reside on-premises or off-premises.

[0068] Public Cloud: This cloud infrastructure is made available to the general public or large industry organizations and is owned by an organization that sells cloud services.

[0069] Hybrid cloud: This cloud infrastructure is a composite of two or more clouds (private, community, or public) that maintain their unique identity and are bound together by standardized or proprietary technologies that allow data and application portability (e.g., cloud bursting for load balancing between clouds).

[0070] A cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that consists of a network of interconnected nodes.

[0071] Referring now to FIG. 3, an exemplary cloud computing environment 300 is illustrated. As illustrated, the cloud computing environment 300 includes one or more cloud computing nodes 310 with which local computing devices used by cloud consumers (e.g., personal digital assistant (PDA) or mobile phone 340a, desktop computer 340b, laptop computer 340c, and / or automotive computer system 340n) may communicate. The computing nodes 310 may communicate with each other. The computing nodes 310 may be physically or virtually grouped in one or more networks (not shown), such as into a private cloud, community cloud, public cloud, or hybrid cloud, or combinations thereof, as previously described herein. This enables the cloud computing environment 300 to provide an infrastructure, platform, and / or SaaS that does not require cloud consumers to maintain resources on their local computing devices. The types of computing devices 340a-n shown in FIG. 3 are intended to be illustrative only, and it is understood that computing node 310 and cloud computing environment 300 can communicate with any type of computer-controlled device via any type of network and / or network-addressable connection (e.g., a connection using a web browser).

[0072] Referring now to Figure 4, a set of functional abstraction layers provided by cloud computing environment 300 (Figure 3) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 4 are intended to be illustrative only, and that embodiments of the present invention are not limited thereto. As shown in the figure, the following layers and corresponding functions are provided:

[0073] Hardware and software layer 460 includes hardware and software components. Examples of hardware components include mainframe 461, reduced instruction set computer (RISC) architecture-based server 462, server 463, blade server 464, storage device 465, and network and network components 466. In some embodiments, software components include network application server software 467 and database software 468.

[0074] The virtualization layer 470 comprises an abstraction layer that can provide virtual entities such as virtual servers 471, virtual storage 472, virtual networks including virtual private networks 473, virtual applications and operating systems 474, and virtual clients 475.

[0075] By way of example, management layer 480 may provide the following functions: Resource provisioning 481 dynamically procures computing and other resources used to execute tasks within the cloud computing environment. Metering and pricing 482 tracks costs as resources are utilized within the cloud computing environment and generates and sends bills for the utilization of those resources. By way of example, those resources may include application software licenses. Security verifies the identity of cloud consumers and tasks and protects data and other resources. User portal 483 provides consumers and system administrators with access to the cloud computing environment. Service level management 484 allocates and manages cloud computing resources to meet required service levels. Service level agreement (SLA) planning and execution 485 proactively prepares and procures cloud computing resources in anticipation of upcoming demands in accordance with SLAs.

[0076] The Workload Layer 490 illustrates examples of functionality available in a cloud computing environment. Examples of workloads and functionality that may be provided from this layer include mapping and navigation 491, software development and lifecycle management 492, virtual classroom instruction delivery 493, data analytics processing 494, transaction processing 495, and AIVRS 496 with which KCAM is integrated or operatively coupled.

[0077] 5 illustrates a schematic diagram of an example computing node 500. In one or more embodiments, computing node 500 is an example of a suitable cloud computing node. Computing node 500 is not intended to suggest any limitation as to the scope of use or functionality of the embodiments of the invention described herein. Computing node 500 may perform any of the functions described within this disclosure.

[0078] Computing node 500 includes computer system 512, which can operate in numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, suitable for use with computer system 512 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of these systems or devices.

[0079] Computer system 512 may be described in the general context of computer system executable instructions, such as program modules being executed by the computer system. Typically, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer system 512 may be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.

[0080] As shown in FIG. 5, computer system 512 is illustrated in the form of a general-purpose computing device. Components of computer system 512 may include, but are not limited to, one or more processors 516, memory 528, and a bus 518 coupling various system components, including memory 528, to processor 516. As defined herein, a "processor" means at least one hardware circuit configured to execute instructions. The hardware circuit may be an integrated circuit. Examples of processors include, but are not limited to, a central processing unit (CPU), an array processor, a vector processor, a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA), an application-specific integrated circuit (ASIC), a programmable logic circuit, and a controller.

[0081] The execution of instructions of a computer program by a processor includes running the program. As defined herein, "run" and "execute" include a sequence of actions or events performed by a processor in accordance with one or more machine-readable instructions. "Running" and "executing," as defined herein, refer to the active performance of actions or events by a processor. The terms "run," "running," "execute," and "executing" are used synonymously herein.

[0082] Bus 518 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures, including, by way of example only, but not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, a Peripheral Component Interconnects (PCI) bus, and a PCI Express (PCIe) bus.

[0083] Computer system 512 typically includes a variety of computer system readable media. Such media can be any available media that can be accessed by computer system 512 and can include both volatile and nonvolatile media, removable and non-removable media.

[0084] Memory 528 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 530 and / or cache memory 532. Computer system 512 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example, a storage system 534 may be provided for reading from and writing to non-removable, non-volatile magnetic media and / or solid-state drives (not shown, typically referred to as "hard drives"). Although not shown, a magnetic disk drive may be provided for reading from and writing to removable, non-volatile magnetic disks (e.g., "floppy disks"), and an optical disk drive may be provided for reading from and writing to removable, non-volatile optical disks, such as CD-ROMs, DVD-ROMs, or other optical media. In such examples, each may be connected to bus 518 by one or more data media interfaces. As further shown and described below, memory 528 may include at least one program product comprising a series of (e.g., at least one) program modules configured to perform the functions of embodiments of the present invention.

[0085] For example, a program / utility 540 including a set of (at least one) program modules 542 may be stored in memory 528, including, but not limited to, an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or a combination thereof, may include an implementation of a network environment. The program modules 542 typically perform the functions and / or methods of embodiments of the invention described herein. For example, one or more of the program modules may include a portion of the AIVRS 596, within which the KCAM is integrated or operably coupled.

[0086] The programs / utilities 540 are executable by the processor 516. The programs / utilities 540, and any data items used, created, manipulated, or a combination thereof, performed by the computer system 512, are functional data structures that convey functionality when employed by the computer system 512. As defined within this disclosure, a "data structure" is the physical implementation of a data model's organization of data in physical memory. As such, a data structure is formed of specific electrical or magnetic structural elements in memory. The data structure imposes a physical organization on data stored in memory when used by an application program executed using the processor.

[0087] The computer system 512 may communicate with one or more external devices 514, such as a keyboard, pointing device, display 524, one or more devices that allow a user to interact with the computer system 512, or any device that allows the computer system 512 to communicate with one or more other computing devices (e.g., a network card, modem, etc.), or a combination thereof. Such communication may occur through an input / output (I / O) interface 522. Additionally, the computer system 512 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, through a network adapter 520. As shown, the network adapter 520 communicates with other components of the computer system 512 through a bus 518. It should be understood that other hardware and / or software components, not shown, may be used with the computer system 512. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.

[0088] While computing node 500 is used to illustrate an example of a cloud computing node, it should be understood that computer systems using architectures the same as or similar to those described in connection with FIG. 5 to perform various operations described herein may be used in implementations other than cloud computing. In this regard, the example embodiments described herein are not intended to be limited to cloud computing environments. Computing node 500 is an example of a data processing system. As defined herein, "data processing system" means one or more hardware systems configured to process data, each including at least one processor and memory programmed to initiate operations.

[0089] Computing node 500 is an example of computer hardware. Computing node 500 may include fewer components than those shown in FIG. 5 or additional components not shown in FIG. 5 depending on the particular type of device and / or system in which it is implemented. The particular operating system and / or applications may vary depending on the type of device and / or system, such as the types of I / O devices included. Furthermore, one or more of the example components may be incorporated into or otherwise form part of another component. For example, a processor may include at least some memory.

[0090] Computing node 500 is also an example of a server. As defined herein, a "server" refers to a data processing system configured to share services with one or more other data processing systems. As defined herein, a "client device" refers to a data processing system that requests shared services from a server, and a user interacts directly with the client device. Examples of client devices include, but are not limited to, workstations, desktop computers, computer terminals, mobile computers, laptop computers, netbook computers, tablet computers, smartphones, personal digital assistants, smart watches, smart glasses, gaming devices, set-top boxes, and smart televisions. In one or more embodiments, the various user devices described herein may be client devices. Network infrastructure such as routers, firewalls, switches, and access points are not client devices as the term "client device" is defined herein.

[0091] 6 illustrates an exemplary portable device 600 in accordance with one or more embodiments described within this disclosure. The portable device 600 may include a memory 602, one or more processors 604 (e.g., an image processor, a digital signal processor, a data processor), and an interface circuit 606.

[0092] In one embodiment, the memory 602, the processor 604, and / or the interface circuitry 606 are implemented as separate components. In another embodiment, the memory 602, the processor 604, and / or the interface circuitry 606 are integrated into one or more integrated circuits. The various components of the portable device 600 may be coupled, for example, by one or more communication buses or signal lines (e.g., interconnects and / or wires). In one embodiment, the memory 602 may be coupled to the interface circuitry 606 via a memory interface (not shown).

[0093] To facilitate the functions and / or operations described herein, including generation of sensor data, sensors, devices, subsystems, and / or input / output (I / O) devices may be coupled to interface circuit 606. Various sensors, devices, subsystems, and / or I / O devices may be coupled to interface circuit 606 directly or through one or more intervening I / O controllers (not shown).

[0094] For example, a position sensor 610, a light sensor 612, and a proximity sensor 614 may be coupled to the interface circuit 606 to facilitate orientation, lighting, and proximity functions, respectively, of the portable device 600. A position sensor 610 (e.g., a GPS receiver and / or a GPS processor) may be connected to the interface circuit 606 to provide geopositioning sensor data. An electronic magnetometer 618 (e.g., an integrated circuit chip) may be connected to the interface circuit 606 to provide sensor data that can be used to determine the direction of magnetic north for directional navigation purposes. An accelerometer 620 may be connected to the interface circuit 606 to provide sensor data that can be used to determine changes in speed and direction of the device's movement in three dimensions. An altimeter 622 (e.g., an integrated circuit) may be connected to the interface circuit 606 to provide sensor data that can be used to determine altitude. A voice recorder 624 may be connected to the interface circuit 606 to store recorded speech.

[0095] The camera subsystem 626 may be coupled to a light sensor 628. The light sensor 628 may be implemented using any of a variety of technologies. Examples of the light sensor 628 include, for example, a charged coupled device (CCD) light sensor, a complementary metal-oxide semiconductor (CMOS) light sensor, etc. The camera subsystem 626 and the light sensor 628 may be used to facilitate camera functions, such as recording images and / or video clips (hereinafter, "image data"). In one aspect, the image data is a subset of the sensor data.

[0096] Communication functions may be facilitated by one or more wireless communication subsystems 630. The wireless communication subsystems 630 may include radio frequency receivers and transmitters, optical (e.g., infrared) receivers and transmitters, etc. The particular design and implementation of the wireless communication subsystem 630 may depend on the particular type of portable device 600 being implemented and / or the communication network with which the portable device 600 is intended to operate.

[0097] By way of example, the wireless communications subsystem 630 may be designed to operate over one or more mobile networks (e.g., GSM, GPRS, EDGE), a Wi-Fi network that may include a WiMax network, or a short-range wireless network (e.g., a Bluetooth network), or any combination thereof. The wireless communications subsystem 630 may implement a hosting protocol such that the portable device 600 may be configured as a base station for other wireless devices.

[0098] An audio subsystem 632 may be coupled to a speaker 634 and a microphone 636 to facilitate voice-enabled functions such as voice recognition, voice replication, digital recording, voice processing, and telephony. The audio subsystem 632 may generate audio type sensor data. In one or more embodiments, the microphone 636 may be utilized as a mask sensor.

[0099] I / O devices 638 may be coupled to the interface circuit 606. Examples of I / O devices 638 include, for example, a display device, a touch-sensitive display device, a trackpad, a keyboard, a pointing device, a communication port (e.g., a USB port), a network adapter, a button, or other physical control. A touch-sensitive device, such as a display screen and / or a pad, is configured to detect contact, movement, interruption of contact, etc. using any of a variety of touch sensitivity technologies. Exemplary touch-sensitive technologies include, for example, capacitive, resistive, infrared, and surface acoustic wave technologies, other proximity sensor arrays, or other elements for determining one or more locations of contact with the touch-sensitive device. One or more of the I / O devices 638 may be adapted to control functions of sensors, subsystems, etc. of the portable device 600.

[0100] Portable device 600 further includes a power supply 640. Power supply 640 can provide power to various elements of portable device 600. In one embodiment, power supply 640 is implemented as one or more batteries. The batteries may be implemented using any of a wide variety of battery technologies, whether disposable (e.g., replaceable) or rechargeable. In another embodiment, power supply 640 is configured to obtain power from an external power source and provide power (e.g., DC power) to elements of portable device 600. In the case of rechargeable batteries, power supply 640 can further include circuitry capable of charging the one or more batteries when coupled to an external power source.

[0101] Memory 602 may include random access memory (e.g., volatile memory) and / or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, flash memory, etc. Memory 602 may store an operating system 652, such as LINUX, UNIX®, a mobile operating system, an embedded operating system, etc. Operating system 652 may include instructions for handling system services and for performing hardware-dependent tasks.

[0102] The memory 602 may store additional program code 654. Examples of other program code 654 may include instructions for facilitating communication with one or more additional devices, one or more computers, or one or more servers, or a combination thereof; processing instructions for facilitating graphic user interface processing; sensor-related functions; telephone-related functions; electronic messaging-related functions; web browsing-related functions; media processing-related functions; GPS and navigation-related functions; security functions; camera-related functions, including webcam and / or web video functions; and the like. The program code may include program code for implementing an intelligent virtual assistant (IVA) in the portable device, the intelligent virtual assistant being implemented in IVA program code 656 executing on the processor 604. The IVA may, for example, interact with an AIVRS running on a remote cloud-based server. The memory 602 may also store one or more other applications 662.

[0103] The various types of instructions and / or program code described are provided for purposes of illustration and not limitation. The program code may be implemented as separate software programs, procedures, or modules. The memory 602 may include additional or fewer instructions. Furthermore, various functions of the portable device 600 may be implemented in hardware and / or software, including one or more signal processing and / or application specific integrated circuits.

[0104] The program code stored in memory 602 and any data used, generated, or manipulated by portable device 600, or a combination thereof, are functional data structures that, when employed as part of the device, impart functionality to the device. Further examples of functional data structures include, for example, sensor data, data obtained by user input, data obtained by querying external data sources, baseline information, etc. The term "data structure" refers to the physical implementation of a data model's organization of data in physical memory. As such, a data structure is formed of specific electrical or magnetic structural elements in memory. A data structure imposes a physical organization on data stored in memory for use by a processor.

[0105] In particular embodiments, one or more of the various sensors and / or subsystems described with reference to portable device 600 may be separate devices coupled to or communicatively linked to portable device 600 via wired or wireless connections. For example, one or more (or all) of position sensor 610, light sensor 612, proximity sensor 614, gyroscope 616, magnetometer 618, accelerometer 620, altimeter 622, voice recorder 624, camera subsystem 626, audio subsystem 632, etc. may be implemented as separate systems or subsystems operably coupled to portable device 600 via I / O device(s) 638 and / or wireless communication subsystem 630.

[0106] Portable device 600 may include fewer components than those shown in Figure 6 or may include additional components other than those shown in Figure 6, depending on the particular type of system in which it is implemented. Furthermore, the particular operating system and / or applications and / or other program code included may vary according to the type of system. Furthermore, one or more of the illustrated components may be incorporated into or otherwise form part of another component. For example, a processor may include at least some memory.

[0107] Portable device 600 is provided for purposes of illustration and not limitation. Devices and / or systems configured to perform the operations described herein may have architectures different from that shown in FIG. 6. The architecture may be a simplified version of portable device 600 and may include a processor and memory storing instructions. The architecture may include one or more sensors as described herein. Portable device 600 or a similar system may collect data using various sensors of the device or sensors coupled to the device. However, it should be understood that portable device 600 may include fewer sensors or additional sensors. In this disclosure, data generated by a sensor is referred to as “sensor data.”

[0108] Examples of implementations of portable device 600 include, for example, a smartphone or other mobile or cellular phone, a wearable computing device (e.g., a smartwatch), a dedicated medical device, or other suitable handheld, wearable, or comfortably portable electronic device that can sense and process signals and data detected by sensors. It will be understood that embodiments can be deployed as a standalone device or as multiple devices in a distributed client / server network system. For example, in certain embodiments, a smartwatch can be operably coupled to a mobile device (e.g., a smartphone). The mobile device may or may not be configured to interact with a remote server and / or computer system.

[0109] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. Nevertheless, several definitions are now presented that apply throughout this document.

[0110] As defined herein, the singular forms "a," "an," and "the" include the plural forms as well, unless the context clearly indicates otherwise.

[0111] As defined herein, "another" means at least a second or more.

[0112] As defined herein, unless expressly stated otherwise, "at least one," "one or more," and "or or combinations thereof" are open-ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, and C," "one or more of A, B, or C," and "A, B, or C, or combinations thereof" means A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B, and C.

[0113] As defined herein, "automatically" means without user intervention.

[0114] As defined herein, "comprises," "including," "comprises," or "comprising," or combinations thereof, indicate the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.

[0115] As defined herein, "when" means "in response to" or "responsive to," depending on the context. Thus, the phrase "when it is determined" may be interpreted to mean "in response to determining" or "responsive to determining," depending on the context. Similarly, the phrase "when a stated condition or event is detected" may be interpreted to mean "upon detecting a stated condition or event" or "in response to detecting a stated condition or event" or "responsive to detecting a stated condition or event," depending on the context.

[0116] As defined herein, the terms "in one embodiment," "embodiment," "in one or more embodiments," "in a particular embodiment," or similar phrases mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment described within the disclosure. Thus, throughout this disclosure, appearances of the foregoing phrases and / or similar phrases may, but do not necessarily, all refer to the same embodiment.

[0117] As defined herein, the phrases "in response to" and "responsive to" mean to respond or react readily to an action or event. Thus, when a second action is performed "in response to" or "responsive to" a first action, a causal relationship exists between the occurrence of the first action and the occurrence of the second action. The phrases "in response to" or "responsive to" indicate a causal relationship.

[0118] As defined herein, "real-time" means a level of processing responsiveness that a user or system perceives as sufficiently immediate with respect to a particular process or decision being made, or that allows the processor to keep up with some external process.

[0119] As defined herein, "substantially" means that the recited property, parameter, or value need not be achieved exactly, but rather that deviations or variations may occur, including, for example, tolerances, measurement errors, limitations in measurement precision, or other factors known to those skilled in the art, in an amount that does not render the property unable to produce the effect it is intended to produce.

[0120] As defined herein, "user," "individual," and "guest" each refer to a human being.

[0121] In this specification, terms such as first, second, etc. may be used to refer to various elements. Unless otherwise stated or the context clearly indicates otherwise, these terms are used only to distinguish one element from another and should not be construed as limiting these elements.

[0122] The present invention may be a system, method, or computer program product, or any combination thereof, at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium containing computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0123] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, punch cards or mechanically encoded devices such as ridge structures in grooves in which instructions are recorded, and any suitable combination thereof. As used herein, a computer-readable storage medium should not itself be construed as a transitory signal such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted over a wire.

[0124] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). This network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.

[0125] Computer-readable program instructions for carrying out the operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages ​​such as Smalltalk®, C++, and procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, to carry out aspects of the present invention, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to customize the electronic circuitry.

[0126] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0127] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to create a machine, where the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium and capable of directing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions implementing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0128] Furthermore, the computer-readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device to generate a computer-implemented process, thereby causing a series of operable steps to be performed on the computer, other programmable apparatus, or other device, such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0129] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, comprising one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may actually be executed as a single step, executed concurrently, executed substantially concurrently in a partially or fully overlapping manner in time, or executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks included in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.

[0130] The description of various embodiments of the present invention is presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many changes and modifications will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used in this specification are selected to best explain the principles of the embodiments, practical applications, or technical improvements beyond those found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. electronically perceiving the physical presence of a second user using an artificial intelligence (AI) voice response system of the first user; Communicating a voice request generated by the AI ​​voice response system of the first user requesting access authorization to an electronically stored knowledge corpus by the AI ​​voice response system of the second user and receiving a voice response from the second user; instantiating, by the first user's AI voice response system, an electronic communication session with the second user's AI voice response system based on the voice response, wherein the electronic communication session is initiated by the first user's AI voice response system via an electronic communication connection with the second user's portable device; retrieving a selected portion of the knowledge corpus from the second user's AI voice response system by the first user's AI voice response system via a data communications network, wherein the selected portion of the knowledge corpus is selected based on the second user's voice response to the voice request communicated by the first user's AI voice response system; and initiating an action by one or more IoT devices in response to a voice prompt interpreted by the first user's AI voice response system based on the selected portion of the knowledge corpus.

2. 2. The computer-implemented method of claim 1, wherein the communicating is performed in response to: searching, by the first user's AI voice response system, data stored in electronic memory; and determining that the data stored in electronic memory lacks electronically stored prior authorization and the selected portion of the knowledge corpus.

3. performing said retrieval of said selected portion of said knowledge corpus over said data communications network in response to determining that data stored in electronic memory includes electronically stored prior authorization; 2. The computer-implemented method of claim 1, further comprising: in response to determining that the data stored in the electronic memory includes the selected portion of the knowledge corpus, ceasing to retrieve the selected portion over the data communications network.

4. 2. The computer-implemented method of claim 1, wherein the perceiving includes recognizing a voice prompt of the first user by an AI voice response system of the first user and identifying a reference to the second user by the AI ​​voice response system of the first user.

5. 2. The computer-implemented method of claim 1, wherein the perceiving comprises performing visual recognition of the second user based on an image captured by an IoT camera operably coupled to an AI voice response system of the first user.

6. 2. The computer-implemented method of claim 1, further comprising obtaining information of a second user by the first user's AI voice response system based on data captured by the first user's AI voice response system from a first user's device communicatively coupled to the first user's AI voice response system.

7. 10. The computer-implemented method of claim 1, wherein the first user's AI voice response system discards the retrieved selected portion of the knowledge corpus after a predetermined time interval has elapsed without receiving a hold prompt from the second user.

8. 1. A system comprising: a first user artificial intelligence (AI) voice response system operably coupled to a processor configured to initiate an action, the action comprising: Electronically perceiving the physical presence of a second user; Communicating a voice request to access an electronically stored knowledge corpus by an AI voice response system of the second user and receiving a voice response from the second user; instantiating, by the first user's AI voice response system, an electronic communication session with the second user's AI voice response system based on the voice response, wherein the electronic communication session is initiated via an electronic communication connection with the second user's portable device; Retrieving a selected portion of the knowledge corpus from the second user's AI voice response system via a data communications network, the selected portion of the knowledge corpus being selected based on the second user's voice response to the request; Initiating an action by one or more IoT devices in response to a voice prompt interpreted by the first user's AI voice response system based on the selected portion of the knowledge corpus; and Including, the system.

9. 10. The system of claim 8, wherein the communicating is performed in response to retrieving data stored in electronic memory and determining that the data stored in electronic memory lacks electronically stored prior authorization and the selected portion of the knowledge corpus.

10. the processor: performing said retrieval of said selected portion of said knowledge corpus over said data communications network in response to determining that data stored in electronic memory includes electronically stored prior authorizations; 9. The system of claim 8, configured to initiate further actions including ceasing said retrieval of said selected portion over said data communications network in response to determining that data stored in said electronic memory includes said selected portion of said knowledge corpus.

11. The system of claim 8 , wherein the perceiving includes recognizing a voice prompt of a first user and identifying a reference to the second user within the voice prompt.

12. The system of claim 8 , wherein the perceiving comprises performing visual recognition of the second user based on an image captured by a camera operably coupled to the system.

13. 10. The system of claim 8, wherein the processor is configured to initiate further operations including obtaining information of a second user based on data captured by a device of a first user communicatively coupled to the system.

14. A computer program product which, when run on a computer, causes the computer-implemented method of any one of claims 1 to 7 to be carried out.

Citation Information

Patent Citations

  • A unified framework for device configuration, interaction and control, and related methods, devices and systems

    JP2016502137A

  • Management server, device control system, device, terminal device, control method of management server, and control program

    JP2019036811A

  • Voice interactive system, method and program

    JP2019090944A

  • Unified framework for device configuration, interaction and control, and associated methods, devices and systems

    WO2014078480A1