System, method, and computer program for recommending guidance instructions to a user (self-learning artificial intelligence voice response based on user behavior during dialogue)

The AI user activity guidance system addresses inaccuracies in voice response systems by offering real-time feedback and guidance, improving task performance and efficiency.

JP7776227B2Active Publication Date: 2025-11-26INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021188054
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-20
Filing Date
2021-11-18
Publication Date
2025-11-26
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

Voice response systems often provide inaccurate actions and responses due to AI misinterpreting voice commands, leading to user frustration and inefficiency in tasks like playing a musical instrument or performing crafts.

Method used

An AI user activity guidance system that monitors user actions in real-time, provides tactile alerts for incorrect actions, and offers guidance through augmented images and spoken instructions to correct errors without interrupting the task.

Benefits of technology

Enables users to perform tasks accurately by correcting errors on the fly, reducing the time required to complete tasks and enhancing user experience by providing immediate feedback and guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007776227000001
    Figure 0007776227000001
  • Figure 0007776227000002
    Figure 0007776227000002
  • Figure 0007776227000003
    Figure 0007776227000003
Patent Text Reader

Abstract

To solve a problem that action taken by a voice response system and response provided to a user may sometimes be inaccurate based on user's intended context because AI used by a voice response system misunderstands words that configured a voice command.SOLUTION: A system is provided for recommending guidance instructions to a user. The system includes a memory having computer readable instructions and a processor for executing the same. The computer readable instructions control the processor to monitor an ongoing task that includes at least one operation performed by the user, generate image data indicative of the ongoing task, and perform operation of displaying the ongoing task based on the image data. The system analyzes the ongoing task to generate an extended image. The extended image is overlaid on the image data so that the extended image is displayed concurrently with the ongoing task in order to instruct the user to advance the ongoing task.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to artificial intelligence (AI) computing systems, and more particularly to AI recommendation of guiding instructions to a user. [Background technology]

[0002] Voice response systems are very popular. Typically, a voice response system monitors ambient sounds for voice commands from a user. The voice commands instruct the voice response system to perform a particular action and provide a voice response to the user. In some cases, the action taken by the voice response system and the response provided to the user may be inaccurate based on the user's intended context because the AI ​​used by the voice response system misunderstood the words that constituted the voice command. Summary of the Invention [Problem to be solved by the invention]

[0003] The actions taken by the voice response system and the responses provided to the user may be inaccurate based on the user's intended context because the AI ​​used by the voice response system misinterpreted the words that constituted the voice command. [Means for solving the problem]

[0004] According to a non-limiting embodiment, a system for recommending guidance instructions to a user is provided. The system includes a memory having computer-readable instructions and a processor for executing the computer-readable instructions. The computer-readable instructions control the processor to monitor an ongoing task including at least one action performed by a user, generate image data indicative of the ongoing task, and perform operations to display the ongoing task based on the image data. The system analyzes the ongoing task and generates an augmented image. The augmented image is overlaid on the image data such that the augmented image is displayed simultaneously with the ongoing task to instruct the user to advance the ongoing task.

[0005] According to another non-limiting embodiment, a method for recommending guidance instructions to a user is provided, the method including monitoring a task in progress including at least one action performed by a user, generating image data indicative of the task in progress, and displaying the task in progress based on the image data. The method further includes analyzing the task in progress, generating an augmented image, and overlaying the augmented image on the image data such that the augmented image is displayed simultaneously with the task in progress to instruct the user to advance the task in progress.

[0006] According to yet another non-limiting embodiment, a computer program product for recommending guidance instructions to a user is provided. The computer program product includes a computer-readable storage medium having program instructions embodied thereon. The program instructions are readable by the processing circuit to cause the processing circuit to perform operations including monitoring an ongoing task including at least one action performed by a user, generating image data indicative of the ongoing task, and displaying the ongoing task based on the image data. The operations further include analyzing the ongoing task, generating an augmented image, and overlaying the augmented image on the image data such that the augmented image is displayed concurrently with the ongoing task to instruct the user to advance the ongoing task.

[0007] Additional features and advantages are realized through the teachings of the present invention. Other embodiments and aspects of the invention are described in detail herein and are considered a part of the claimed invention. For a better understanding of the present invention, together with its advantages and features, reference is made to the description and drawings. [Brief explanation of the drawings]

[0008] The subject matter which is regarded as the invention is particularly pointed out and distinctly claimed in the claims at the conclusion of this specification. The foregoing and other features and advantages of the invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings.

[0009] [Figure 1] 1 illustrates a cloud computing environment in accordance with one or more embodiments of the present invention.

[0010] [Figure 2] 1 illustrates abstraction model layers in accordance with one or more embodiments of the present invention.

[0011] [Figure 3] 1 illustrates an exemplary computer system in which one or more embodiments of the present invention may be implemented.

[0012] [Figure 4A] 1 illustrates an AI user activity guidance system according to one or more embodiments of the present invention.

[0013] [Figure 4B] 1 illustrates an AI user activity guidance system according to one or more embodiments of the present invention.

[0014] [Figure 5] 1 illustrates learned modeled user activity according to one or more embodiments of the present invention.

[0015] [Figure 6]1 illustrates a user activity being performed erroneously, in accordance with one or more embodiments of the present invention.

[0016] [Figure 7] 10 shows an image of an incorrectly performed user activity overlaid on a correct modeled user activity, in accordance with one or more embodiments of the present invention.

[0017] [Figure 8] 1 illustrates learned modeled user activity in accordance with one or more embodiments of the present invention.

[0018] [Figure 9] 1 illustrates a user activity being performed erroneously, in accordance with one or more embodiments of the present invention.

[0019] [Figure 10] 10 shows an image of an incorrectly performed user activity overlaid with a correct modeled user activity, in accordance with one or more embodiments of the present invention.

[0020] [Figure 11] 10 illustrates an image of a performed user activity included in a task, according to one or more embodiments of the present invention.

[0021] [Figure 12] 12 illustrates an image of the performed user activity shown in FIG. 11 overlaid on the next user activity included in the task, in accordance with one or more embodiments of the present invention.

[0022] [Figure 13] FIG. 1 is a flow diagram illustrating a method performed by an AI user activity guidance system for recommending guidance instructions to a user, in accordance with one or more embodiments of the present invention.

[0023] [Figure 14] 1 illustrates a machine learning system that may be utilized to implement various embodiments of the present invention.

[0024] [Figure 15] 15 illustrates a learning phase that may be implemented by the machine learning system shown in FIG. 14.

[0025] [Figure 16] 1 illustrates an exemplary computing system in which various embodiments of the present invention may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0026] Various embodiments of the present invention are described herein with reference to the associated drawings. Alternative embodiments of the present invention may be devised without departing from the scope of the present invention. Various connection and positional relationships (e.g., above, below, adjacent, etc.) are described between elements in the following description and in the drawings. These connection or positional relationships, or combinations thereof, may be direct or indirect unless otherwise specified, and the present invention is not intended to be limited in this respect. Thus, connections between entities may refer to direct or indirect connections, and positional relationships between entities may be direct or indirect positional relationships. Furthermore, various tasks and process steps described herein may be incorporated into a broader procedure or process having additional steps or functions not described in detail herein.

[0027] The following definitions and abbreviations may be used for interpreting the claims and the specification. As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," "contains," or "containing," or any other variation thereof, are intended to cover an exclusive inclusion. For example, a composition, mixture, process, method, article, or device that includes a list of elements is not necessarily limited to only those elements and may include a list of other elements not expressly listed or inherent to such composition, mixture, process, method, article, or device.

[0028] Additionally, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" may be understood to include any integer number greater than or equal to one, i.e., 1, 2, 3, 4, etc. The term "plurality" may be understood to include any integer number greater than or equal to two, i.e., 2, 3, 4, 5, etc. The term "connected" may include both an indirect "connected" and a direct "connected."

[0029] The terms "about," "substantially," "approximately," and variations thereof are intended to include the degree of error associated with measurement of a particular quantity based on equipment available at the time of filing of this application. For example, "about" can include a range of ±8%, or 5%, or 2% of a given value.

[0030] For purposes of brevity, conventional techniques relating to making and using aspects of the present invention may or may not be described in detail herein. In particular, various aspects of computing systems and specific computer programs for implementing various technical features described herein are well known. Thus, for purposes of brevity, many conventional implementation details are mentioned only briefly herein or omitted entirely without providing details of well-known systems and / or processes.

[0031] Although this disclosure includes detailed descriptions related to cloud computing, it should be understood that implementation of the teachings recited herein is not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0032] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0033] The characteristics are as follows:

[0034] On-Demand Self-Service: Cloud consumers can unilaterally provision computing capacity, such as server time and network storage, automatically as needed, without requiring human interaction with the service provider.

[0035] Broad Network Access: Capabilities are available over the network and accessed through standard mechanisms (eg, cell phones, laptops, and PDAs) facilitating use by heterogeneous thin or thick client platforms.

[0036] Resource Pool: A provider's computing resources are pooled and serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated according to demand. Location independence means that consumers generally have no control or knowledge over the exact location of the resources provided, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0037] Rapid Elasticity: Capacity can be rapidly and elastically provisioned, in some cases automatically, for rapid scale out, and rapidly released for rapid scale in. To the consumer, the capabilities available for delivery often appear unlimited and can be purchased at any time and in any quantity.

[0038] Measured Services: Cloud systems automatically control and optimize resource usage by leveraging measurement capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource utilization can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services used.

[0039] The service model is as follows:

[0040] Software as a Service (SaaS): The consumer is offered the ability to use a provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through a thin-client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, with the possible exception of limited user-specific application configuration settings.

[0041] Platform as a Service (PaaS): The capability offered to consumers is the deployment of applications created using programming languages ​​and tools supported by the provider on cloud infrastructure created or acquired by the consumer. The consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but does exercise control over the deployed applications and, in some cases, the application hosting environment configuration.

[0042] Infrastructure as a Service (IaaS): The ability offered to consumers is to provision processing, storage, network, and other basic computing resources onto which they can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does exercise control over the operating systems, storage, applications deployed, and sometimes limited control over selection of network components (e.g., host firewalls).

[0043] The deployment model is as follows:

[0044] Private Cloud: Cloud infrastructure is operated solely for the organization. It can be managed by the organization or a third party and can exist on-premise or off-premise.

[0045] Community Cloud: Cloud infrastructure is shared by multiple organizations and supports a specific community with shared interests (e.g., roles, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-premises or off-premises.

[0046] Public Cloud: Cloud infrastructure is made available to the general public or large industry groups and is owned by organizations that sell cloud services.

[0047] Hybrid Cloud: A combination of two or more clouds (private, community, or public) that remain distinct entities, but are joined together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application portability.

[0048] A cloud computing environment is a service oriented environment with an emphasis on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0049] Turning now to an overview more specifically related to aspects of the present invention, hand-related activities such as cooking, sewing, creative crafts, and operating a musical instrument typically involve performing several steps or actions to advance an ongoing task or complete a desired task. For example, a song typically includes several different chord or note arrangements. To play a desired song, a user must perform several different actions to operate the instrument in a manner that achieves the correct chords or notes required to complete the task, i.e., properly play the song. However, when attempting to play a song, a user of the instrument may not realize that they are playing an incorrect chord or note. Similarly, when first learning a new song, they may not know the next chord or note progression required to advance the task correctly, i.e., continue playing the song. In the prior art, a user is required to stop the task and shift their attention from the instrument to the sheet music to determine the next chord or note progression. This action is typically performed several times before the action is memorized, increasing the time it takes to complete the task, i.e., to properly play the song.

[0050] One or more non-limiting embodiments described herein provide an AI user activity guidance system capable of performing a given task, monitoring user actions performed in real time, and learning various different user actions necessary to recommend guidance to the user on how to correctly perform one or more actions to advance or complete the task. In one or more non-limiting embodiments, the AI ​​user activity guidance system performs imaging of the user's actions while the user performs the task in real time and detects actions being performed incorrectly. A display is provided that displays images of the user performing the task actions in real time. In response to detecting the incorrectly performed actions, the AI ​​user activity guidance system generates a tactile alert alerting the user that the current action is being performed incorrectly and generates a recommendation output indicating a correction to the incorrectly performed action. The recommendation output includes spoken instructions guiding the user on how to correct the incorrectly performed action, an augmented image showing the correct action overlaid on a displayed image of the task in progress, or a combination thereof. In this manner, the user can easily correct their actions without stopping the task in progress and / or diverting their attention from the task in progress.

[0051] Referring now to FIG. 1 , an exemplary cloud computing environment 50 is illustrated. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10, with which local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, and / or a vehicle computer system 54N, may communicate. The nodes 10 may also communicate with each other. They may be physically or virtually grouped in one or more networks (not shown), such as a private, community, public, or hybrid cloud as described herein above, or a combination thereof. This enables the cloud computing environment 50 to provide infrastructure, platform, and / or software-as-a-service services that do not require cloud consumers to maintain resources on local computing devices. The types of computing devices 54A-N illustrated in FIG. 1 are intended for illustrative purposes only, and it will be understood that the computing nodes 10 and the cloud computing environment 50 may communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).

[0052] Referring now to Figure 2, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 1) is shown. It should be understood upfront that the components, layers, and functions shown in Figure 2 are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0053] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (reduced instruction set computer) architecture-based servers 62, servers 63, blade servers 64, storage devices 65, and networks and network components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0054] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities may be provided: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.

[0055] In one example, management layer 80 may provide the functions described below. Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of those resources. In one example, these resources may include application software licenses. Security provides identity authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 provides cloud computing resource allocation and management so that requested service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-configuration and procurement for cloud computing resources where future demand is predicted according to SLAs.

[0056] The workload tier 90 provides examples of functions for which a cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this tier include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and training artificial intelligence (AI) for voice response systems 96.

[0057] Turning now to a more detailed description of aspects of the present invention, Figure 3 depicts a high-level block diagram illustrating an example of a computer-based system 300 useful for implementing one or more embodiments of the present invention. While one exemplary computer system 300 is shown, computer system 300 includes a communications path 326 connecting computer system 300 to additional systems, which may include one or more wide area networks (WANs) or local area networks (LANs), such as the Internet, an intranet, or a wireless communications network, or a combination thereof. Computer system 300 and the additional systems communicate (e.g., to communicate data therebetween) via communications path 326.

[0058] Computer system 300 includes one or more processors, e.g., processor 302. Processor 302 is connected to a communications infrastructure 304 (e.g., a communications bus, crossover bar, or network). Computer system 300 may include a display interface 306 that transfers graphics, text, and other data from communications infrastructure 304 (or from a frame buffer, not shown) for display on a display unit 308. Computer system 300 may also include a main memory 310, preferably random access memory (RAM), and may also include a secondary memory 312. Secondary memory 312 may include, e.g., a hard disk drive 314 or a removable storage drive 316, representing, e.g., a floppy disk drive, a magnetic tape drive, or an optical disk drive, or a combination thereof. Removable storage drive 316 reads from and / or writes to a removable storage unit 318 in a manner well known to those skilled in the art. The removable storage unit 318 represents, for example, a floppy disk, compact disk, magnetic tape, or optical disk that is read by and written to by the removable storage drive 316. As will be appreciated, the removable storage unit 318 includes a computer-readable medium having stored therein computer software or data, or a combination thereof.

[0059] In some alternative embodiments of the present invention, secondary memory 312 may include other similar means for allowing computer programs or other instructions to be loaded into the computer system. Such means may include, for example, removable storage unit 320 and interface 322. Examples of such means may include program packages and package interfaces (e.g., those found in video game devices), removable memory chips (e.g., EPROM or PROM) and associated sockets, and other removable storage units 320 and interfaces 322 that allow software and data to be transferred from removable storage unit 320 to computer system 300.

[0060] Computer system 300 may also include a communications interface 324. Communications interface 324 allows software and data to be transferred between the computer system and external devices. Examples of communications interface 324 may include a modem, a network interface (such as an Ethernet card), a communications port, or a PCM-CIA slot and card. The software and data transferred via communications interface 324 may be in the form of signals receivable by communications interface 324, which may be, for example, electronic, electromagnetic, optical, or other signals. These signals are provided to communications interface 324 via communications path (i.e., channel) 326. Communications path 326 carries signals and may be implemented using wire or cable, optical fiber, telephone line, cellular phone link, RF link, or other communications channel, or a combination thereof.

[0061] In this disclosure, the terms "computer program medium," "computer usable medium," and "computer-readable medium" are generally used to refer to media such as main memory 310 and secondary memory 312, removable storage drive 316, and hard disks installed in hard disk drive 314. Computer programs (also called computer control logic) are stored in main memory 310 or secondary memory 312, or a combination thereof. Computer programs may also be received via communications interface 324. Such computer programs, when operating, enable the computer system to perform the features of the present disclosure described herein. Specifically, when operating, the computer programs enable processor 302 to perform the features of the computer system. Thus, such computer programs represent the controller of the computer system.

[0062] In an exemplary embodiment, a system for training artificial intelligence (AI) for a voice response system is provided. In the exemplary embodiment, the voice response system is configured to monitor ambient sounds for voice commands from a user. The voice response system interprets the voice command using AI, and based on the interpretation, the voice response system provides a response to the user and / or performs an action requested by the user. The voice response system is also configured to monitor the user's reaction to the response provided by the voice response system or the action taken by the voice response system. The reaction may be monitored using a microphone or a camera, or a combination thereof, in communication with the voice response system. The voice response system analyzes the user's reaction and determines whether the user is satisfied with the response provided or the action taken. If the user is not satisfied with the response provided or the action taken, the voice response system updates the AI ​​model used to interpret the voice command.

[0063] Referring now to FIG. 4A, an AI user activity guidance system 400 according to a non-limiting embodiment is shown. The AI ​​user activity guidance system 400 includes a computing system 12 in signal communication with functional components, including, but not limited to, one or more computing devices 430, an imaging device 432, and a voice-activated hub 420. Each of the devices, components, modules, or functions, or combinations thereof, described in FIGS. 1-3 may also apply to the devices, components, modules, and functions of FIG. 4A. In addition, one or more of the operations and stages of FIGS. 1-3 may also be included in one or more operations or actions of FIG. 4A.

[0064] Referring to Figure 4B, another non-limiting embodiment of AI user activity guidance system 400 is shown. In the non-limiting embodiment shown in Figure 4B, several components of AI user activity guidance system 400 are integrated into a single smart wearable computing device 430 (e.g., smart glasses 430). AI user activity guidance system 400 operates in a similar manner to AI user activity guidance system 400, and therefore the details of the individual components described below are also applicable to AI user activity guidance system 400.

[0065] The computing device 430 may include, but is not limited to, a television, smartphone, desktop computer, laptop computer, tablet, smartwatch, smart wearable device, or other electronic / wireless communication device, or a combination thereof, which may have one or more processors, memory, or wireless communication technology for displaying / streaming audio / video data. The computing device 430 may receive inputs (e.g., touch input, spoken input, mouse clicks, etc.) that may control the AI ​​user activity guidance system 400. In one or more non-limiting embodiments, these inputs may indicate the type of task to be performed by the user. Tasks include, for example, cooking, sewing, creative crafts, and musical instrument operation (i.e., playing an instrument). In terms of musical instrument operation, for example, the input may indicate the song to be played, the type of instrument used to play the song, or the tuning of the instrument, or a combination thereof. Additionally, the input may include the difficulty level of the task to be performed. In terms of instrument operation, for example, a beginner level of difficulty associated with playing a song through an instrument may involve correctly playing a barre chord or a power chord, while an advanced level of playing the same song may involve correctly playing the same chord as an open major / minor chord.

[0066] The imaging device 432 may include, but is not limited to, a camera or a video recorder. Thus, the imaging device 432 may capture images of the task in progress 434 (i.e., a task being performed in real time). The intelligent service 402 may work in conjunction with the imaging device 432 to perform image recognition operations. For example, the intelligent service 402 may monitor images of the task in progress 432 and detect whether an action is being performed correctly or incorrectly. The intelligent service 402 may monitor images of the task in progress 432 and predict the next action to be included in the task in progress 432.

[0067] The voice-activated hub 420 may include, for example, a personal assistant Internet of Things (IoT) computing device, and may detect voice-activated commands or queries, or a combination thereof, and output spoken language including answers, prompts, recommendations, etc.

[0068] Computer system / server 12 is again shown, which may incorporate intelligent services 402 or "intelligent recommendation of guidance instruction services 402" (e.g., artificial intelligence simulated humanoid assistant "AISHA"). As shown in FIG. 4A, computer system / server 12 may provide virtualized computing services (i.e., virtualized computing, virtualized storage, virtualized networking, etc.) to one or more computing devices described herein. More specifically, computer system / server 12 may provide virtualized computing, virtualized storage, virtualized networking, and other virtualized services running on a hardware board.

[0069] 4A (e.g., intelligent recommendation 402 of the guidance instruction service) is in communication with and / or associated with computing device 430, imaging device 432, and voice-activated hub 420. Thus, intelligent recommendation 402 of the guidance instruction service, computing device 430, imaging device 432, and guidance voice-activated hub 420 may each be associated with and / or communicate with each other via one or more communication methods, such as a computing network, a wireless communication network, or other network means that enable communication (each collectively referred to as “network 18” in FIG. 4A ). According to a non-limiting embodiment, intelligent recommendation 402 of the guidance instruction service may be installed on voice-activated hub 420 or computing device 430, or a combination thereof. Alternatively, intelligent recommendation 402 of the guidance instruction service may be located external to voice-activated hub 420 or computing device 430, or a combination thereof (e.g., via a cloud computing server).

[0070] The intelligent recommendation of the guidance service 402 may incorporate a processing unit 16 that performs various computing data processing and other functions according to various non-limiting embodiments of the present invention. Domain knowledge 412 (e.g., a database that may include an ontology) is shown along with a guidance component 404, an analysis component 406, a monitoring component 408, a machine learning component 410, a recognition component 414, or an augmented reality (AR) component 416, or a combination thereof. In one or more non-limiting embodiments, one of the guidance component 404, the analysis component 406, the monitoring component 408, the machine learning component 410, the recognition component 414, the domain knowledge 412, or the augmented reality (AR) component 416, or a combination thereof, is implemented as an electronic hardware controller including a memory and a processor configured to execute algorithms and computer-readable program instructions stored in the memory. Additionally, the guidance component 404, the analysis component 406, the monitoring component 408, the machine learning component 410, the recognition component 414, the domain knowledge 412, or the augmented reality (AR) component 416, or any combination thereof, may all be incorporated or integrated into a single controller.

[0071] Domain knowledge 412 may include and / or be associated with an ontology of concepts, keywords, expressions that represent a domain of knowledge. The thesaurus or ontology may be used as a database and may be used by the machine learning component 410 (e.g., a recognition component) to identify semantic relationships between observed and / or unobserved variables. According to non-limiting embodiments, the term "domain" is intended to have its ordinary meaning. Additionally, the term "domain" may include a discipline, a system or collection of materials, information, content, or other resources related to a particular subject or subjects, or a combination thereof. A domain may refer to information related to any particular subject or combination of selected subjects.

[0072] The term ontology is also intended to have its ordinary meaning. According to non-limiting embodiments, the term ontology, in its broadest sense, may include anything that can be modeled as an ontology, including, but not limited to, taxonomies, thesauri, vocabularies, etc. For example, an ontology may include information or content related to a domain of interest or content of a particular class or concept. An ontology may be continually updated with information synchronized with a source, adding information from the source to the ontology as models, attributes of the models, or associations between models within the ontology. In one or more non-limiting embodiments, the domain knowledge stores learned model behaviors that include images showing exemplary or correct performance of behaviors included in a given task. In terms of musical instruments, for example, the learned model behaviors may include images showing how to play the correct chords or notes on a given instrument.

[0073] Additionally, domain knowledge 412 may include one or more external resources, such as, for example, links to one or more internet domains, web pages, etc. For example, textual data may be hyperlinked to web pages that may describe, explain, or provide additional information related to the textual data. Thus, summaries may be enhanced through links to external resources that further explain, instruct, illustrate, or provide context and / or additional information to support decisions, alternative suggestions, alternative options, or criteria, or a combination thereof.

[0074] The analysis component 406 of the computer system / server 12 may cooperate with the processing unit 16 to accomplish various embodiments of the present invention. For example, the analysis component 406 may undergo various data analysis functions to analyze data communicated from one or more devices, such as, for example, the voice-activated hub 420 or the computing device 430, or a combination thereof.

[0075] The analysis component 406 may receive and analyze each physical characteristic associated with the media data (e.g., audio data or video data, or a combination thereof). The analysis component 406 may cognitively receive, detect, or both, the audio data or video data, or a combination thereof, for the guidance instructions component 404.

[0076] The analysis component 406, the monitoring component 408, or the machine learning component 410, or a combination thereof, may access and monitor one or more audio or video data sources or a combination thereof (e.g., websites, audio storage systems, video storage systems, cloud computing systems, etc.) to provide audio, video, or text data or a combination thereof for providing guiding instructions for performing a task. The analysis component 406 may cognitively analyze data obtained from domain knowledge 412, one or more online sources, cloud computing systems, text corpora, or a combination thereof. The analysis component 406 or the machine learning component 410, or a combination thereof, may use natural language processing (“NLP”) to extract one or more keywords, phrases, instructions, or transcripts (e.g., transcribe audio data into text data), or a combination thereof.

[0077] As part of discovering data, the analytics component 406, the monitoring component 408, and / or the machine learning component 410 can identify audio data, video data, text data, and / or contextual factors associated with the audio data, video data, and / or text data, or a combination thereof, from one or more sources. The machine learning component 410 can also initiate machine learning operations to learn contextual factors associated with the audio data, video data, or text data, or a combination thereof, that are associated with guidance instructions for performing a task, such as, for example, assembling or repairing an item, or a combination thereof (e.g., assembling a new bicycle or repairing a computer).

[0078] The monitoring component 408 can monitor the ongoing performance of the task 434 via the imaging device 432 and one or more user voice / words 423 (via the microphone 433 or the voice hub 420, or a combination thereof) in real time while the task 434 is being performed. The recognition component 414 can recognize that a user is performing the task 434 on an item using the voice-activated hub 420 or the computing device 430, or a combination thereof. For example, the voice-activated hub 420, the computing device 430, or the imaging device 432, or a combination thereof, can identify one or more activities, body movements or characteristics, or a combination thereof (e.g., facial recognition, facial expressions, hand / foot gestures, etc.), behavior, audio data (e.g., voice detection or recognition, or a combination thereof), surrounding environment, or other defined parameters / features that may identify, locate, and / or recognize the user and / or the task being performed by the user. The analysis component 406, working with the monitoring component, can perform image recognition to recognize each object in a given image or video clip among the objects in each frame, respectively.

[0079] In one or more non-limiting embodiments, the monitoring component 408 can detect an erroneously performed action included in the task in progress 434. In response to detecting the erroneous action or its next action, or a combination thereof, the monitoring component can instruct one or more computing devices 430 to generate a haptic alert 436 that alerts the user that the current action is being performed erroneously. Similarly, the monitoring component 408 can predict the next action to be performed in the task in progress 434 and generate a haptic alert 436 that alerts the user of the next action to be performed.

[0080] The navigation prompts component 404 may provide one or more navigation prompts to assist in the execution of a selected task in response to the identified contextual factors. The navigation prompts may be textual, audio, or video data, or a combination thereof. For example, the voice-activated hub 420 may audibly communicate the navigation prompts 422. The computing device 430 may provide the navigation prompts 450 along with image / video data 485 displayed by a graphical user interface ("GUI") of the computing device 430, or as spoken instructions output from a speaker 431, or a combination thereof.

[0081] The guiding instructions component 404 may cognitively guide a user to perform a selected task using one or more guiding instructions 422. The guiding instructions component 404 may provide a sequence of guiding instructions 422 obtained from domain knowledge, one or more online sources, a cloud computing system, a text corpus, or a combination thereof. The guiding instructions component 404 may provide media data from one or more online sources, a cloud computing system, or a combination thereof.

[0082] The guidance instructions component 404 may validate each stage of the one or more guidance instructions 422 to assist in performing the selected task. The guidance instructions component 404 may also identify a level of difficulty (e.g., level of stress, frustration, anxiety, excitement, or other emotional response) achieved by the user while performing a set of tasks associated with one or more guidance instructions 422 delivered via streaming media, and may pause / stop / unsubscribe from the streaming media for a selected period of time, or provide the user with a modified set of guidance instructions 422 to guide the user through the increased instruction level, or a combination thereof.

[0083] The guidance instructions component 404 may provide additional guidance information related to guidance instructions 422 collected from domain knowledge, one or more online sources, cloud computing systems, text corpora, or combinations thereof, for performing a selected task. For example, if an initial set of instructions is insufficient for the user, an additional set may be provided that may further explain one or more of the original instructions.

[0084] According to one or more non-limiting embodiments, the guidance instruction component 404 may also work in conjunction with the AR component 416 to provide enhanced guidance instructions 422 for assistance in performing a selected task. More specifically, the AR component 416 may generate an augmented image 500 overlaid on real-time image / video data 485 displayed in a graphical user interface (GUI) 429. The augmented image 500 is thus displayed along with the task in progress 434 such that the augmented image 500 shows the user how to correct an incorrectly performed action. In this manner, the user can easily correct their action without stopping the task in progress 434 and / or diverting their attention from the task in progress. In another non-limiting embodiment, the augmented image 500 may show the user how to correctly perform the next action included in the task in progress 434.

[0085] The intelligent recommendation of the guidance instructions service 402 may adjust the tone, volume, pace, or frequency of the voice of the audio / media data of the guidance instructions 422, or a combination thereof, based on the user's speed / pace of following the guidance instructions 422. Also, words, phrases, or complete sentences (e.g., all or part of a conversation) by others associated with the audio data, or a combination thereof, may be transcribed into text form based on NLP extraction operations (e.g., NLP-based keyword extraction). The text data may be transferred, transmitted, stored, or further processed so that the same audio / video data (e.g., all or part of a conversation) can be heard or listened to simultaneously provide a text version of the guidance instructions.

[0086] As previously indicated, the intelligent recommendations of the guided prompt service 402 can also communicate with other linked devices, such as, for example, the voice-activated hub 420, the computing device 430, or the imaging device 432, or a combination thereof. Additionally, the analytics component 406 or the machine learning component 410, or a combination thereof, can access one or more online data sources, such as, for example, social media networks, websites, or data sites, to provide one or more guided prompts 422 to assist in the performance of a selected task in response to identified contextual factors. That is, the analytics component 406, the recognition component 414, or the machine learning component 410, or a combination thereof, can learn and observe a user's degree or level of attention, the user's level of difficulty in performing a task, the type of response, and / or feedback on various topics and / or guided prompts 422. The user's learning and observed behavior can be linked to various data sources providing personal information, social media data, or user profile information for learning, probabilities, or determining, or a combination thereof, the confidence associated with the performance of the guided prompts 422.

[0087] In one or more non-limiting embodiments, observations of other people's (users') responses to the same method can be used by the AI ​​user activity guidance system 400 to understand learning and recommend actions to the user. Iterative feedback between crowdsourced data and the focused user can then be utilized to make intelligent decisions about the usage path and projections for a given task. Additionally, the AI ​​user activity guidance system 400 can actively modify the predicted path for an ongoing task based on the learned success rate.

[0088] According to non-limiting embodiments, the machine learning component 410 may be implemented by a wide variety of methods or combinations of methods, such as supervised learning, unsupervised learning, time-lag learning, reinforcement learning, etc. Some non-limiting examples of supervised learning that may be used with the present technology include AODE (averaged averaged univariate estimation of dependence), artificial neural networks, backpropagation, Bayesian statistics, naive Bayes classifiers, Bayesian networks, Bayesian knowledge bases, case-based reasoning, decision trees, inductive logic programming, Gaussian process regression, gene expression programming, group methods of data manipulation (GMDH), learning automata, learning vector quantization, minimum message length (decision trees, decision graphs, etc.), lazy learning, learning by example, nearest neighbor methods, analogical modeling, probabilistic learning, etc. Approximately correct (PAC) learning, ripple down rules, knowledge acquisition methodologies, symbolic machine learning algorithms, quasi-symbolic machine learning algorithms, support vector machines, random forests, classifier ensembles, bootstrap aggregating (bagging), boosting (meta-algorithms), regular classification, regression analysis, information fuzzy networks (IFNs), statistical classification, linear classifiers, Fisher linear discriminant, logistic regression, perceptrons, support vector machines, quadratic classifiers, k-nearest neighbors, hidden Markov models, and boosting. Some non-limiting examples of unsupervised learning that may be used with the present technology include artificial neural networks, data clustering, expectation maximization, self-organizing maps, radial basis function networks, vector quantization, generative topographic maps, information bottleneck methods, IBSEAD (Distributed Autonomous Entity System Based Dialogue), association rule learning, apriori algorithm, eclat algorithm, FP-growing algorithm, hierarchical clustering, single-link clustering, concept clustering, divisive clustering, k-means algorithm, fuzzy clustering, and reinforcement learning. Some non-limiting examples of lagged learning may include Q-learning and learning automata. Specific details regarding any of the supervised, unsupervised, lagged, or other machine learning examples described in this paragraph are known and within the scope of this disclosure.Also, when deploying one or more machine learning models, the computing device may first be tested in a controlled environment before being deployed in a public environment, and even when deployed in a public environment (e.g., outside of a controlled testing environment), the computing device may be monitored for compliance.

[0089] According to non-limiting embodiments, the intelligent recommendations of the guidance instruction service 402 may perform one or more calculations according to a mathematical operation or function, which may include one or more mathematical operations (e.g., analytically or computationally solving differential or partial differential equations using addition, subtraction, multiplication, division, standard deviation, mean, average, probability, probabilistic modeling using statistical distributions, by finding minimum, maximum, or similar thresholds such as combined variables). Thus, as used herein, a calculation operation may include all or a portion of one or more mathematical operations.

[0090] According to a non-limiting embodiment, if the task the user wishes to perform is initially undetectable, the user may provide activity data as input to the intelligent recommendation of the guided instruction service 402 (e.g., verbally via the voice-activated hub 420 and / or microphone 433 and / or via the GUI 429 interface of the computing device 430) so that the intelligent recommendation of the guided instruction service 402 may begin object scanning, instruction scanning (after downloading to the corpus if not already performed), and guiding the user with step-by-step instructions based on monitoring the user's activity.

[0091] As described herein, the AI ​​user activity guidance system 400 includes an AR component 416 configured to generate augmented images 500 overlaid on real-time image / video data to provide enhanced guidance instructions 422 to assist the user in performing a task. Figures 5, 6, and 7 collectively illustrate an example task as if the user were playing the guitar.

[0092] Referring initially to FIG. 5, a learned modeled user activity is illustrated in accordance with one or more embodiments of the present invention. The learned modeled user activity 600 in this example is a learned correctly played guitar barre chord 600 (e.g., a barre C chord 600). The correctly played guitar barre chord 600 may be learned by the machine learning component 410 described herein and stored in the domain knowledge 412 for future reference by the AI ​​user activity guidance system 400 (e.g., the AR component 416). A chord diagram 602 corresponding to the correctly played guitar barre chord 600 may also be stored in the domain knowledge 412 for future reference by the AI ​​user activity guidance system 400.

[0093] 6 illustrates a GUI 429 displaying image / video data 485 of an ongoing task 434 (e.g., a user playing a desired song on a guitar) performed in real time, where the user is incorrectly performing a user activity included in task 434. In this example, the incorrect user activity is an incorrectly performed barre chord on the guitar (e.g., an incorrect barre C chord). In one or more non-limiting embodiments, the GUI 429 may also display a chord diagram 602 including an indicator 604 indicating the actual guitar string being played by the user. In one or more embodiments, the display of the indicator (e.g., color, shape, etc.) may be altered to indicate which particular guitar string is being incorrectly played.

[0094] 7 illustrates the GUI 429 displaying an augmented image 500 overlaid on top of real-time image / video data 485. As described herein, the augmented image 500 shows the user how to correct an incorrectly performed action. In this example, the augmented image 500 shows the user how to correctly play a guitar barre chord (e.g., a correct barre C chord) to correctly progress through or complete a task, e.g., to correctly play a song. Thus, the user can easily correct their action without stopping and / or diverting their attention from the guitar. In one or more non-limiting embodiments, the GUI 429 may also display a chord diagram 602 including correction indicators 606 indicating which guitar strings should be played to correctly play the song.

[0095] 8, 9, and 10 collectively illustrate an example task as if the user were playing the guitar, according to another non-limiting embodiment. As mentioned herein, the user can input a difficulty level corresponding to the selected task to be performed. For example, FIGS. 5, 6, and 7 above may correspond to a user inputting a request to play a particular song corresponding to a beginner level. Thus, the AI ​​user activity guidance system 400 can obtain a learned modeled image of barre chords from the domain knowledge 412.

[0096] However, if the user inputs a request to play a song at a more advanced level, the AI ​​user activity guidance system 400 can retrieve from the domain knowledge 412 (see FIG. 8 ) a learned modeled image 600 of an open major / minor chord (e.g., an open guitar C chord), which may be more complex to play compared to a barre chord.

[0097] 9, GUI 429 displays image / video data 485 of an ongoing task 434 (e.g., a user playing a desired song on a guitar) performed in real time, where the user is incorrectly playing an open C chord. Chord diagram 602 includes an indicator 604 that indicates the user is holding down the wrong string.

[0098] 10, GUI 429 displays an augmented image 500 overlaid on top of real-time image / video data 485. As described herein, augmented image 500 shows the user how to correctly play an open chord (e.g., a correct open C chord) to successfully progress through or complete a task at a more advanced difficulty level. GUI 429 also displays a chord diagram 602 that includes correction indicators 606 that indicate how to modify an action, e.g., how to correctly play an open C chord.

[0099] As described herein, the AI ​​user activity guidance system 400 can monitor images of the task in progress 432 and predict the next action involved in the task in progress 432. Accordingly, an augmented image 500 can be generated that notifies the user of the next action to be performed in the task in progress 432.

[0100] 11 and 12 collectively illustrate an AI user activity guidance system 400 that predicts the next guitar chord to be played in an ongoing song being played by a user. In FIG. 11 , for example, GUI 429 displays image / video data 485 of a user playing an open A chord included in the song the user is playing in real time. As the song progresses, AI user activity guidance system 400 recognizes that the next chord in the song is an open C chord. Therefore, AI user activity guidance system 400 actively generates an augmented image 500, as shown in FIG. 12 . Augmented image 500 is overlaid on image / video data 485 to inform or instruct the user on how to transition from a current action (e.g., an open A chord) to the next action in the task (e.g., an open C chord). In this way, the user can continue to accurately play the song without diverting attention from the ongoing task.

[0101] Referring now to FIG. 13 , a method for recommending guidance instructions to a user by the AI ​​user activity guidance system 400 is shown, according to one or more embodiments of the present invention. The method begins at operation 800, where, at operation 802, the AI ​​user activity guidance system 400 determines a task to be performed by the user. The task may include multiple user actions performed by the user and may be determined in response to receiving user input (e.g., touch input, voice input, etc.) indicating the task. At operation 804, the AI ​​user activity guidance system 400 obtains one or more learned model actions included in the task. The learned model actions may be obtained from domain knowledge 412. In one or more non-limiting embodiments, the learned model actions include images showing exemplary or correct performance of the actions. At operation 806, the AI ​​user activity guidance system 400 generates image data of the ongoing task being performed by the user in real time. The image data may include, for example, a video stream generated by a camera monitoring the ongoing task.

[0102] Moving to operation 808, the AI ​​user activity guidance system 400 analyzes the image data of the ongoing task and determines, in operation 810, whether the current action included in the ongoing task is being performed correctly. If the action is being performed correctly, the AI ​​user activity guidance system 400 determines whether the task is complete, i.e., whether all actions included in the task have been performed. If the task is complete, the method ends. Otherwise, the AI ​​user activity guidance system 400 proceeds to operation 824 to determine the next action included in the task, which is described in more detail below.

[0103] However, if the behavior is performed incorrectly, the AI ​​user activity guidance system 400 generates a haptic warning indicating that the current behavior is being performed incorrectly at operation 812. At operation 814, the AI ​​user activity guidance system 400 accesses domain knowledge 412 to obtain a learned modeled image of the correct behavior, which is used to augment the image data. At operation 816, the AI ​​user activity guidance system 400 augments the image data by overlaying the learned modeled image on top of the image data. Thus, a user viewing the GUI 429 can become aware of how to correct the current incorrectly performed behavior. At operation 818, the AI ​​user activity guidance system 400 can analyze the image data to determine whether the user has adjusted their performance to correct their behavior based on the augmented image. If the incorrect behavior has not been corrected, the method returns to operation 816 and continues overlaying the learned modeled image on top of the image data until the user corrects the incorrect behavior. Once the incorrect behavior is corrected, the method proceeds to operation 820 to determine whether the task is complete. Once the task is completed, the method ends at operation 822.

[0104] If the task is not completed, the AI ​​user activity guidance system 400 determines the next action included in the task at operation 824. At operation 826, the AI ​​user activity guidance system 400 generates a haptic alert notifying the user that the next action in the task should be performed. At operation 828, the AI ​​user activity guidance system 400 accesses domain knowledge 412 to obtain a learned modeled image of the next action included in the task, which is used to augment the image data. At operation 830, the AI ​​user activity guidance system 400 augments the image data by overlaying the learned modeled image of the next action on top of the image data. Thus, a user viewing the GUI 429 can quickly move on to the next action included in the task without diverting attention from the ongoing task. The method returns to operation 810, where the AI ​​user activity guidance system 400 analyzes whether the next action was performed correctly, and the method continues as described above.

[0105] Further details of machine learning techniques that may be used to implement portions of computer system / server 12 are now provided. Various types of computer control functions described herein (e.g., inferences, judgments, decisions, recommendations, etc. of computer system / server 12) may be implemented using machine learning or natural language processing techniques, or a combination thereof. Generally, machine learning techniques operate on so-called "neural networks," which may be implemented as programmable computers configured to run a set of machine learning algorithms. Neural networks incorporate knowledge from diverse disciplines, including neurophysiology, cognitive science / psychology, physics (statistical mechanics), control theory, computer science, artificial intelligence, statistics / mathematics, pattern recognition, computer vision, parallel processing, and hardware (e.g., digital / analog / VLSI / optics).

[0106] The basic function of neural networks and their machine learning algorithms is to recognize patterns by interpreting unstructured sensor data through some kind of machine learning algorithm. Unstructured real-world data in its natural form (e.g., images, sound, text, or time series data) is converted into a numerical form (e.g., vectors with magnitude and direction) that can be understood and manipulated by a computer. The machine learning algorithm performs multiple iterations of learning-based analysis on real-world data vectors until patterns (or relationships) contained in the real-world data vectors are discovered and learned. The learned patterns / relationships serve as predictive models that can be used to perform a variety of tasks, including, for example, real-world data classification (or labeling) and real-world data clustering. Classification tasks often rely on the use of labeled datasets to train neural networks (i.e., models) to recognize correlations between labels and data. This is known as supervised learning. Examples of classification tasks include detecting people / faces in images, recognizing facial expressions in images (e.g., angry, happy, etc.), identifying objects in images (e.g., stop signs, pedestrians, lane markers, etc.), recognizing gestures in video, detecting musical instruments and instrument manipulations, detecting hand activities (e.g., cooking, cross-stitching, sewing, etc.), detecting voices, detecting voices in audio, identifying specific speakers, transcribing speech to text, etc. Clustering tasks identify similarities between objects, allowing objects to be grouped according to common characteristics and differentiated from other groups of objects. These groups are known as "clusters."

[0107] An example of a machine learning technique that may be used to implement aspects of the present invention is described with reference to Figures 14 and 15. A machine learning model configured and arranged in accordance with an embodiment of the present invention is described with reference to Figure 14. A detailed description of an example of a computing system and network architecture capable of implementing one or more of the embodiments of the present invention described herein is provided with reference to Figure 16.

[0108] FIG. 14 shows a block diagram illustrating a classifier system 1200 capable of implementing various aspects of the invention described herein. More specifically, the functionality of system 1200 is used in embodiments of the invention to generate various models and sub-models that can be used to implement the computer functionality of embodiments of the invention. System 1200 includes multiple data sources 1202 in communication with classifier 1210 over network 1204. In some aspects of the invention, data sources 1202 may bypass network 1204 and feed directly to classifier 1210. Data sources 1202 provide data / information inputs that are evaluated by classifier 1210 according to embodiments of the invention. Data sources 1202 also provide data / information inputs that can be used by classifier 1210 to train and / or update model 1216 created by classifier 1210. Data sources 1202 can be implemented as a wide variety of data sources, including, but not limited to, sensors configured to collect real-time data, data repositories (including training data repositories), cameras, and outputs from other classifiers. The network 1204 can be any type of communication network, including, but not limited to, a local network, a wide area network, a private network, the Internet, and the like.

[0109] The classifier 1210 may be implemented as an algorithm executed by a programmable computer, such as the processing system 1400 (shown in FIG. 16 ). As shown in FIG. 14 , the classifier 1210 includes a set of machine learning (ML) algorithms 1212, a natural language processing (NLP) algorithm 1214, and a model 1216, which is a relationship (or predictive) algorithm generated (or learned) by the ML algorithm 1212. The algorithms 1212, 1214, and 1216 of the classifier 1210 are shown separately for ease of illustration and explanation. In embodiments of the invention, the functions performed by the various algorithms 1212, 1214, and 1216 of the classifier 1210 may be distributed differently than shown. For example, if the classifier 1210 is configured to perform an entire task having subtasks, the set of ML algorithms 1212 may be segmented so that some of the ML algorithms 1212 perform each subtask and some of the ML algorithms 1212 perform the entire task. Furthermore, in some embodiments of the present invention, NLP algorithms 1214 may be integrated within ML algorithms 1212.

[0110] The NLP algorithms 1214 include speech recognition functionality that enables the classifier 1210, and more specifically the ML algorithms 1212, to receive natural language data (text and audio) and apply elements of language processing, information retrieval, and machine learning to derive meaning from the natural language input and potentially take action based on the derived meaning. The NLP algorithms 1214 used in accordance with embodiments of the present invention may also include speech synthesis functionality that enables the classifier 1210 to translate results 1220 into natural language (text and audio) and communicate aspects of the results 1220 as natural language communications.

[0111] The NLP and ML algorithms 1214, 1212 receive and evaluate input data (i.e., training data and data to be analyzed) from the data sources 1202. The ML algorithm 1212 includes the functionality necessary to interpret and utilize the format of the input data. For example, if the data sources 1202 include image data, the ML algorithm 1212 may include vision software configured to interpret the image data. The ML algorithm 1212 applies machine learning techniques to the received training data (e.g., data received from one or more of the data sources 1202, images or sounds extracted from a video stream, or a combination thereof) to create / train / update one or more models 1216 over time that model the overall tasks and subtasks that the classifier 1210 is designed to complete.

[0112] 14 and 15 , FIG. 15 illustrates an example of a learning phase 1300 performed by the ML algorithm 1212 to generate the model 1216 described above. In the learning phase 1300, the classifier 1210 extracts features from the training data and converts the features into a vector representation that can be recognized and analyzed by the ML algorithm 1212. The feature vectors are analyzed by the ML algorithm 1212 to “classify” the training data relative to a target model (e.g., a correct “model” of the behavior involved in the task or model’s task) and discover relationships between and across the classified training data. Examples of suitable implementations of the ML algorithm 1212 include, but are not limited to, neural networks, support vector machines (SVMs), logistic regression, decision trees, hidden Markov models (HMMs), etc. The learning or training performed by the ML algorithm 1212 may be supervised, unsupervised, or a hybrid that includes aspects of supervised and unsupervised learning. Supervised learning is when training data is already available and classified / labeled. Unsupervised learning is when training data is not classified / labeled and must be developed through iterations of the classifier 1210 and ML algorithm 1212. Unsupervised learning can utilize additional learning / training methods including, for example, clustering, anomaly detection, neural networks, deep learning, etc.

[0113] When the model 1216 has been sufficiently trained by the ML algorithm 1212, a data source 1202 that generates "real-world" data is accessed, and the "real-world" data is applied to the model 1216 to generate a usable version of the results 1220. In some embodiments of the invention, the results 1220 may be returned to the classifier 1210 and used by the ML algorithm 1212 as additional training data to update and / or refine the model 1216.

[0114] In an aspect of the present invention, the ML algorithm 1212 and the model 1216 can be configured to apply a confidence level (CL) to various ones of their respective results / judgments (including result 1220) in order to improve the overall accuracy of a particular result / judgment. If the ML algorithm 1212 or the model 1216 or both make or generate a judgment for a result where the value of the CL is below a predetermined threshold (TH) (i.e., CL < TH), the result / judgment is classified as having a sufficiently low "confidence" and can justify the conclusion that the judgment / result is not valid, and this conclusion can be used to determine when, how, and / or whether the judgment / result should be handled in downstream processing. If CL > TH, the judgment / result can be considered valid, and this conclusion can be used to determine when, how, and / or whether the judgment / result should be handled in downstream processing. Various pre-determined TH levels can be provided. The judgment / result with CL > TH can be ranked from the highest CL > TH to the lowest CL > TH in order to prioritize when, how, and / or whether the judgment / result should be handled in downstream processing.

[0115] In an aspect of the present invention, the classifier 1210 can be configured to apply a confidence level (CL) to the result 1220. When the classifier 1210 determines that the CL in the result 1220 is below a predetermined threshold (TH) (i.e., CL < TH), the result 1220 is classified as being sufficiently low and can justify the classification of "no confidence" in the result 1220. If CL > TH, the result 1220 is classified as being sufficiently high and can justify the determination that the result 1220 is valid. Various pre-determined TH levels can be provided such that the results 1220 with CL > TH can be ranked from the highest CL > TH to the lowest CL > TH.

[0116] The classifier 1210, and more specifically the functions performed by the ML algorithm 1212, may be organized as a weighted directed graph, where the nodes are artificial neurons (e.g., modeled after neurons in the human brain) and weighted directed edges connect the nodes. The directed graph of the classifier 1210 may be organized so that certain nodes form input layer nodes, certain nodes form hidden layer nodes, and certain nodes form output layer nodes. The input layer nodes connect to hidden layer nodes, and the hidden layer nodes connect to output layer nodes. Each node is connected to all nodes in adjacent layers by connection paths, which may be shown as directional arrows, each with a connection strength. Multiple input layers, multiple hidden layers, and multiple output layers may be provided. When multiple hidden layers are provided, the classifier 1210 may perform unsupervised deep learning to perform its assigned tasks.

[0117] Similar to the functioning of the human brain, each input layer node receives inputs without connection strength adjustment and node summation. Each hidden layer node receives its inputs from all input layer nodes according to the connection strengths associated with the associated connection paths. Similar connection strength multiplication and node summation are performed for hidden and output layer nodes.

[0118] The weighted directed graph of the classifier 1210 "learns" by processing data records (e.g., output from the data source 1202) one by one and comparing the record's initial, arbitrary classification with the record's known, actual classification. Using a training methodology known as "backpropagation" (i.e., "backpropagation of error"), errors from the initial classification of the first record are fed back into the weighted directed graph of the classifier 1210 and used to modify the weighted connections of the weighted directed graph a second time, and this feedback process continues iteratively. During the training phase of the weighted directed graph of the classifier 1210, the correct classification of each record is known, and therefore, the output nodes can be assigned "correct" values. For example, nodes corresponding to the correct class are assigned a node value of "1" (or 0.9), and others are assigned a node value of "0" (or 0.1). In this way, the calculated values ​​of the weighted directed graph for the output nodes can be compared to these "correct" values ​​and an error term for each node can be calculated (i.e., the "delta" rule). These error terms are then used to adjust the weights in the hidden layer so that the output values ​​in the next iteration are closer to the "correct" values.

[0119] 16 illustrates a high-level block diagram of a computer system 1400 that may be used to implement one or more computer processing operations in accordance with aspects of the present invention. While one exemplary computer system 1400 is shown, the computer system 1400 includes a communication path 1425 connecting the computer system 1400 to additional systems (not shown), which may include one or more wide area networks (WANs) or local area networks (LANs), such as the Internet, an intranet, or a wireless communication network, or a combination thereof. The computer system 1400 and the additional systems communicate via the communication path 1425, for example, to communicate data therebetween. In some embodiments of the present invention, the additional systems may be implemented as one or more cloud computing systems 50. The cloud computing systems 50 may supplement, support, or replace some or all (in any combination) of the functionality of the computer system 1400, including any and all computing systems described in this detailed description that may be implemented using the computer system 1400. Furthermore, some or all of the functionality of the various computing systems described in this detailed description may be implemented as nodes of the cloud computing system 50.

[0120] Computer system 1400 includes one or more processors, such as processor 1402. Processor 1402 is connected to a communications infrastructure 1404 (e.g., a communications bus, crossover bar, or network). Computer system 1400 may include a display interface 1406 that transfers graphics, text, and other data from communications infrastructure 1404 (or from a frame buffer, not shown) for display on a display unit 1408. Computer system 1400 may also include main memory 1410, preferably random access memory (RAM), and may also include secondary memory 1412. Secondary memory 1412 may include, for example, a hard disk drive 1414 or a removable storage drive 1416, or a combination thereof, representing, for example, a floppy disk drive, magnetic tape drive, or optical disk drive. Removable storage drive 1416 reads from and / or writes to a removable storage unit 1418 in a manner well known to those skilled in the art. The removable storage unit 1418 represents, for example, a floppy disk, compact disk, magnetic tape, or optical disk, flash drive, solid-state memory, etc. that is read by and written to by the removable storage drive 1416. As will be appreciated, the removable storage unit 1418 includes a computer-readable medium having stored therein computer software or data, or a combination thereof.

[0121] In alternative embodiments of the present invention, secondary memory 1412 may include other similar means for allowing computer programs or other instructions to be loaded into the computer system. Such means may include, for example, removable storage unit 1420 and interface 1422. Examples of such means may include program packages and package interfaces (e.g., those found in video game devices), removable memory chips (e.g., EPROM or PROM) and associated sockets, and other removable storage units 1420 and interfaces 1422 that allow software and data to be transferred from removable storage unit 1420 to computer system 1400.

[0122] Computer system 1400 may also include a communications interface 1424. The communications interface 1424 allows software and data to be transferred between the computer system and external devices. Examples of communications interface 1424 may include a modem, a network interface (such as an Ethernet card), a communications port, or a PCM-CIA slot and card. The software and data transferred via communications interface 1424 may be in the form of signals receivable by communications interface 1424, which may be, for example, electronic, electromagnetic, optical, or other signals. These signals are provided to communications interface 1424 over communications path (i.e., channel) 1425. Communications path 1425 carries signals and may be implemented using wire or cable, optical fiber, telephone line, cellular phone link, RF link, or other communications channel, or a combination thereof.

[0123] Technical advantages include improved capabilities of an AI computing system that can learn a variety of different user actions required to perform a given task, monitor the user's actions as they are performed in real time to accomplish the desired task, and recommend guidance to the user on how to correctly perform one or more of the actions to accomplish the task. In one or more non-limiting embodiments, the AI ​​user activity guidance system performs imaging of the user's actions as the user performs the task in real time and detects actions that are being performed incorrectly. A display is provided that displays images of the user performing the task actions in real time. In response to detecting the incorrectly performed action, the AI ​​user activity guidance system generates a tactile alert alerting the user that the current action is being performed incorrectly and generates a recommendation output indicating a correction to the incorrectly performed action. The recommendation output may include spoken instructions guiding the user on how to correct the incorrectly performed action, an augmented image overlaid on an image shown on the display indicating the correct action, or a combination thereof. In this manner, the user can easily correct their actions without stopping and / or diverting their attention from the task in progress. Thus, the AI ​​user activity guidance system described herein facilitates users to complete tasks more quickly while avoiding errors in the completed tasks.

[0124] The present invention may be a system, a method, or a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon that cause a processor to perform aspects of the present invention.

[0125] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. An exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. Computer-readable storage medium, as used herein, should not be construed as being a transitory signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses through a fiber optic cable), or an electrical signal transmitted through a wire.

[0126] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.

[0127] The computer-readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, C++, or the like, or conventional procedural programming languages ​​such as the “C” programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry to perform aspects of the present invention.

[0128] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0129] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, whereby the instructions executing via the processor of the computer or other programmable data processing apparatus form means for implementing the function(s) / act(s) specified in the flowchart and / or block diagram block(s). These computer-readable program instructions may also be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium having stored thereon instructions comprises an article of manufacture containing instructions that implement an aspect of the function(s) / act(s) specified in the flowchart and / or block diagram block(s).

[0130] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device and cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device to generate a computer-implemented process, whereby the instructions executing on the computer, other programmable apparatus, or other device implement the functions / acts identified in the flowchart and / or block diagram blocks or within the blocks.

[0131] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially in parallel, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or actions or executes a combination of dedicated hardware and computer instructions.

[0132] The description of various embodiments of the present invention is presented for illustrative purposes and is not intended to be exhaustive or to be limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications of or technical improvements to the technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. a memory having computer readable instructions; and one or more processors for executing the computer-readable instructions, the computer-readable instructions comprising: receiving input from a user indicating a task including at least one action to be performed by the user; monitoring an ongoing task including the at least one action performed by the user in response to receiving the input; generating image data indicative of the task in progress; displaying the ongoing task based on the image data; and analyzing the ongoing task; and generating an augmented image; overlaying the augmented image on the image data such that the augmented image is displayed simultaneously with the task in progress to instruct the user to proceed with the task in progress; controlling the one or more processors to perform operations including: system.

2. The system of claim 1 , wherein the augmented image is generated in response to detecting that the at least one action is being performed erroneously.

3. The system of claim 2 , wherein the augmented image is a modified image showing a modification of the at least one movement.

4. The system of claim 3 , further comprising generating a tactile alert in response to detecting that the at least one action is being performed erroneously.

5. The system of claim 1 , wherein the augmented image is generated in response to determining a next action from the at least one action included in the task.

6. The system of claim 5 , wherein the augmented image is an image showing the next action.

7. The system of claim 6 , further comprising generating a haptic alert in response to determining the next action.

8. 1. A method for recommending guidance instructions to a user, comprising: receiving input from a user indicating a task including at least one action to be performed by the user; responsive to receiving the input, monitoring an ongoing task including the at least one action performed by the user; generating image data indicative of the task in progress; displaying the ongoing task based on the image data; analyzing the ongoing task; generating an augmented image; overlaying the augmented image on the image data such that the augmented image is displayed simultaneously with the task in progress to instruct the user to proceed with the task in progress; A method comprising:

9. detecting that the at least one action is being performed erroneously; generating the augmented image in response to detecting that the at least one action has been erroneously performed; The method of claim 8 further comprising:

10. The method of claim 9 , wherein the augmented image is a modified image showing a modification of the at least one movement.

11. The method of claim 10 , further comprising generating a tactile alert in response to detecting that the at least one action is being performed erroneously.

12. determining a next action from the at least one action included in the task; generating the augmented image in response to determining the next action from the at least one action included in the task; 12. The method of any one of claims 8 to 11, further comprising:

13. The method of claim 12 , wherein the augmented image is an image that indicates the next action.

14. The method of claim 13 , further comprising generating a haptic alert in response to determining the next action.

15. On the computer, receiving input from a user indicating a task including at least one action to be performed by the user; responsive to receiving the input, monitoring an ongoing task including the at least one action performed by the user; generating image data indicative of the ongoing task; displaying the ongoing task based on the image data; analyzing the ongoing task; generating an augmented image; overlaying the augmented image on the image data so that the augmented image is displayed simultaneously with the task in progress to instruct the user to proceed with the task in progress; A computer program for recommending guidance instructions to a user.

16. The computer, detecting that the at least one operation is being performed erroneously; generating the augmented image in response to detecting that the at least one operation has been erroneously performed; 16. The computer program of claim 15, further comprising:

17. The computer program of claim 16 , wherein the augmented image is a modified image illustrating a modification of the at least one movement.

18. The computer, generating a tactile warning in response to detecting that the at least one action has been erroneously performed; 18. The computer program of claim 17, further comprising:

19. The computer, determining a next action from the at least one action included in the task; generating the augmented image of the next action in response to determining the next action from the at least one action included in the task; 19. A computer program product according to any one of claims 15 to 18, further comprising:

20. The computer, generating a haptic alert in response to determining said next action.

20. The computer program of claim 19, further comprising:

Citation Information

Patent Citations

  • Information processing device and method, and program

    JP2012066026A

  • Movement-information processing device

    JP2015061577A

  • USER GUIDANCE SYSTEMS AND METHODS, USING AUGMENTED REALITY DEVICES

    JP2018503416A

  • User operation evaluation system and computer program

    JP2020092944A

  • Virtual or augmented reality rehabilitation

    US20180315247A1