Self-learning artificial intelligence voice response based on user behavior during interaction

The AI ​​user activity guidance system solves the problem of misunderstanding in voice response systems by using real-time monitoring and augmented reality technology to guide users, thereby achieving real-time correction of user actions and improving task efficiency.

CN114519098BActive Publication Date: 2025-12-30INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111372119.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-20
Filing Date
2021-11-18
Publication Date
2025-12-30
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

Existing voice response systems often misunderstand user voice commands when performing user tasks, resulting in inaccurate actions and responses. Users need to interrupt the task to check the instructions to correct the erroneous actions, which increases the completion time.

Method used

The AI-powered user activity guidance system monitors user actions in real time, generates augmented image guidance, corrects incorrect actions in real time, and provides tactile and verbal suggestions, while augmented reality technology overlays correct action guidance.

Benefits of technology

Users can correct erroneous actions without interrupting the task, improving task completion efficiency, reducing the number of attention shifts, and increasing task accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519098B_ABST
    Figure CN114519098B_ABST
Patent Text Reader

Abstract

Embodiments of the invention relate to self-learning artificial intelligence voice responses based on user behavior during an interaction. A system for recommending guidance instructions to a user is provided. The system includes a memory having computer readable instructions and a processor for executing the computer readable instructions. The computer readable instructions control the processor to perform the following operations: monitor an ongoing task including at least one action performed by the user; generate image data depicting the ongoing task; and display the ongoing task based on the image data. The system analyzes the ongoing task and generates an augmented image. The augmented image is superimposed on the image data such that the augmented image is displayed concurrently with the ongoing task to guide the user in advancing the ongoing task.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This invention relates generally to artificial intelligence (AI) computing systems, and more specifically to AI recommendations for guiding instructions to users.

[0002] Voice response systems are becoming increasingly popular. Typically, a voice response system monitors ambient audio in response to a user's voice commands. The voice command guides the system to perform a specific action and provide an auditory response to the user. In some cases, the actions taken by the voice response system and the responses provided to the user may be inaccurate based on the user's expected context because the AI ​​used by the system misunderstands the words that make up the voice command. Summary of the Invention

[0003] According to a non-limiting embodiment, a system is provided for recommending guidance instructions to a user. The system includes a memory having computer-readable instructions and a processor for executing the computer-readable instructions. The computer-readable instructions control the processor to perform the following operations: monitor an ongoing task including at least one action performed by the user; generate image data depicting the ongoing task; and display the ongoing task based on the image data. The system analyzes the ongoing task and generates an enhanced image. The enhanced image is overlaid on the image data such that the enhanced image and the ongoing task are displayed simultaneously to guide the user through the ongoing task.

[0004] According to another non-limiting embodiment, a method for recommending guidance instructions to a user includes: monitoring an ongoing task, the ongoing task including at least one action performed by the user; generating image data depicting the ongoing task; and displaying the ongoing task based on the image data. The method further includes analyzing the ongoing task, generating an enhanced image, and overlaying the enhanced image onto the image data such that the enhanced image and the ongoing task are displayed simultaneously to guide the user through the ongoing task.

[0005] According to another non-limiting embodiment, a computer program product is provided for recommending guidance instructions to a user. The computer program product includes a computer-readable storage medium having program instructions embodied therein. The program instructions are readable by processing circuitry to cause the processing circuitry to perform operations including: monitoring an ongoing task comprising at least one action performed by a user; generating image data depicting the ongoing task; and displaying the ongoing task based on the image data. The operations further include: analyzing the ongoing task, generating an enhanced image, and overlaying the enhanced image onto the image data such that the enhanced image and the ongoing task are displayed simultaneously to guide the user through the ongoing task.

[0006] Additional features and advantages are achieved through the technology of this invention. Other embodiments and aspects of the invention described in detail herein are considered part of the claimed invention. For a better understanding of the advantages and features of the invention, refer to the specification and drawings. Attached Figure Description

[0007] The subject matter considered to be the invention is specifically pointed out and clearly claimed in the appended claims. The foregoing and other features and advantages of the invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings, wherein:

[0008] Figure 1 A cloud computing environment according to one or more embodiments of the present invention is described;

[0009] Figure 2 An abstract model layer is described according to one or more embodiments of the present invention;

[0010] Figure 3 An exemplary computer system capable of implementing one or more embodiments of the present invention is described;

[0011] Figure 4A An AI user activity guidance system according to one or more embodiments of the present invention is described;

[0012] Figure 4B An AI user activity guidance system according to one or more embodiments of the present invention is described;

[0013] Figure 5 The learned, modeled user activities are depicted according to one or more embodiments of the present invention;

[0014] Figure 6 The invention describes user activities that are incorrectly executed according to one or more embodiments of the invention.

[0015] Figure 7 Images depict user activities that are incorrectly executed, superimposed on correctly modeled user activities, according to one or more embodiments of the present invention;

[0016] Figure 8 Describe the learned, modeled user activities according to one or more embodiments of the present invention;

[0017] Figure 9 The invention describes user activities that are incorrectly executed according to one or more embodiments of the invention.

[0018] Figure 10 Images depict user activities that are incorrectly executed, superimposed on correctly modeled user activities, according to one or more embodiments of the present invention;

[0019] Figure 11 Images depicting user activities performed within a task according to one or more embodiments of the present invention;

[0020] Figure 12 The invention describes one or more embodiments thereof. Figure 11 The image shown is an image of the user activity being performed, overlaid with the next user activity included in the task;

[0021] Figure 13 This is a flowchart illustrating a method for recommending guidance instructions to a user, executed by an AI user activity guidance system according to one or more embodiments of the present invention.

[0022] Figure 14 A machine learning system that can be used to implement various embodiments of the present invention is described;

[0023] Figure 15 It describes what can be made by Figure 14 The machine learning system shown implements the learning phase; and

[0024] Figure 16 Exemplary computing systems capable of implementing various embodiments of the present invention are described. Detailed Implementation

[0025] Various embodiments of the invention are described herein with reference to the accompanying drawings. Alternative embodiments of the invention may be designed without departing from the scope of the invention. Various connections and positional relationships (e.g., above, below, adjacent, etc.) are illustrated between elements in the following description and drawings. Unless otherwise stated, these connections and / or positional relationships may be direct or indirect, and the invention is not schematically limited in this respect. Thus, coupling of entities may refer to direct or indirect coupling, and positional relationships between entities may be direct or indirect positional relationships. Furthermore, the various tasks and process steps described herein may be incorporated into a more comprehensive procedure or process with additional steps or functions not described in detail herein.

[0026] The following definitions and abbreviations are used to interpret the claims and description. As used herein, the terms “comprise,” “comprising,” “include,” “including,” “has,” “having,” “contain,” or “containing,” or any other variations thereof, are intended to cover a non-exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such a composition, mixture, process, method, article, or apparatus.

[0027] Additionally, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" can be understood to include any integer greater than or equal to one, i.e., one, two, three, four, etc. The term "multiple" can be understood to include any integer greater than or equal to two, i.e., two, three, four, five, etc. The term "connection" can include both indirect "connection" and direct "connection."

[0028] The terms “about,” “substantially,” “approximately,” and variations thereof are intended to include the degree of error associated with a measurement based on a specific quantity of equipment available at the time of filing of this application. For example, “about” may include a range of ±8%, 5%, or 2% of a given value.

[0029] For the sake of brevity, conventional techniques associated with the making and use of various aspects of the present invention may or may not be described in detail herein. Specifically, various aspects of computing systems and particular computer programs used to implement the various technical features described herein are well known. Therefore, for the sake of brevity, many conventional implementation details are only briefly mentioned herein, or omitted entirely, without providing well-known system and / or process details.

[0030] It should be understood that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings cited herein is not limited to cloud computing environments. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed hereafter.

[0031] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model may include at least five features, at least three service models, and at least four deployment models.

[0032] The features are as follows:

[0033] On-demand self-service: Cloud consumers can automatically and unilaterally supply computing power, such as server time and network storage, on demand, without requiring human interaction with the service provider.

[0034] Extensive network access: Capabilities are available on the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0035] Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned as needed. Location independence is significant because consumers typically do not have control or knowledge of the exact location of the provided resources, but can specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0036] Rapid flexibility: Capacity can be supplied quickly and flexibly, and in some cases automatically, to shrink rapidly and expand rapidly. For the consumer side, the available supply capacity often appears unlimited and can be purchased in any quantity at any time.

[0037] Measuring services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both the providers and consumers of the services being utilized.

[0038] The service model is as follows:

[0039] Software as a Service (SaaS): The capability provided to consumers is the ability to use the provider's applications running on cloud infrastructure. These applications are accessible from various client devices through thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage devices, or even the individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.

[0040] Platform as a Service (PaaS): This provides consumers with the ability to deploy applications created or acquired by the consumer onto cloud infrastructure. These applications are created using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage devices, but they have control over the deployed applications and the configuration of any application hosting environment.

[0041] Infrastructure as a Service (IaaS): The capability provided to consumers is the provision of processing, storage, networking, and other basic computing resources that enable consumers to deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but have control over the operating system, storage devices, deployed applications, and, possibly, limited control over the selection of networking components (e.g., host firewalls).

[0042] The deployment model is as follows:

[0043] Private cloud: Cloud infrastructure operated solely by an organization. It can be managed by the organization or a third party and can exist on-site or off-site.

[0044] Community cloud: A cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can exist on-site or off-site.

[0045] Public cloud: Cloud infrastructure that is made available to the public or large groups of industries and is owned by an organization that sells cloud services.

[0046] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain distinct entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursts for load balancing between clouds).

[0047] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. The core of cloud computing is its infrastructure, which includes a network of interconnected nodes.

[0048] Turning now to an overview more specifically related to aspects of the invention, hand-related activities (such as cooking, sewing, creative handicrafts, and musical instrument manipulation) typically involve performing several steps or actions to advance an ongoing task or accomplish a desired task. For example, a song often comprises several different arrangements of chords or notes. To perform the song as desired, the user must perform several different actions to manipulate the instrument in a manner necessary to achieve the correct chords or notes required to complete the task (i.e., to play the song correctly). However, when attempting to play a song, the user of the instrument may not realize they are playing the wrong chords or notes. Similarly, when learning a new song for the first time, the user may not know the next chord or note progression required to accurately advance the task (i.e., to continue playing the song). Conventional techniques require the user to stop the task and shift their attention from the instrument to the sheet music to determine the next chord or note progression. This action is typically performed several times before the user memorizes it, thus increasing the time required for the user to complete the task (i.e., to play the song correctly).

[0049] The one or more non-limiting embodiments described herein provide an AI user activity guidance system capable of learning various user actions necessary to perform a given task, monitoring the actions performed by the user in real time, and recommending guidance to the user on how to correctly perform one or more actions to advance or complete the task. In one or more non-limiting embodiments, the AI ​​user activity guidance system performs imaging of user actions as the user performs the task in real time and detects incorrectly performed actions. A display is provided showing an image of the user performing the actions of the task in real time. In response to detecting an incorrectly performed action, the AI ​​user activity guidance system generates a tactile warning to the user that the current action is incorrectly performed, and generates a recommended output indicating the correction of the incorrectly performed action. The recommended output includes verbal guidance on how to correct the incorrectly performed action and / or an enhanced image indicating the correct action superimposed on top of the displayed image of the ongoing task. In this way, the user can easily correct their actions without stopping the ongoing task and / or diverting their attention from the ongoing task.

[0050] Now for reference Figure 1The diagram illustrates an illustrative cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10, with local computing devices used by cloud consumers (such as, for example, personal digital assistants (PDAs) or mobile phones 54A, desktop computers 54B, laptop computers 54C, and / or automotive computer systems 54N) capable of communicating with the cloud computing nodes 10. The nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 50 to provide cloud consumers with infrastructure, platforms, and / or software-as-a-service that eliminates the need for them to maintain resources on their local computing devices. It should be understood that... Figure 1 The types of computing devices 54A-N shown are intended to be illustrative only, and computing node 10 and cloud computing environment 50 can communicate with any type of computerized device via any type of network and / or network-addressable connection (e.g., using a web browser).

[0051] Now for reference Figure 2 This demonstrates a cloud computing environment of 50 ( Figure 1 This provides a set of functional abstractions. It should be understood beforehand that... Figure 2 The components, layers, and functions shown are intended to be illustrative only, and embodiments of the invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0052] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: a mainframe 61; a RISC (Reduced Instruction Set Computer) based server 62; a server 63; a blade server 64; a storage device 65; and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0053] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 71; virtual storage device 72; virtual network 73, including virtual private network; virtual application and operating system 74; and virtual client 75.

[0054] In one instance, management layer 80 can provide the functionalities described below. Resource Provisioning 81 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and Pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User Portal 83 provides access to the cloud computing environment for consumers and system administrators. Service Level Management 84 provides allocation and management of cloud computing resources to ensure that required service levels are met. Service Level Agreement (SLA) Planning and Fulfillment 85 provides pre-scheduling and procurement of cloud computing resources, anticipating future demand for those resources according to the SLA.

[0055] Workload layer 90 provides examples of the functionality that can be leveraged in a cloud computing environment. Examples of workloads and functions that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analytics and processing 94; transaction processing 95; and artificial intelligence (AI) for training voice response systems 96.

[0056] We now turn to a more detailed description of various aspects of the invention. Figure 3 A high-level block diagram of an example of a computer-based system 300 that can be used to implement one or more embodiments of the present invention is shown. Although an exemplary computer system 300 is shown, the computer system 300 includes a communication path 326 connecting the computer system 300 to an additional system and may include one or more wide area networks (WANs) and / or local area networks (LANs), such as the Internet, intranets(s), and / or wireless communication networks(s). The computer system 300 and the additional system communicate via the communication path 326 (e.g., to transfer data between them).

[0057] Computer system 300 includes one or more processors, such as processor 302. Processor 302 is connected to communication infrastructure 304 (e.g., a communication bus, cross-over bar, or network). Computer system 300 may include display interface 306, which forwards graphics, text, and other data from communication infrastructure 304 (or from a frame buffer, not shown) for display on display unit 308. Computer system 300 also includes main memory 310, preferably random access memory (RAM), and may also include secondary memory 312. Secondary memory 312 may include, for example, hard disk drive 314 and / or removable storage drive 316, which represents, for example, a floppy disk drive, magnetic tape drive, or optical disk drive. Removable storage drive 316 reads from and / or writes to removable storage unit 318 in a manner known to those skilled in the art. Removable storage unit 318 represents, for example, a floppy disk, compact disk, magnetic tape, or optical disk read and written by removable storage drive 316. As will be appreciated, the removable storage unit 318 includes a computer-readable medium in which computer software and / or data are stored.

[0058] In some alternative embodiments of the invention, secondary memory 312 may include other similar means for allowing computer programs or other instructions to be loaded into a computer system. Such means may include, for example, removable memory unit 320 and interface 322. Examples of such means may include packages and package interfaces (such as those found in video game devices), removable memory chips (such as EPROM or PROM) and associated sockets, and other removable memory units 320 and interfaces 322 that allow software and data to be transferred from removable memory unit 320 to computer system 300.

[0059] Computer system 300 may also include a communication interface 324. Communication interface 324 allows software and data to be transferred between the computer system and external devices. Examples of communication interface 324 may include a modem, a network interface (such as an Ethernet card), a communication port, or a PCM-CIA slot and card. The software and data transmitted via communication interface 324 are in the form of signals, which may be, for example, electronic, electromagnetic, optical, or other signals that can be received by communication interface 324. These signals are provided to communication interface 324 via a communication path (i.e., channel) 326. Communication path 326 carries signals and may be implemented using wires or cables, optical fibers, telephone lines, cellular telephone links, RF links, and / or other communication channels.

[0060] In this disclosure, the terms "computer program medium," "computer-usable medium," and "computer-readable medium" are used to refer to media such as main memory 310 and secondary memory 312, removable storage drive 316, and hard disks installed in hard disk drive 314. A computer program (also referred to as computer control logic) is stored in main memory 310 and / or secondary memory 312. The computer program can also be received via communication interface 324. Such a computer program, when executed, enables the computer system to perform the features of this disclosure as discussed herein. Specifically, the computer program, when executed, enables processor 302 to perform features of the computer system. Thus, such a computer program represents the controller of the computer system.

[0061] In an exemplary embodiment, a system for training artificial intelligence (AI) in a voice response system is provided. In this exemplary embodiment, the voice response system is configured to monitor ambient audio in response to voice commands from a user. The voice response system uses AI to interpret the voice commands, and based on this interpretation, the voice response system provides a response to the user and / or performs an action requested by the user. The voice response system is also configured to monitor the user's reaction to the response provided by the voice response system or the action taken. A microphone and / or camera communicating with the voice response system may be used to monitor the reaction. The voice response system analyzes the user's reaction and determines whether the user is satisfied with the provided response or the action taken. If the user is dissatisfied with the provided response or the action taken, the voice response system updates the AI ​​model used to interpret the voice commands.

[0062] Now go to Figure 4A The diagram illustrates an AI user activity guidance system 400 according to a non-limiting embodiment. The AI ​​user activity guidance system 400 includes a computing system 12 that signals in communication with functional components. These functional components include, but are not limited to, one or more computing devices 430, an imaging device 432, and a voice-activated hub 420. Figures 1-3 Each of the devices, components, modules, and / or functions described herein may also be applied to Figure 4A The equipment, components, modules, and functions. Similarly, Figures 1-3 One or more of the operations and steps may also be included. Figure 4A In one or more operations or actions.

[0063] refer to Figure 4B This illustrates an AI user activity guidance system 400 according to another non-limiting embodiment. Figure 4BIn the non-limiting embodiment shown, several components of the AI ​​user activity guidance system 400 are integrated into a single smart wearable computing device 430 (e.g., smart glasses 430). The details of the individual components described below also apply to the AI ​​user activity guidance system 400 when it operates in a similar manner to the AI ​​user activity guidance system 400.

[0064] Computing device 430 includes, but is not limited to, televisions, smartphones, desktop computers, laptop computers, tablet computers, smartwatches, smart wearable devices, and / or other electronic / wireless communication devices that may have one or more processors, memories, and / or wireless communication technologies for displaying / streaming audio / video data. Computing device 430 may receive input (e.g., touch input, verbal input, mouse clicks, etc.) that can control the AI ​​user activity guidance system 400. In one or more non-limiting embodiments, the input may indicate the type of task to be performed by the user. Tasks include, for example, cooking, sewing, creative handicrafts, and instrument manipulation (i.e., playing an instrument). Regarding instrument manipulation, for example, the input may indicate the song to be played, the type of instrument to be used to play the song, and / or the tuning of the instrument. Furthermore, the input may include the difficulty level of the task to be performed. For instrument manipulation, for example, a beginner level of difficulty associated with playing a song via an instrument may include correctly executing closed or powerful chords, while an advanced level of playing the same song may include correctly executing the same chords as open major / minor chords.

[0065] Imaging device 432 includes, but is not limited to, a camera or video recorder. Therefore, imaging device 432 is capable of capturing images 434 of an ongoing task (i.e., a task performed in real time). Intelligent service 402 can work in conjunction with imaging device 432 to perform image recognition operations. For example, intelligent service 402 can monitor images of the ongoing task 432 and detect whether actions are performed correctly or incorrectly. Intelligent service 402 can monitor images of the ongoing task 432 and predict the next action to be included in the ongoing task 432.

[0066] The voice-activated hub 420 includes, for example, a personal assistant Internet of Things (IoT) computing device. The voice-activated hub 420 can detect voice-activated commands and / or queries and output spoken language including answers, instructions, recommendations, etc.

[0067] Computer system / server 12 is shown again, which can incorporate intelligent service 402 or "guided instruction intelligent recommendation service 402" (e.g., artificial intelligence simulating a human assistant "AISHA"). Figure 4AAs shown, computer system / server 12 can provide virtualized computing services (i.e., virtualized computing, virtualized storage, virtualized networking, etc.) to one or more computing devices, as described herein. More specifically, computer system / server 12 can provide virtualized computing, virtualized storage, virtualized networking, and other virtualization services executed on a hardware substrate.

[0068] Figure 4A The intelligent service 402 depicted (e.g., the guidance instruction intelligent recommendation service 402) communicates and / or is associated with the computing device 430, the imaging device 432, and the voice-activated hub 420. Therefore, the guidance instruction intelligent recommendation service 402, the computing device 430, the imaging device 432, and the voice-activated hub 420 can communicate via one or more communication methods (such as a computing network, a wireless communication network, or other network devices capable of communication). Figure 4A The networks (collectively referred to as "Network 18") are interconnected and / or communicate with each other. According to a non-limiting embodiment, the guidance instruction intelligent recommendation service 402 may be locally installed on the voice-activated hub 420 and / or computing device 430. Alternatively, the guidance instruction intelligent recommendation service 402 may be located outside the voice-activated hub 420 and / or computing device 430 (e.g., via a cloud computing server).

[0069] The guidance instruction intelligent recommendation service 402 can be incorporated into the processing unit 16 to perform various calculations, data processing, and other functionalities according to various non-limiting embodiments of the invention. Domain knowledge 412 (e.g., may include a database of ontology) is shown together with the guidance instruction component 404, analysis component 406, monitoring component 408, machine learning component 410, recognition component 414, and / or augmented reality (AR) component 416. In one or more non-limiting embodiments, one of the guidance instruction component 404, analysis component 406, monitoring component 408, machine learning component 410, recognition component 414, domain knowledge 412, and / or augmented reality (AR) component 416 is configured as an electronic hardware controller, which includes a memory and a processor configured to execute algorithms and computer-readable program instructions stored in the memory. Alternatively, the guidance instruction component 404, analysis component 406, monitoring component 408, machine learning component 410, recognition component 414, domain knowledge 412, and / or augmented reality (AR) component 416 can all be embedded or integrated into a single controller.

[0070] Domain knowledge 412 may include ontology representing concepts, keywords, expressions, and / or associated with a knowledge domain. The thesaurus or ontology may be used as a database and may also be used to identify semantic relationships between variables observed and / or unobserved by the machine learning component 410 (e.g., a cognitive component). According to a non-limiting embodiment, the term "domain" is intended to have its general meaning. Additionally, the term "domain" may include a system's area of ​​expertise or a collection of materials, information, content, and / or other resources related to one or more specific topics. A domain may refer to information related to any specific topic or a combination of selected topics.

[0071] A term ontology is also intended to be a term with its general meaning. According to one non-limiting embodiment, a term ontology, in its broadest sense, can include anything that can be modeled as an ontology, including but not limited to taxonomies, thesaurus, vocabularies, etc. For example, an ontology can include information or content related to a domain of interest or a particular category or concept. The ontology can be continuously updated using information synchronized with a source, adding information from the source as a model, a model's attribute, or an association between models within the ontology. In one or more non-limiting embodiments, the domain knowledge store learned model actions, which include images indicating exemplary or correct execution of actions included in a given task. For example, in the case of a musical instrument, learned model actions could include images showing how to play the correct chords or notes on a given instrument.

[0072] Additionally, domain knowledge 412 may include one or more external resources, such as links to one or more Internet domains, web pages, etc. For example, text data may be hyperlinked to web pages that can describe, explain, or provide additional information related to the text data. Thus, the summary can be enhanced via links to external resources that further explain, instruct, interpret, provide context and / or additional information to support decision-making, alternative suggestions, alternative choices, and / or criteria.

[0073] The analysis component 406 of the computer system / server 12 can work in conjunction with the processing unit 16 to implement various embodiments of the present invention. For example, the analysis component 406 can undergo different data analysis functions to analyze data transmitted from one or more devices (e.g., voice-activated hub 420 and / or computing device 430).

[0074] Analysis component 406 can receive and analyze each physical property associated with media data (e.g., audio data and / or video data). Analysis component 406 can cognitively receive and / or detect audio data and / or video data used to guide instruction component 404.

[0075] Analysis component 406, monitoring component 408, and / or machine learning component 410 can access and monitor one or more audio and / or video data sources (e.g., websites, audio storage systems, video storage systems, cloud computing systems, etc.) to provide audio data, video data, and / or text data for providing guidance instructions to perform tasks. Analysis component 406 can cognitively analyze data retrieved from domain knowledge 412, one or more online sources, cloud computing systems, text corpora, or combinations thereof. Analysis component 406 and / or machine learning component 410 can use natural language processing (“NLP”) to extract one or more keywords, phrases, instructions, and / or transcripts (e.g., transcribing audio data into text data).

[0076] As part of the detection data, the analysis component 406, the monitoring component 408, and / or the machine learning component 410 can identify audio data, video data, text data, and / or contextual factors associated with audio data, video data, and / or text data, or combinations thereof, from one or more sources. Furthermore, the machine learning component 410 can initiate machine learning operations to learn contextual factors associated with the audio data, video data, and / or text data associated with instructions used to perform tasks, such as assembling and / or repairing / restoring items (e.g., assembling a new bicycle or repairing a computer).

[0077] The monitoring component 408 can monitor the execution of the ongoing task 434 in real time via the imaging device 432 and one or more user voices / words 423 (via microphone 433 and / or voice center 420) while the task 434 is being performed. The recognition component 414 can identify the user performing the task 434 on an object using the voice-activated center 420 and / or computing device 430. For example, the voice-activated center 420, computing device 430, and / or imaging device 432 can identify one or more activities, body movements and / or features (e.g., facial recognition, facial expressions, hand / foot gestures, etc.), behaviors, audio data (e.g., voice detection and / or recognition), environmental context, or other defined parameters / features that can identify, locate, and / or recognize the user and / or the task performed by the user. The analysis component 406 can work in conjunction with the monitoring component to perform image recognition to identify all objects in a given image or video clip among objects in each frame.

[0078] In one or more non-limiting embodiments, the monitoring component 408 can detect incorrectly performed actions included in the ongoing task 434. In response to detecting an incorrect action and / or the next one, the monitoring component can instruct one or more computing devices 430 to generate a haptic warning 436 alerting the user that the current action is being performed incorrectly. Similarly, the monitoring component 408 can predict the next action to be performed in the ongoing task 434 and generate a haptic warning 436 alerting the user to the next action to be performed.

[0079] The guidance instruction component 404 can provide one or more guidance instructions to assist in performing a selected task based on identified contextual factors. Guidance instructions can be text data, audio data, and / or video data. For example, a voice-activated hub 420 can audibly deliver guidance instruction 422. The computing device 430 can provide guidance instruction 450 together with image / video data 485 displayed by the graphical user interface (“GUI”) of the computing device 430 and / or as a voice command output from speaker 431.

[0080] The guidance instruction component 404 can cognitively guide a user to perform a selected task using one or more guidance instructions 422. The guidance instruction component 404 can provide a sequence of guidance instructions 422 retrieved from domain knowledge, one or more online sources, cloud computing systems, text corpora, or combinations thereof. The guidance instruction component 404 can provide media data from one or more online sources, cloud computing systems, or combinations thereof.

[0081] The guidance instruction component 404 can verify each step of one or more guidance instructions 422 used to assist in performing a selected task. The guidance instruction component 404 can also identify the level of difficulty (e.g., level of stress, frustration, anxiety, excitement, or other emotional response) achieved by the user when performing a set of tasks associated with one or more guidance instructions 422 delivered via streaming media, pause / stop / terminate the delivery of streaming media for a selected period of time, and / or provide the user with a modified set of guidance instructions 422 to guide the user with an enhanced level of instruction.

[0082] The guidance instruction component 404 can provide additional guidance information related to the guidance instruction 422, collected from domain knowledge, one or more online sources, cloud computing systems, text corpora, or a combination thereof, for performing the selected task. For example, if the first set of instructions is insufficient for the user, an additional set can be provided that can further explain one or more of the original instructions.

[0083] According to one or more non-limiting embodiments, the guidance instruction component 404 may also work with the AR component 416 to provide enhanced guidance instructions 422 for assisting in the performance of a selected task. More specifically, the AR component 416 may generate an enhanced image 500 overlaid on real-time image / video data 485 displayed on the graphical user interface (GUI) 429. Thus, the enhanced image 500 is displayed along with the ongoing task 434, such that the enhanced image 500 shows the user how to correct incorrectly performed actions. In this way, the user can easily correct their actions without stopping the ongoing task 434 and / or diverting their attention from the ongoing task 434. In another non-limiting embodiment, the enhanced image 500 may show the user how to correctly perform the next action included in the ongoing task 434.

[0084] The intelligent recommendations of the guidance instruction service 402 can adjust the tone, volume, pace, and / or frequency of the audio / media data of the guidance instruction 422 based on the speed / pacing of the user following the guidance instruction 422. Furthermore, words, phrases, and / or complete sentences (e.g., all or part of a conversation) from other parties to the audio data can be transcribed into text form based on NLP extraction operations (e.g., NLP-based keyword extraction). The text data can be relayed, transmitted, stored, or further processed so that the same audio / video data (e.g., all or part of a conversation) can be heard or listened to simultaneously with the text version of the guidance instruction.

[0085] As previously indicated, the guidance instruction intelligent recommendation service 402 can also communicate with other linked devices, such as, for example, a voice-activated hub 420, a computing device 430, and / or an imaging device 432. Furthermore, the analysis component 406 and / or the machine learning component 410 can access one or more online data sources, such as social media networks, websites, or data sites for providing one or more guidance instructions 422 to assist in the performance of a selected task based on identified contextual factors. That is, the analysis component 406, the identification component 414, and / or the machine learning component 410 can learn and observe the degree or level of user attention, the level of difficulty the user faces when performing the task, the type of response, and / or feedback on various topics and / or guidance instructions 422. The user's learned and observed behavior can be linked to various data sources that provide personal information, social media data, or user profile information to learn, build, and / or determine confidence factors related to the performance of guidance instructions 422.

[0086] In one or more non-limiting embodiments, the AI ​​user activity guidance system 400 can learn and recommend actions to users by observing how others (users) react to the same methods. Iterative feedback between crowdsourced data and observed users can then be used to make intelligent decisions about the usage path and projection for a given task. Furthermore, the AI ​​user activity guidance system 400 can proactively modify the expected path of an ongoing task based on the learning success rate.

[0087] According to a non-limiting embodiment, the machine learning component 410 described herein can be performed by a variety of methods or combinations thereof, such as supervised learning, unsupervised learning, temporal difference learning, reinforcement learning, and so on. Some non-limiting examples of supervised learning that can be used with this technique include AODE (Average Single Dependency Estimator), artificial neural networks, backpropagation, Bayesian statistics, Naive Bayes classifiers, Bayesian networks, Bayesian knowledge bases, case-based reasoning, decision trees, inductive logic programming, Gaussian process regression, gene expression programming, group methods for data processing (GMDH), learning automata, learning vector quantization, minimum message length (decision trees, decision graphs, etc.), lazy learning, instance-based learning, nearest neighbor algorithms, simulation modeling, possibly approximating correct (PAC) learning, chaining rules, knowledge acquisition methods, symbolic machine learning algorithms, sub-symbolic machine learning algorithms, support vector machines, random forests, multi-classifier ensembles, guided clustering (bagging), boosting (meta-algorithms), ordinal classification, regression analysis, information fuzzy networks (IFN), statistical classification, linear classifiers, Fisher linear discriminant analysis, logistic regression, perceptrons, support vector machines, quadratic classifiers, k-nearest neighbors, hidden Markov models, and boosting. Some non-limiting examples of unsupervised learning that can be used with this technique include artificial neural networks, data clustering, expectation maximization, self-organizing maps, radial basis function networks, vector quantization, topographic mapping, information bottleneck methods, IBSEAD (Interaction Based on Distributed Autonomous Entities), association rule learning, prior algorithms, Eclat algorithm, FP-growth algorithm, hierarchical clustering, single-link clustering, concept clustering, partitioning clustering, k-means algorithm, fuzzy clustering, and reinforcement learning. Some non-limiting examples of temporal difference learning may include Q-learning and learning automata. Specific details regarding any of the examples of supervised, unsupervised, temporal difference, or other machine learning described in this paragraph are known and within the scope of this disclosure. Moreover, when deploying one or more machine learning models, the computing device can be tested in a controlled environment before being deployed in a public environment. Furthermore, even when deployed in a public environment (e.g., outside of a controlled testing environment), the compliance of the computing device can be monitored.

[0088] According to a non-limiting embodiment, the guiding instruction intelligent recommendation service 402 may perform one or more calculations based on mathematical operations or functions that may involve one or more mathematical operations (e.g., analytically or computationally solving differential equations or partial differential equations, using addition, subtraction, division, multiplication, standard deviation, mean, average, percentage, statistical modeling using statistical distributions, finding the minimum, maximum, or similar threshold of combined variables, etc.). Therefore, as used herein, computational operations may include all or part of one or more mathematical operations.

[0089] According to a non-limiting embodiment, if the task a user wants to perform cannot be detected initially, the user can (e.g., verbally via a voice-activated hub 420 and / or microphone 433 and / or via an interactive GUI 429 interface of computing device 430) provide activity data to the guidance instruction intelligent recommendation service 402 as input, so that the guidance instruction intelligent recommendation service 402 can begin object scanning, instruction scanning (or after downloading to the corpus if not already downloaded), and guide the user with step-by-step instructions based on monitoring the user's activities.

[0090] As described herein, the AI ​​user activity guidance system 400 includes an AR component 416 configured to generate an enhanced image 500 overlaid on real-time image / video data to provide enhanced guidance instructions 422 to assist the user in performing tasks. Figure 5 , Figure 6 and Figure 7 Together they depict an example of the task while a user is playing the guitar.

[0091] First go to Figure 5 According to one or more embodiments of the present invention, a learned, modeled user activity is depicted. In this example, the learned, modeled user activity 600 is a learned correctly played guitar closed chord 600 (e.g., a closed C chord 600). The correctly played guitar closed chord 600 can be learned by the machine learning component 410 as described herein and stored in domain knowledge 412 for future reference by the AI ​​user activity guidance system 400 (e.g., AR component 416). The chord diagram 602 corresponding to the correctly played guitar closed chord 600 can also be stored in domain knowledge 412 for future reference by the AI ​​user activity guidance system 400.

[0092] Figure 6A GUI 429 depicts image / video data 485 showing an ongoing task 434 being performed in real time (e.g., a user playing a desired song on a guitar), where the user is incorrectly performing user activities included in task 434. In this example, incorrect user activity is incorrectly performing guitar chords (e.g., an incorrect closed C chord). In one or more non-limiting embodiments, the GUI 429 may also display a chord diagram 602 including indicators 604 indicating the actual guitar strings being played by the user. In one or more embodiments, the display of the indicators (e.g., color, shape, etc.) may be varied to indicate which specific guitar strings are being played incorrectly.

[0093] Figure 7 A GUI 429 is depicted displaying an enhanced image 500 overlaid on real-time image / video data 485. As described herein, the enhanced image 500 shows how a user can correct incorrectly performed actions. In this example, the enhanced image 500 shows how a user can correctly perform guitar chord progressions (e.g., a correct closed C chord) to properly advance or complete a task (e.g., play a song correctly). Therefore, the user can easily correct their actions without stopping and / or taking their attention away from the guitar. In one or more non-limiting embodiments, the GUI 429 may also display a chord chart 602 including a correction indicator 606 that indicates the guitar strings that should be played to correctly play a song.

[0094] Figure 8 , Figure 9 and Figure 10 Together, an example of a task performed by a user while playing a guitar is depicted according to another non-limiting embodiment. As mentioned herein, the user can input a difficulty level corresponding to the selected task to be performed. For example, the previously described... Figure 5 , Figure 6 and Figure 7 It can correspond to a user's request to play a specific song at a beginner level. Therefore, the AI ​​user activity guidance system 400 can obtain a learned model of closed chords from the domain knowledge 412.

[0095] However, when a user inputs a request to play a song at a higher level, the AI ​​user activity guidance system 400 can draw upon domain knowledge 412 (see [link to relevant documentation]). Figure 8 The learned model of open major / minor chords (e.g., open guitar C chords) is obtained as an image 600, which can be more complex to play than closed chords.

[0096] exist Figure 9In the GUI 429, image / video data 485 of an ongoing task 434 being performed in real time (e.g., a user playing a desired song on a guitar) is displayed, in which the user is incorrectly executing an open C chord. Chord diagram 602 includes indicators 604 showing which strings the user is pressing incorrectly.

[0097] exist Figure 10 In this context, GUI 429 displays an enhanced image 500 overlaid on real-time image / video data 485. As described herein, the enhanced image 500 shows how a user can correctly execute open chords (e.g., a correct open C chord) to properly advance the task or complete it at a higher difficulty level. GUI 429 also displays a chord diagram 602 including a correction indicator 606 that indicates how to correct actions, such as playing an open C chord correctly.

[0098] As described herein, the AI ​​user activity guidance system 400 can monitor an image of an ongoing task 432 and predict the next action included in the ongoing task 432. Therefore, an enhanced image 500 can be generated to inform the user of the next action to be performed in the ongoing task 432.

[0099] Figure 11 and Figure 12 Together, they depicted an AI user activity guidance system 400, which predicts the next guitar chord to be played in an ongoing song performed by the user. Figure 11 For example, GUI 429 displays image / video data 485 of the user playing an open A chord, which is included in the song the user is playing in real time. As the song progresses, the AI ​​user activity guidance system 400 identifies the next chord in the song as an open C chord. Therefore, as... Figure 12 As shown, the AI ​​user activity guidance system 400 proactively generates an enhanced image 500. The enhanced image 500 is overlaid on the image / video data 485 to inform or guide the user on how to transition from the current action (e.g., an open A chord) to the next action in the task (e.g., an open C chord). In this way, the user can accurately continue playing the song without taking their attention away from the ongoing task.

[0100] Now go to Figure 13According to one or more embodiments of the present invention, a method is shown for recommending guidance instructions to a user by an AI user activity guidance system 400. The method begins at operation 800, and at operation 802, the AI ​​user activity guidance system 400 determines a task to be performed by the user. The task may include multiple user actions to be performed by the user and may be determined in response to receiving input from the user instructing the task (e.g., touch input, voice input, etc.). At operation 804, the AI ​​user activity guidance system 400 obtains one or more learned model actions included in the task. The learned model actions may be obtained from domain knowledge 412. In one or more non-limiting embodiments, the learned model actions include images indicating exemplary or correct execution of the action. At operation 806, the AI ​​user activity guidance system 400 generates image data of the ongoing task performed by the user in real time. The image data may include, for example, a video stream generated by a camera monitoring the ongoing task.

[0101] Moving to operation 808, the AI ​​user activity guidance system 400 analyzes the image data of the ongoing task and, at operation 810, determines whether the current action included in the ongoing task has been correctly performed. If the action has been correctly performed, the AI ​​user activity guidance system 400 determines whether the task has been completed, i.e., whether all actions included in the task have been performed. When the task is completed, the method ends. Otherwise, the AI ​​user activity guidance system 400 proceeds to operation 824 to determine the next action included in the task, which is discussed in more detail below.

[0102] However, when an action is performed incorrectly, the AI ​​user activity guidance system 400 generates a tactile warning at operation 812, indicating that the current action is being performed incorrectly. At operation 814, the AI ​​user activity guidance system 400 accesses domain knowledge 412 to obtain a learned modeling image of the correct action, which will be used to augment the image data. At operation 816, the AI ​​user activity guidance system 400 augments the image data by overlaying the learned modeling image onto the image data. Thus, the user viewing GUI 429 can discern how to correct the currently incorrectly performed action. At operation 818, the AI ​​user activity guidance system 400 analyzes the image data to determine if the user has adjusted their execution based on the augmented image to correct their action. If the incorrect action has not yet been corrected, the method returns to operation 816 and continues to overlay the learned modeling image onto the image data until the user corrects the incorrect action. When the incorrect action has been corrected, the method proceeds to operation 820 to determine if the task has been completed. When the task has been completed, the method terminates at operation 822.

[0103] When the task is not completed, at operation 824, the AI ​​user activity guidance system 400 determines the next action included in the task. At operation 826, the AI ​​user activity guidance system 400 generates a haptic warning, which notifies the user that the next action in the task is to be performed. At operation 828, the AI ​​user activity guidance system 400 accesses domain knowledge 412 to obtain a learned modeling image of the next action included in the task, which will be used to augment the image data. At operation 830, the AI ​​user activity guidance system 400 augments the image data by overlaying the learned modeling image of the next action onto the image data. Thus, the user viewing GUI 429 can quickly move to the next action included in the task without taking their attention away from the ongoing task. The method returns to operation 810, where the AI ​​user activity guidance system 400 analyzes whether the next action has been correctly performed, and the method continues as described above.

[0104] Additional details will now be provided regarding machine learning techniques that can be used to implement parts of the computer system / server 12. The various types of computer control functionalities described herein (e.g., estimation, determination, decision-making, recommendation, etc. of the computer system / server 12) can be implemented using machine learning and / or natural language processing techniques. Typically, machine learning techniques operate on so-called “neural networks,” which can be implemented as a programmable computer configured to run a set of machine learning algorithms. Neural networks incorporate knowledge from various disciplines, including neurophysiology, cognitive science / psychology, physics (statistical mechanics), control theory, computer science, artificial intelligence, statistics / mathematics, pattern recognition, computer vision, parallel processing, and hardware (e.g., digital / analog / VLSI / optics).

[0105] The fundamental function of neural networks and their machine learning algorithms is to identify patterns by interpreting unstructured sensor data via a form of machine perception. Unstructured real-world data, in its native form (e.g., images, sounds, text, or time-series data), is transformed into a numerical form (e.g., vectors with magnitude and direction) that can be understood and manipulated by a computer. Machine learning algorithms perform multiple iterations of learning-based analysis on the real-world data vectors until the patterns (or relationships) contained within the real-world data vectors are revealed and learned. The learned patterns / relationships act as predictive models that can be used to perform a variety of tasks, including, for example, the classification (or labeling) and clustering of real-world data. Classification tasks typically rely on training a neural network (i.e., a model) using a labeled dataset to identify correlations between labels and data. This is known as supervised learning. Examples of classification tasks include detecting people / faces in images, recognizing facial expressions in images (e.g., anger, happiness, etc.), identifying objects in images (e.g., stop signs, pedestrians, lane markings, etc.), recognizing gestures in videos, detecting musical instruments and instrument manipulation, detecting hand activities (e.g., cooking, cross-stitching, sewing, etc.), detecting speech, detecting speech in audio, identifying specific speakers, and transcribing speech into text. Clustering tasks identify the similarity between objects by grouping them according to those common characteristics and distinguishing them from other groups of objects. These groups are called "clusters."

[0106] Reference Figure 14 and 15 Examples of machine learning techniques that can be used to implement various aspects of the present invention are described. References will be made to... Figure 14 This describes a machine learning model configured and arranged according to embodiments of the present invention. References will be made to... Figure 16 A detailed description of example computing systems and network architectures capable of implementing one or more embodiments of the invention described herein is provided.

[0107] Figure 14A block diagram illustrating a classifier system 1200 capable of implementing various aspects of the invention described herein is depicted. More specifically, the functionality of system 1200 in embodiments of the invention is used to generate various models and sub-models that can be used to implement the computer functionality in embodiments of the invention. System 1200 includes multiple data sources 1202 communicating with classifier 1210 via network 1204. In some aspects of the invention, data sources 1202 may bypass network 1204 and be fed directly into classifier 1210. According to embodiments of the invention, data sources 1202 provide data / information inputs to be evaluated by classifier 1210. Data sources 1202 also provide data / information inputs that can be used by classifier 1210 to train and / or update models(s)1216 created by classifier 1210. Data sources 1202 can be implemented as a wide variety of data sources, including but not limited to sensors, data repositories (including training data repositories), cameras, and outputs from other classifiers configured to collect real-time data. Network 1204 can be any type of communication network, including but not limited to local area networks, wide area networks, private networks, the Internet, etc.

[0108] The classifier 1210 can be implemented by a programmable computer (such as a processing system 1400). Figure 16 The algorithm executed is shown in the diagram. Figure 14 As shown, classifier 1210 includes a set of machine learning (ML) algorithms 1212; natural language processing (NLP) algorithms 1214; and multiple models 1216 as relation (or prediction) algorithms generated (or learned) by the ML algorithms 1212. For ease of illustration and explanation, the algorithms 1212, 1214, and 1216 of classifier 1210 are depicted separately. In embodiments of the invention, the functions performed by the various algorithms 1212, 1214, and 1216 of classifier 1210 may be distributed differently than shown. For example, when classifier 1210 is configured to perform an overall task with sub-tasks, the set of ML algorithms 1212 may be divided such that a portion of ML algorithms 1212 performs each sub-task, and a portion of ML algorithms 1212 performs the overall task. Additionally, in some embodiments of the invention, NLP algorithm 1214 may be integrated within ML algorithm 1212.

[0109] NLP algorithm 1214 includes a speech recognition function that allows classifier 1210, and more specifically ML algorithm 1212, to receive natural language data (text and audio) and apply elements of language processing, information retrieval, and machine learning to derive meaning from the natural language input and potentially take actions based on the derived meaning. NLP algorithm 1214, used according to aspects of the invention, may also include a speech synthesis function that allows classifier 1210 to translate results(s)1220 into natural language (text and audio) to convey aspects of results(s)1220 as natural language communication.

[0110] NLP and ML algorithms 1214 and 1212 receive and evaluate input data (i.e., training data and data in analysis) from data source 1202. ML algorithm 1212 includes the functionality necessary to interpret and utilize the format of the input data. For example, if data source 1202 includes image data, ML algorithm 1212 may include visual recognition software configured to interpret image data. ML algorithm 1212 applies machine learning techniques to the received training data (e.g., data received from one or more data sources in data source 1202, images and / or sound extracted from a video stream, etc.) to create / train / update one or more models 1216 over time, which one or more models 1216 model the overall task and sub-tasks to be performed by classifier 1210.

[0111] Now let's refer to each other. Figure 14 and Figure 15 , Figure 15 An example of a learning phase 1300 performed by ML algorithm 1212 to generate the aforementioned model 1216 is depicted. In the learning phase 1300, classifier 1210 extracts features from the training data and transforms the features into vector representations that can be recognized and analyzed by ML algorithm 1212. The feature vectors are analyzed by ML algorithm 1212 to “classify” the training data against the target model (e.g., correcting “the model that includes actions in a task or the task of the model”) and to reveal the relationships between and among the classified training data. Examples of suitable implementations of ML algorithm 1212 include, but are not limited to, neural networks, support vector machines (SVM), logistic regression, decision trees, hidden Markov models (HMM), etc. The learning or training performed by ML algorithm 1212 can be supervised, unsupervised, or a mixture of aspects of supervised and unsupervised learning. Supervised learning occurs when training data is already available and has been classified / labeled. Unsupervised learning occurs when training data has not been classified / labeled and therefore must be developed through iterations of classifier 1210 and ML algorithm 1212. Unsupervised learning can utilize additional learning / training methods, including, for example, clustering, anomaly detection, neural networks, deep learning, etc.

[0112] When model 1216 is fully trained by ML algorithm 1212, data source 1202 that generates “real-world” data is accessed, and the “real-world” data is applied to model 1216 to generate a usable version of result 1220. In some embodiments of the invention, result 1220 may be fed back to classifier 1210 and used by ML algorithm 1212 as additional training data for updating and / or refining model 1216.

[0113] In various aspects of the invention, ML algorithm 1212 and model 1216 can be configured to apply confidence levels (CL) to individual results / determinations (including result 1220) in their results / determinations to improve the overall accuracy of a particular result / determination. When the CL value of a determination or result generated by ML algorithm 1212 and / or model 1216 is lower than a predetermined threshold (TH) (i.e., CL < TH), the result / determination can be classified as having a sufficiently low “confidence” to justify the conclusion that the determination / result is invalid, and this conclusion can be used to determine when, how, and / or whether to process the determination / result in downstream processing. If CL > TH, then the determination / result can be considered valid, and this conclusion can be used to determine when, how, and / or whether to process the determination / result in downstream processing. Many different predetermined TH levels can be provided. Determinations / results with CL > TH can be ranked from the highest CL > TH to the lowest CL > TH to determine the priority of when, how, and / or whether to process the determination / result in downstream processing.

[0114] In an aspect of the invention, classifier 1210 can be configured to apply a confidence level (CL) to result 1220. When classifier 1210 determines that the CL in result 1220 is below a predetermined threshold (TH) (i.e., CL < TH), result 1220 can be classified as low enough to justify the "no confidence" classification in result 1220. If CL > TH, then result 1220 can be classified as high enough to justify the valid determination of result 1220. Many different predetermined TH levels can be provided, allowing results 1220 with CL > TH to be ranked from the highest CL > TH to the lowest CL > TH.

[0115] The functions performed by classifier 1210 (and more specifically by ML algorithm 1212) can be organized as a weighted directed graph, where nodes are artificial neurons (e.g., modeled based on neurons in the human brain), and where weighted directed edges connect nodes. The directed graph of classifier 1210 can be organized such that some nodes form input layer nodes, some nodes form hidden layer nodes, and some nodes form output layer nodes. Input layer nodes are coupled to hidden layer nodes, and hidden layer nodes are coupled to output layer nodes. Each node is connected to every node in the adjacent layers via a connection path, which can be depicted as directed arrows with respective connection strengths. Multiple input layers, multiple hidden layers, and multiple output layers can be provided. When multiple hidden layers are provided, classifier 1210 can perform unsupervised deep learning to perform the assigned tasks of classifier 1210(s).

[0116] Similar to the functionality of the human brain, each input layer node receives input without connection strength adjustment and without node summation. Each hidden layer node receives its input from all input layer nodes based on the connection strength associated with the relevant connection paths. Similar connection strength multiplication and node summation are performed on both hidden layer nodes and output layer nodes.

[0117] Weighted directed classifier 1210 Figure 1 Data records are processed one by one (e.g., output from data source 1202), and they "learn" by comparing an initial arbitrary classification of a record with the known actual classification of the record. Using a training method called "backpropagation" (i.e., "backpropagation of error"), the error of the initial classification from the first record is fed back into the weighted directed graph of classifier 1210 and used to modify the weighted connections of the weighted directed graph again, and this feedback process continues for many iterations. During the training phase of the weighted directed graph of classifier 1210, the correct classification for each record is known, and these output nodes can therefore be assigned "correct" values. For example, the node value corresponding to the correct class is "1" (or 0.9), and the node value of other nodes is "0" (or 0.1). Therefore, it is possible to compare the calculated values ​​of the weighted directed graph of the output nodes with these "correct" values ​​and compute an error term (i.e., an "incremental (delta)" rule) for each node. These error terms are then used to adjust the weights in the hidden layers so that in the next iteration, the output values ​​will be closer to the "correct" values.

[0118] Figure 16A high-level block diagram of a computer system 1400 is depicted, which can be used to implement one or more computer processing operations according to aspects of the present invention. Although an exemplary computer system 1400 is shown, the computer system 1400 includes a communication path 1425 connecting the computer system 1400 to an additional system (not shown) and may include one or more wide area networks (WANs) and / or local area networks (LANs), such as the Internet, intranet(s), and / or wireless communication networks(s). The computer system 1400 and the additional system communicate via the communication path 1425, for example, to transfer data between them. In some embodiments of the invention, the additional system may be implemented as one or more cloud computing systems 50. The cloud computing system 50 may supplement, support, or replace some or all of the functionality of the computer system 1400 (in any combination), including any and all computing systems described in this specific embodiment that can be implemented using the computer system 1400. Additionally, some or all of the functionality of the various computing systems described in this specific embodiment may be implemented as nodes of the cloud computing system 50.

[0119] Computer system 1400 includes one or more processors, such as processor 1402. Processor 1402 is connected to communication infrastructure 1404 (e.g., a communication bus, crossbar switch, or network). Computer system 1400 may include a display interface 1406 that forwards graphics, text, and other data from communication infrastructure 1404 (or from a frame buffer, not shown) for display on display unit 1408. Computer system 1400 also includes main memory 1410, preferably random access memory (RAM), and may also include secondary memory 1412. Secondary memory 1412 may include, for example, hard disk drive 1414 and / or removable storage drive 1416, which represents, for example, a floppy disk drive, magnetic tape drive, or optical disk drive. Removable storage drive 1416 reads from and / or writes to removable storage unit 1418 in a manner known to those skilled in the art. Removable storage unit 1418 represents, for example, a floppy disk, compact disk, magnetic tape or optical disk, flash memory drive, solid-state memory, etc., read from and written to by removable storage drive 1416. As will be appreciated, the removable storage unit 1418 includes a computer-readable medium in which computer software and / or data are stored.

[0120] In an alternative embodiment of the invention, secondary memory 1412 may include other similar means for allowing computer programs or other instructions to be loaded into a computer system. Such means may include, for example, removable memory unit 1420 and interface 1422. Examples of such means may include packages and package interfaces (e.g., packages and package interfaces found in video game devices), removable memory chips (such as EPROM or PROM) and associated sockets, and other removable memory units 1420 and interfaces 1422 that allow software and data to be transferred from removable memory unit 1420 to computer system 1400.

[0121] Computer system 1400 may also include a communication interface 1424. Communication interface 1424 allows software and data to be transferred between the computer system and external devices. Examples of communication interface 1424 may include a modem, a network interface (such as an Ethernet card), a communication port, or a PCM-CIA slot and card. The software and data transferred via communication interface 1424 are in the form of signals, which may be, for example, electronic, electromagnetic, optical, or other signals that can be received by communication interface 1424. These signals are provided to communication interface 1424 via a communication path (i.e., channel) 1425. Communication path 1425 carries signals and may be implemented using wires or cables, optical fibers, telephone lines, cellular telephone links, RF links, and / or other communication channels.

[0122] The technical benefits include improved functionality of the AI ​​computing system, which can learn a variety of different user actions necessary to perform a given task, monitor the actions performed by the user in real time to achieve the desired task, and recommend guidance to the user on how to correctly perform one or more of the actions required to achieve the task. In one or more non-limiting embodiments, the AI ​​user activity guidance system performs imaging of the user's actions as the user performs the task in real time and detects actions that are performed incorrectly. A display showing images of the user's actions performing the task in real time is provided. In response to detecting an incorrectly performed action, the AI ​​user activity guidance system generates a tactile warning to the user that the current action is performed incorrectly, and generates recommended output indicating the correction of the incorrectly performed action. The recommended output includes verbal guidance on how to correct the incorrectly performed action and / or an enhanced image indicating the correct action, which is superimposed on top of the image shown on the display. In this way, the user can easily correct their actions without stopping the ongoing task and / or diverting their attention from the ongoing task. Therefore, the AI ​​user activity guidance system described herein facilitates the user's ability to complete tasks more quickly while avoiding errors in the completed task.

[0123] The present invention can be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.

[0124] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital universal disc (DVD), memory sticks, floppy disks, mechanically encoded devices (such as punched cards or raised structures in grooves with instructions recorded thereon), and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0125] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the corresponding computing / processing device.

[0126] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++, as well as conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet provided by an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry devices, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuitry devices in order to perform aspects of this invention.

[0127] This document describes aspects of the invention with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0128] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more boxes of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices in a particular manner, such that the computer-readable storage medium having the instructions stored therein includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0129] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to produce a computer-implemented process on the computer, other programmable apparatus, or other device, such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may identify a portion of a module, segment, or instruction, which includes one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions marked in the blocks may occur in a different order than indicated in the figures. For example, depending on the functionality involved, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0131] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements to technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

Claims

1. A system for recommending guidance instructions to a user, comprising: a memory having computer readable instructions; and one or more processors for executing the computer readable instructions, the computer readable instructions controlling the one or more processors to perform operations comprising: monitoring an ongoing task comprising at least one action performed by a user, the at least one action being an action among a plurality of actions to be performed by the user in progressing to correctly complete the ongoing task; generating image data depicting the ongoing task; displaying a current action of the ongoing task being performed by the user based on the image data; analyzing the ongoing task and predicting a next action in the progression; generating an augmented image depicting at least a portion of the ongoing task, the augmented image being an image depicting correct performance of the predicted next action; and superimposing the augmented image on the image data such that the current action and the predicted next action are displayed simultaneously to guide the user in advancing the ongoing task.

2. The system of claim 1, further comprising generating a haptic alert in response to detecting that the at least one action is not being performed correctly.

3. The system of claim 1, wherein the augmented image is generated in response to determining the next action among the at least one action included in the task.

4. The system of claim 1, further comprising generating a haptic alert in response to determining the next action.

5. A method for recommending guidance instructions to a user, the method comprising: monitoring an ongoing task comprising at least one action performed by a user, the at least one action being an action among a plurality of actions to be performed by the user in progressing to correctly complete the ongoing task; generating image data depicting the ongoing task; displaying a current action of the ongoing task being performed by the user based on the image data; analyzing the ongoing task and predicting a next action in the progression; generating an augmented image depicting at least a portion of the ongoing task, the augmented image being an image depicting correct performance of the predicted next action; and superimposing the augmented image on the image data such that the current action and the predicted next action are displayed simultaneously to guide the user in advancing the ongoing task.

6. The method of claim 5, further comprising generating a haptic alert in response to detecting that the at least one action is not being performed correctly.

7. The method of claim 5, further comprising: generating the augmented image in response to determining the next action among the at least one action included in the task.

8. The method of claim 5, further comprising generating a haptic alert in response to determining the next action.

9. A computer program product for recommending guidance instructions to a user, the computer program product comprising: ​ program instructions readable by a processing circuit to cause the processing circuit to perform operations of the method of any of claims 5-8.

Citation Information

Patent Citations

  • Information processing device, information processing method, and program

    CN102542247A

  • Automobile skylight control method and electronic equipment

    CN110936797A

  • Virtual or augmented reality rehabilitation

    US20180315247A1