Dynamic adjustment of commands in an augmented reality experience

The system addresses real-time error detection and correction in augmented reality by using image processing and machine-learned models to identify and rectify errors, enhancing user interaction and accuracy in text-based tasks.

JP2025538934APending Publication Date: 2025-12-03GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025522692
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-08-29
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Existing systems struggle with real-time error detection and correction in augmented reality environments, particularly in text-based tasks, leading to confusion and propagation of errors due to complex and difficult responses.

Method used

A computing system utilizing image processing, semantic understanding, and machine-learned models to detect errors in images and provide real-time user interface elements for correction, including optical character recognition and natural language processing to identify and rectify errors in text.

Benefits of technology

Enables real-time error detection and correction in augmented reality environments, providing interactive guidance for users to resolve errors efficiently and accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025538934000001_ABST
    Figure 2025538934000001_ABST
Patent Text Reader

Abstract

Systems and methods for augmented reality tutoring may utilize optical character recognition, natural language processing, and / or augmented reality rendering to provide real-time notifications for completing a determined task. The systems and methods may include utilizing one or more machine-learned models trained for quantitative reasoning and may include providing multiple different user interface elements at different times.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority claims This application is based on and claims priority to U.S. Application No. 17 / 969,303, filed October 19, 2022, which is incorporated by reference in its entirety.

[0002] The present disclosure relates generally to error detection and corrective action communication through an augmented reality experience. More specifically, the present disclosure relates to acquiring image data, processing the image data to determine that an error exists in the image data, determining a corrective action, and providing one or more user interface elements for display to indicate the corrective action. [Background technology]

[0003] Determining errors in an environment and figuring out how to correct the determined errors can be difficult. Errors in text can be particularly difficult to detect. Furthermore, if an error is not detected, it can lead to the propagation of further errors, which leads to further confusion. The lack of real-time error detection can cause users to spend time on a problem without understanding when and where they went wrong.

[0004] Furthermore, some errors and / or problems may involve difficult and complex responses to resolve the problem or error. Such difficult and complex responses may be confusing to users, and attempts to resolve the error may result in further errors. Summary of the Invention

[0005] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the description that follows, or may be learned from the description, or may be learned by practice of the embodiments.

[0006] One exemplary aspect of the present disclosure is directed to a computing system. The system may include one or more processors and one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations may include acquiring image data. The image data may describe one or more images. In some implementations, the one or more images may describe an environment. The operations may include processing the image data to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of the one or more images. The operations may include determining errors in the one or more images based at least in part on the semantic data. The operations may include determining corrective actions based on the semantic data and the errors. In some implementations, the corrective actions may describe at least one of replacing the errors or an action to correct the errors. The operations may include providing a user interface element for display based on the corrective action. The user interface element may include informational data describing the corrective action.

[0007] In some embodiments, determining errors in one or more images based at least in part on the semantic data may include obtaining a specific machine-learned model based on the semantic data and processing the image data with the specific machine-learned model to detect errors. The errors may include a mismatch with semantic understanding. In some embodiments, the errors may include deviations from a multi-part process. The multi-part process may be associated with the semantic data. In some embodiments, determining corrective actions based on the semantic data and the errors may include detecting a location of the errors in an environment, determining an errorless dataset associated with the semantic data and the one or more images, and determining replacement data from the errorless dataset based on the location of the errors in the environment.

[0008] In some embodiments, the error may be determined using an error detection model. The error detection model may generate text data based on optical character recognition, parse the text data based on one or more features in the environment, and process each parsed segment of the plurality of parsed segments to determine the error. In some embodiments, the error detection model may be trained with a plurality of mathematical proofs. The error detection model may include an optical character recognition model and a natural language processing model.

[0009] In some implementations, the image data may be generated by one or more image sensors of the mobile computing device. The user interface elements may be provided for display by the mobile computing device. The mobile computing device may be a smart wearable.

[0010] Another exemplary aspect of the present disclosure is directed to a computer-implemented method. The method may include acquiring, by a computing system including one or more processors, image data. The image data may describe one or more images. In some implementations, the one or more images may describe one or more pages. The method may include processing, by the computing system, the image data with an optical character recognition model to generate text data. In some implementations, the text data may describe text on one or more pages. The method may include determining, by the computing system, a prompt based on the text data. The prompt may describe a request for a response. The method may include determining, by the computing system, a multipart response to the prompt. The multipart response may include multiple individual responses associated with the prompt. The method may include acquiring, by the computing system, additional image data. The additional image data may describe one or more additional images. In some implementations, the one or more additional images may describe one or more pages having user-generated text. The method may include processing, by the computing system, the additional image data with an optical character recognition model to generate the additional text data. The additional text data may describe the user-generated text on the one or more pages. The method may include determining, by a computing system, that the user-generated text deviates from the multipart response and providing, by the computing system, a notification. The notification may describe the user-generated text having the error.

[0011] In some embodiments, determining, by the computing system, that the user-generated text deviates from the multi-part response may include determining that the user-generated text is inconsistent with the multi-part response. Determining, by the computing system, that the user-generated text deviates from the multi-part response may include determining that the user-generated text lacks one or more particular characteristics of the multi-part response. In some embodiments, the one or more pages may include one or more questions. The user-generated text may include user responses to the one or more questions. The method may include, by the computing system, processing the image data with a machine-learned model to determine a prompt and a multi-part response. In some embodiments, the method may include, by the computing system, processing additional image data with a machine-learned model to determine that the user-generated text deviates from the multi-part response.

[0012] Another exemplary aspect of the present disclosure is directed to one or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations may include obtaining image data. The image data may describe one or more images. In some implementations, the one or more images may describe one or more pages. The one or more pages may include a plurality of characters. The operations may include processing the image data to generate semantic data. The semantic data may describe a semantic understanding of at least some of the plurality of characters. The operations may include determining that the plurality of characters contain an error based at least in part on the semantic data. The error may describe text that is at least one of the opposite of the semantic understanding or an inaccuracy. The operations may include determining a corrective action based on the semantic data and the error. In some implementations, the corrective action may describe at least one of replacing the error or an action to correct the error. The operations may include providing a user interface element for display based on the corrective action. The user interface element may include informational data describing the corrective action.

[0013] In some implementations, the user interface element may include one or more popup elements that describe multiple sub-actions for performing corrective actions. The user interface element may include an inline overlay. The inline overlay may be utilized to enhance at least one of the one or more images or the one or more additional images to generate one or more enhanced images. In some implementations, the one or more enhanced images may include an inline overlay superimposed on at least a portion of the one or more pages. The inline overlay may describe the corrective action. In some implementations, the user interface element may include enhancing one or more of the images to indicate the location of the error.

[0014] Other aspects of the present disclosure are directed to various systems, apparatus, non-transitory computer-readable media, user interfaces, and electronic devices.

[0015] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the detailed description, serve to explain associated principles.

[0016] Detailed descriptions of embodiments directed to those skilled in the art are set forth herein with reference to the accompanying drawings. [Brief explanation of the drawings]

[0017] [Figure 1A] 1 illustrates a block diagram of an exemplary computing system for performing augmented reality tutoring, according to an exemplary embodiment of the present disclosure. [Figure 1B] 1 illustrates a block diagram of an exemplary computing device for performing augmented reality tutoring, according to an exemplary embodiment of the present disclosure. [Figure 1C] 1 illustrates a block diagram of an exemplary computing device for performing augmented reality tutoring, according to an exemplary embodiment of the present disclosure. [Figure 2A] 1 illustrates an example of an exemplary augmented reality experience, according to an exemplary embodiment of the present disclosure. [Figure 2B] 1 illustrates an example of an exemplary augmented reality experience, according to an exemplary embodiment of the present disclosure. [Figure 2C] 1 illustrates an example of an exemplary augmented reality experience, according to an exemplary embodiment of the present disclosure. [Figure 2D] 1 illustrates an example of an exemplary augmented reality experience, according to an exemplary embodiment of the present disclosure. [Figure 2E] 1 illustrates an example of an exemplary augmented reality experience, according to an exemplary embodiment of the present disclosure. [Figure 3] 1 illustrates a block diagram of an exemplary computing system for performing augmented reality tutoring, according to an exemplary embodiment of the present disclosure. [Figure 4] 1 illustrates a block diagram of an exemplary augmented reality tutoring system, according to an exemplary embodiment of the present disclosure. [Figure 5] 1 illustrates a block diagram of an exemplary augmented reality tutoring system, according to an exemplary embodiment of the present disclosure. [Figure 6] FIG. 1 illustrates a flowchart diagram of an exemplary method for performing augmented reality tutoring, according to an exemplary embodiment of the present disclosure. [Figure 7] 1 illustrates a flowchart diagram of an exemplary method for performing image data processing and corrective action determination, according to an exemplary embodiment of the present disclosure. [Figure 8] 1 illustrates a flowchart diagram of an exemplary method for performing video data processing and corrective action determination, according to an exemplary embodiment of the present disclosure. [Figure 9] 1 illustrates an example of an exemplary smart wearable for acquiring image data and providing user interface elements, according to an exemplary embodiment of the present disclosure.

[0018] Reference numbers repeated among the drawings are intended to identify like features in the various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0019] overview Generally, the present disclosure is directed to augmented reality experiences that provide multiple augmented reality assets based on processed image data. In particular, the systems and methods disclosed herein can leverage image processing (e.g., optical character recognition and / or object recognition), semantic understanding, and / or one or more user interface elements (e.g., inline overlays, pop-up bubbles, and / or highlighting) to provide real-time instructions for correcting errors and / or resolving problems. In some implementations, the systems and methods can be utilized for augmented reality tutoring, do-it-yourself (DIY) projects, and / or error detection and correction.

[0020] For example, the systems and methods may include acquiring image data. The image data may describe one or more images. In some implementations, the one or more images may describe an environment. The systems and methods may include processing the image data to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of the one or more images. The systems and methods may include determining errors in the one or more images based at least in part on the semantic data. In some implementations, the systems and methods may include determining corrective actions based on the semantic data and the errors. The corrective actions may describe at least one of replacing the errors or an action to correct the errors. The systems and methods may include providing a user interface element for display based on the corrective action. The user interface element may include information data describing the corrective action.

[0021] Image data may be acquired. The image data may describe one or more images. In some implementations, the one or more images may describe an environment. The environment may include one or more problems. For example, the environment may include questions for the user to respond to. Alternatively and / or additionally, the environment may include objects for completing a DIY project. The image data may be generated by one or more image sensors of a mobile computing device (e.g., a smartphone). In some implementations, the mobile computing device may be a smart wearable (e.g., smart glasses).

[0022] The image data may be processed to generate semantic data. The semantic data may describe the semantic understanding of at least a portion of one or more images. In some implementations, the image data may be processed with a semantic understanding model. The semantic understanding model may include one or more machine-learned models. The semantic understanding model may include a natural language processing model (e.g., one or more large-scale language models trained with multiple examples). In some implementations, the semantic understanding model may include a machine-learned model trained to understand equations and / or other quantitative expressions (e.g., a language model trained for quantitative reasoning such as those discussed in Dyer et al., Minerva: Solving Quantitative Reasoning Problems with Language Models, GOOGLE AI BLOG (June 30, 2022), https: / / ai.googleblog.com / 2022 / 06 / minerva-solving-quantitative-reasoning.html). Additionally and / or alternatively, the image data may be processed with an optical character recognition model to generate text data, which may then be processed with the semantic understanding model. The semantic data may be based on text content, recognized objects, data structure, data layout, information structure, one or more diagrams, received additional input data, the context of the image capture, the type of image capture device, user profile data, and / or one or more other contexts. The semantic data may include one or more queries that summarize a problem (e.g., a question) at the focal point of one or more images.

[0023] The systems and methods may determine errors in one or more images based at least in part on the semantic data. The errors may include inconsistencies with semantic understanding. In some implementations, the errors may include deviations from a multi-part process. The multi-part process may be associated with the semantic data. For example, the multi-part process may include one or more actions to answer a question and / or solve a problem. The errors may be determined based on heuristics, acquired data, and / or the output of a machine-learned model. The errors may be determined based on handwritten text differing from the semantic intent of the printed text. For example, the semantic data may include the semantic intent of the printed text and the semantic understanding of the handwritten text. An error may be determined if the semantic understanding of the handwritten text does not relate to the semantic intent of the printed text.

[0024] In some implementations, determining errors in one or more images based at least in part on the semantic data may include obtaining a particular machine-learned model based on the semantic data and processing the image data with the particular machine-learned model to detect errors. For example, the semantic data may describe a particular problem type (e.g., a literature analysis problem type, a calculus problem, and / or an organic chemistry problem) and a problem-specific machine-learned model (e.g., a literature analysis model, a calculus model, and / or an organic chemistry model). Alternatively and / or additionally, a mathematical engine (e.g., a system of mathematical functions utilized to process problems utilizing one or more processors) may be obtained and utilized based on the semantic data.

[0025] Alternatively and / or additionally, determining a corrective action based on the semantic data and the error may include detecting a location of the error in the environment, determining an error-free data set associated with the semantic data and the one or more images, and determining replacement data from the error-free data set based on the location of the error in the environment.

[0026] In some embodiments, the error may be determined using an error detection model. The error detection model may generate the text data based on optical character recognition. The error detection model may parse the text data based on one or more features in the environment. In some embodiments, the error detection model may process each parsed segment of the plurality of parsed segments to determine the error. The error detection model may be trained with a plurality of mathematical proofs. Additionally and / or alternatively, the error detection model may include an optical character recognition model and a natural language processing model.

[0027] A corrective action may be determined based on the semantic data and the error. The corrective action may describe at least one of replacing the error or an action to correct the error. In some implementations, the corrective action may include indicating a location of the error within the environment and one or more actions to correctly respond to a prompt identified within the environment.

[0028] The system and method may provide a user interface element for display based on the corrective action. The user interface element may include informational data describing the corrective action. In some implementations, the user interface element may be provided for display by a mobile computing device. The user interface element may be provided through an augmented reality experience. The user interface element may include a prompt, an inline comment, a pop-up bubble, and / or one or more highlighting arrows.

[0029] Additionally and / or alternatively, the systems and methods may continuously process the image data to determine and correct user actions in real time via one or more user interface elements provided in response to the determined errors. For example, the systems and methods may include acquiring image data. The image data may describe one or more images. In some implementations, the one or more images may describe one or more pages. The image data may be processed with an optical character recognition model to generate text data. The text data may describe text on one or more pages. The systems and methods may include determining a prompt based on the text data. The prompt may describe a request for a response. The systems and methods may include determining a multipart response to the prompt. The multipart response may include multiple individual responses associated with the prompt. In some implementations, the systems and methods may include acquiring additional image data. The additional image data may describe one or more additional images. The one or more additional images may describe one or more pages having user-generated text (e.g., additional handwritten text and / or user-typed data (e.g., user-generated code and / or user-generated equations)). The additional image data may be processed with an optical character recognition model to generate additional text data. In some implementations, the additional text data may describe user-generated text on one or more pages. The systems and methods may include determining that the user-generated text deviates from the multipart response and providing a notification. The notification may describe the user-generated text having the error.

[0030] The systems and methods may acquire image data. The image data may describe one or more images. The one or more images may describe one or more pages. In some implementations, the one or more pages may include one or more questions. The one or more pages may include printed text and handwritten text. The one or more questions may include a science question that includes a mathematical formula, a writing prompt, and / or one or more diagrams.

[0031] The image data may be processed with an optical character recognition model to generate text data. The text data may describe text on one or more pages. The optical character recognition model may include one or more machine-learned models. The optical character recognition model may include a model trained specifically on handwritten text. The text data may include recognized printed text and / or recognized handwritten text.

[0032] The systems and methods may determine prompts based on text data. The prompts may describe a request for a response. The prompts may be determined based on a semantic understanding of the text on one or more pages. Alternatively and / or additionally, the prompts may be queries generated based on recognized text. The prompts may be determined based on text containing one or more keywords associated with one or more prompts and / or one or more prompt types.

[0033] The system and method may determine a multipart response to a prompt. The multipart response may include multiple individual responses associated with the prompt. The multipart response may be determined based on the output of a machine-learned model, based on heuristics, based on one or more search results received from a search engine and / or one or more knowledge graphs. The multipart response may be based on the output of a machine-learned model trained on one or more textbooks. For example, a machine-learned model may be trained to identify a particular type of problem based on one or more identified features, and the same or a separate model may be trained to generate a proof that describes how to solve the particular problem. The generated proof may be a multipart response, with each line of the proof being part of the response.

[0034] Additional image data may be acquired. The additional image data may describe one or more additional images. The one or more additional images may describe one or more pages having user-generated text (e.g., additional handwritten text and / or user-typed data (e.g., user-generated code and / or user-generated equations)). The user-generated text may include user responses to one or more questions.

[0035] The additional image data may be processed with an optical character recognition model to generate additional text data. The additional text data may describe user-generated text on one or more pages and / or on a computer screen. The user-generated text may describe a user's attempt to answer a prompt (e.g., answer a question).

[0036] The user-generated text may be determined to deviate from a multi-part response. The deviation may be a deviation from a multi-part response such that the user-generated text is the opposite of the multi-part response. For example, the multi-part response may include taking a first action and then a second action, and the user-generated text may include taking the first action and then a third action that is not equivalent to the second action.

[0037] In some implementations, determining that the user-generated text deviates from the multi-part response may include determining that the user-generated text is inconsistent with the multi-part response, e.g., the user-generated text includes a semantic intent that is inconsistent with the semantic intent of one or more parts of the multi-part response.

[0038] Alternatively and / or additionally, determining that the user-generated text deviates from a multi-part response may include determining that the user-generated text lacks one or more particular characteristics of a multi-part response. For example, a multi-part response may include multiplying both sides of an equation by 2, while the user-generated text multiplies only one side by 2.

[0039] The system and method may provide a notification. The notification may describe user-generated text containing the error. The notification may be provided by an augmented reality experience that renders one or more user interface elements to provide the notification. The notification may describe where the error occurred and how to resolve the error.

[0040] In some implementations, the systems and methods may process the image data with a machine-learned model to determine prompts and multi-part responses. The machine-learned model may be a language model trained in quantitative reasoning. In some implementations, the machine-learned model may be trained on one or more subjects, particularly using academic materials (e.g., textbooks and / or academic papers).

[0041] In some implementations, the systems and methods may process the additional image data with a machine-learned model to determine that the user-generated text deviates from a multi-part response.

[0042] The systems and methods may include acquiring image data. The image data may describe one or more images. In some implementations, the one or more images may describe one or more pages. The one or more pages may include a plurality of characters. The image data may be processed to generate semantic data. The semantic data may describe a semantic understanding of at least some of the plurality of characters. The systems and methods may include determining, based at least in part on the semantic data, that the plurality of characters contain an error. The error may describe text that is at least one of a contradiction to the semantic understanding or an inaccuracy. In some implementations, the systems and methods may include determining a corrective action based on the semantic data and the error. The corrective action may describe at least one of a replacement of the error or an action to correct the error. The systems and methods may include providing a user interface element for display based on the corrective action. The user interface element may include information data describing the corrective action.

[0043] The systems and methods may acquire image data. The image data may describe one or more images. The one or more images may describe one or more pages. In some implementations, the one or more pages may include multiple characters. The multiple characters may be part of a problem (e.g., a question, a writing prompt, and / or a task statement). The characters may include letters, numbers, and / or symbols. The one or more pages may include text, pictures, shapes, diagrams, and / or white space.

[0044] The image data may be processed to generate semantic data. The semantic data may describe the meaning of at least a portion of the plurality of characters. In some implementations, the semantic data may be based on text, pictures, shapes, figures, and / or white space.

[0045] The systems and methods may determine that a plurality of characters contains an error based at least in part on the semantic data. The error may describe text that is at least one of the opposite of or inaccurate in semantic understanding. The inaccuracy may be determined by processing the plurality of characters with one or more machine-learned models.

[0046] The system and method may determine a correction action based on the semantic data and the error. The correction action may describe at least one of replacing the error or correcting the error. The correction action may include a delete action (e.g., deleting a subset of multiple characters) and a write action (e.g., writing one or more new characters).

[0047] The system and method may provide a user interface element for display based on the corrective action. The user interface element may include informational data describing the corrective action. In some implementations, the user interface element may include one or more pop-up elements describing multiple sub-actions for performing the corrective action. The user interface element may include an inline overlay. The inline overlay may be utilized to enhance at least one of the one or more images or the one or more additional images to generate one or more enhanced images. The one or more enhanced images may include an inline overlay superimposed on at least a portion of the one or more pages. In some implementations, the inline overlay may describe the corrective action. Additionally and / or alternatively, the user interface element may include enhancing one or more of the images to indicate the location of the error.

[0048] In some embodiments, the systems and methods may include continuous capture of image data, continuous image data processing, continuous error diagnosis, and / or continuous generation of user interface elements for correcting errors. For example, the systems and methods may utilize streaming optical character recognition. User interface elements (e.g., user interface elements for notifications) may be provided in a conversational format so that multi-part responses can be provided incrementally as the user progresses through different parts of the problem resolution.

[0049] In some implementations, one or more machine-learned models may be trained on one or more textbooks, real-world flash cards (e.g., flash cards of foreign languages ​​or scientific names of structures, elements, or compounds), architectural drawings, knowledge graphs, and / or study guides. Additionally and / or alternatively, one or more machine-learned models may be trained on proofs. Training may include training to match the output of a rule engine. Alternatively and / or additionally, training may include black-box optimization. The systems and methods may be utilized for writing tasks, language learning tasks, mathematical problem-solving tasks (e.g., algebra, calculus, and / or discrete mathematics), scientific problem-solving tasks (e.g., physics, organic chemistry, biology, and / or chemistry), architectural design (e.g., tracing lines and keeping track of measurements), and / or surgical procedures. The systems and methods may determine that a prompt includes multiple required criteria, and the user's responses may be processed to determine which criteria are met and which are not.

[0050] The systems and methods disclosed herein may be initiated based on one or more inputs (eg, a voice command, a button selection, and / or a visual cue).

[0051] In some implementations, the systems and methods may include parsing one or more images and / or parsing text data generated by optical character recognition. Parsing may be based on lines, paragraphs, syntax, page structure, image regions, and / or one or more other features.

[0052] In some embodiments, one or more teachers may be provided with a software development kit to tailor the augmented reality tutor to their particular class, course curriculum, and / or teaching style.

[0053] The systems and methods of the present disclosure provide several technical effects and advantages. As one example, the systems and methods can provide an augmented reality tutoring experience. In particular, the systems and methods disclosed herein can leverage optical character recognition, natural language processing, and augmented reality rendering to provide an interactive experience for identifying errors and providing multi-part responses.

[0054] Another technical advantage of the disclosed systems and methods is their ability to leverage one or more machine-learned models to understand an environment and provide a step-by-step process for completing a task. For example, the systems and methods can determine the meaning of the environment, determine a prompt associated with the environment, determine a multi-part response associated with the prompt, and continuously collect data to ensure the user completes the action associated with the multi-part response.

[0055] Other examples of technical effects and benefits relate to improved computational efficiency and improved functionality of computing systems. For example, the systems and methods disclosed herein can leverage memory of determined prompts and multi-part responses to continuously compare additional acquired data with the multi-part responses without having to continuously redetermine the meaning of the environment for error detection.

[0056] Referring now to the drawings, exemplary embodiments of the present disclosure will be described in more detail.

[0057] Exemplary Devices and Systems 1A illustrates a block diagram of an exemplary computing system 100 for performing dynamic adjustment of instructions in an augmented reality experience, according to an exemplary embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled via a network 180.

[0058] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0059] The computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operatively connected processors. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0060] In some implementations, the user computing device 102 can store or include one or more semantic understanding models 120. For example, the semantic understanding models 120 can be or otherwise include various machine-learned models, such as neural networks (e.g., deep neural networks), or other types of machine-learned models, including nonlinear and / or linear models. The neural networks can include feedforward neural networks, recurrent neural networks (e.g., long-short-term memory recurrent neural networks), convolutional neural networks, or other types of neural networks. Exemplary semantic understanding models 120 are described with reference to FIGS. 2A-5.

[0061] In some implementations, one or more semantic understanding models 120 may be received from server computing system 130 over network 180, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing device 102 may implement multiple parallel instances of a single semantic understanding model 120 (e.g., to perform parallel image data processing across multiple instances of images of an environment).

[0062] More specifically, the systems and methods may utilize one or more machine-learned models, including one or more semantic understanding models 120, that may process image data, text data, and / or audio data to generate semantic data associated with the environment. The semantic data may be utilized to detect errors in the environment and generate corrective actions to remedy the errors.

[0063] Additionally or alternatively, one or more semantic understanding models 140 may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the semantic understanding model 140 may be implemented by the server computing system 130 as part of a web service (e.g., a tutoring service). Thus, one or more models 120 may be stored and implemented at the user computing device 102 and / or one or more models 140 may be stored and implemented at the server computing system 130.

[0064] The user computing device 102 may include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may function to implement a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which a user can provide user input.

[0065] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operatively connected processors. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0066] In some implementations, server computing system 130 includes or is otherwise implemented by one or more server computing devices. If server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0067] As described above, the server computing system 130 may store or otherwise include one or more machine-learned semantic understanding models 140. For example, the models 140 may be or otherwise include various machine-learned models. Exemplary machine-learned models include neural networks or other multi-layer nonlinear models. Exemplary neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Exemplary models 140 are described with reference to FIGS. 2A-5.

[0068] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 by interacting with a training computing system 150 that is communicatively coupled via a network 180. The training computing system 150 can be separate from the server computing system 130 or can be part of the server computing system 130.

[0069] Training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operably connected processors. Memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 154 may store data 156 and instructions 158 that are executed by processor 152 to cause training computing system 150 to perform operations. In some implementations, training computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0070] The training computing system 150 may include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored on the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as backpropagation. For example, a loss function may be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent may be used to iteratively update the parameters over several training iterations.

[0071] In some implementations, performing backpropagation may include performing truncated backpropagation through time. The model trainer 160 may perform several generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0072] In particular, model trainer 160 can train semantic understanding models 120 and / or 140 based on a set of training data 162. Training data 162 can include, for example, textbooks, flash cards, academic papers, equations, natural language, books, proofs, and / or homework keys.

[0073] In some implementations, if the user provides consent, training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 with user-specific data received from the user computing device 102. In some examples, this process may be referred to as personalizing the model.

[0074] Model trainer 160 includes computer logic utilized to provide desired functionality. Model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general-purpose processor. For example, in some embodiments, model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other embodiments, model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium, such as RAM, a hard disk, or an optical or magnetic medium.

[0075] Network 180 can be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communications over network 180 can be transmitted over any type of wired and / or wireless connection using a wide variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, Secure HTTP, SSL).

[0076] The machine-learned models described herein may be used in a variety of tasks, applications, and / or use cases.

[0077] In some implementations, the input to the machine-learned model(s) of the present disclosure can be image data. The machine-learned model(s) can process the image data to generate an output. As an example, the machine-learned model(s) can process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an image segmentation output. As another example, the machine-learned model(s) can process the image data to generate an image classification output. As another example, the machine-learned model(s) can process the image data to generate an image data modification output (e.g., a modification of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an upscaled image data output. As another example, the machine-learned model(s) can process the image data to generate a predicted output.

[0078] In some implementations, input to the machine-learned model(s) of the present disclosure can be text or natural language data. The machine-learned model(s) can process the text or natural language data to generate an output. As an example, the machine-learned model(s) can process the natural language data to generate a language-encoded output. As another example, the machine-learned model(s) can process the text or natural language data to generate a latent text embedding output. As another example, the machine-learned model(s) can process the text or natural language data to generate a translation output. As another example, the machine-learned model(s) can process the text or natural language data to generate a classification output. As another example, the machine-learned model(s) can process the text or natural language data to generate a text segmentation output. As another example, the machine-learned model(s) can process the text or natural language data to generate a semantic intent output. As another example, the machine-learned model(s) may process text or natural language data to generate upscaled text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language). As another example, the machine-learned model(s) may process text or natural language data to generate predicted outputs.

[0079] In some implementations, input to the machine-learned model(s) of the present disclosure can be speech data. The machine-learned model(s) can process the speech data to generate an output. As an example, the machine-learned model(s) can process the speech data to generate a speech recognition output. As another example, the machine-learned model(s) can process the speech data to generate a speech translation output. As another example, the machine-learned model(s) can process the speech data to generate a latent embedding output. As another example, the machine-learned model(s) can process the speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of the speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate an upscaled speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, machine-learned model(s) can process speech data to generate predicted outputs.

[0080] In some implementations, input to the machine-learned model(s) of the present disclosure can be latent-coded data (e.g., a latent space representation of the input, etc.). The machine-learned model(s) can process the latent-coded data to generate an output. As an example, the machine-learned model(s) can process the latent-coded data to generate a recognition output. As another example, the machine-learned model(s) can process the latent-coded data to generate a reconstruction output. As another example, the machine-learned model(s) can process the latent-coded data to generate a retrieval output. As another example, the machine-learned model(s) can process the latent-coded data to generate a reclustering output. As another example, the machine-learned model(s) can process the latent-coded data to generate a prediction output.

[0081] In some implementations, input to the machine-learned model(s) of the present disclosure can be statistical data. The machine-learned model(s) can process the statistical data to generate an output. As an example, the machine-learned model(s) can process the statistical data to generate a recognition output. As another example, the machine-learned model(s) can process the statistical data to generate a prediction output. As another example, the machine-learned model(s) can process the statistical data to generate a classification output. As another example, the machine-learned model(s) can process the statistical data to generate a segmentation output. As another example, the machine-learned model(s) can process the statistical data to generate a segmentation output. As another example, the machine-learned model(s) can process the statistical data to generate a visualization output. As another example, the machine-learned model(s) can process the statistical data to generate a diagnostic output.

[0082] In some implementations, input to the machine-learned model(s) of the present disclosure can be sensor data. The machine-learned model(s) can process the sensor data to generate an output. As an example, the machine-learned model(s) can process the sensor data to generate a recognition output. As another example, the machine-learned model(s) can process the sensor data to generate a prediction output. As another example, the machine-learned model(s) can process the sensor data to generate a classification output. As another example, the machine-learned model(s) can process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) can process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) can process the sensor data to generate a visualization output. As another example, the machine-learned model(s) can process the sensor data to generate a diagnostic output. As another example, the machine-learned model(s) can process the sensor data to generate a detection output.

[0083] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data of one or more images and the task is an image processing task. For example, the image processing task can be image classification, and the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images depict an object belonging to the object class. The image processing task can be object detection, and the image processing output identifies one or more regions in one or more images and, for each region, the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, and the image processing output defines, for each pixel in one or more images, a respective likelihood for each category in a predetermined category set. For example, the category set can be foreground and background. As another example, the category set can be object classes. As another example, the image processing task can be depth estimation, and the image processing output defines, for each pixel in one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel of one of the input images, the motion of the scene depicted in pixels between the images in the network input.

[0084] In some cases, the input includes audio data representing a speech utterance and the task is a speech recognition task. The output may include a text output that is mapped to the speech utterance.

[0085] 1A illustrates one exemplary computing system that may be used to implement the present disclosure. Other computing systems may be used as well. For example, in some implementations, user computing device 102 may include model trainer 160 and training dataset 162. In such implementations, model 120 may be trained and used locally on user computing device 102. In some such implementations, user computing device 102 may implement model trainer 160 that personalizes model 120 based on user-specific data.

[0086] 1B illustrates a block diagram of an exemplary computing device 10 for performing operations according to an exemplary embodiment of the present disclosure. The computing device 10 can be a user computing device or a server computing device.

[0087] Computing device 10 includes several applications (e.g., applications 1-N). Each application includes its own machine learning library and machine-learned model(s). For example, each application may include a machine-learned model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0088] 1B , each application may communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application may communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0089] 1C illustrates a block diagram of an exemplary computing device 50 for performing operations according to an exemplary embodiment of the present disclosure. The computing device 50 can be a user computing device or a server computing device.

[0090] Computing device 50 includes several applications (e.g., applications 1-N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the model(s) stored therein) using an API (e.g., a common API across all applications).

[0091] The central intelligence layer includes several machine-learned models. For example, as shown in FIG. 1C , each machine-learned model (e.g., model) can be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications can share a single machine-learned model. For example, in some embodiments, the central intelligence layer can provide a single model (e.g., single model) for all of the applications. In some embodiments, the central intelligence layer is included within or otherwise implemented by the operating system of computing device 50.

[0092] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for computing device 50. As shown in FIG. 1C , the central device data layer can communicate with several other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0093] FIG. 3 illustrates a block diagram of an exemplary computing system 300 for performing augmented reality tutoring, according to an exemplary embodiment of the present disclosure. For example, the systems and methods disclosed herein may include one or more computing devices communicatively coupled via a network 302. In some implementations, the computing system 300 may include one or more image sensors, one or more visual displays, one or more audio sensors, one or more audio output components, one or more storage devices, and / or one or more processors. The computing system 300 may include a smart device 304 (e.g., a smartphone) and / or a smart wearable 306 (e.g., smart glasses). In some implementations, the smart device 304 and the smart wearable 306 may be communicatively coupled via the network 302, a Bluetooth connection, and / or other communication medium. For example, the display, sensors, and / or processor of the smart device 304 may be utilized with the smart wearable 306. In some implementations, the smart device 304 and / or smart wearable 306 may exchange data with one or more server computing systems 308 to implement the systems and methods disclosed herein.

[0094] Exemplary System Configuration and Exemplary Model Configuration 2A-2E illustrate examples of exemplary augmented reality experiences according to exemplary embodiments of the present disclosure. The exemplary augmented reality experiences may be initiated based on one or more inputs (e.g., squeezing a button, touch input to a touchscreen, detecting or classifying an object, and / or audio input). In response to the augmented reality experience being initiated, one or more images 202 may be captured. In some implementations, audio input may be captured.

[0095] One or more images 202 may be processed to determine a prompt 204 at a focal point of the environment (e.g., as shown in FIG. 2A ). The focal point may be determined based on a location within the environment, based on a user's instruction, based on a detected object, based on a user's gaze, and / or based on one or more machine-learned parameters. The prompt 204 may be highlighted, underlined, and / or indicated by one or more other techniques.

[0096] The audio input may be processed to determine speech, which may be provided for display by closed captioning user interface element 206. The augmented reality experience may provide user interface elements to indicate that data is being acquired via image and / or audio sensors.

[0097] The prompts 204 may be processed to determine a multi-part response for a particular prompt 204. The multi-part response may be determined based on one or more machine-learned parameters, based on a knowledge graph, and / or based on data retrieved from a database. Multiple user interface elements may be generated based on a set of actions associated with the multi-part response.

[0098] 2B, prompt 204 remains highlighted and first user interface element 208 is provided for display. First user interface element 208 may describe instructions for completing one or more actions on the first part of the multi-part response.

[0099] Additional image data may then be received. The additional image data may be processed to identify new text data. The new text data may include handwritten text 210. The new text data may be processed to determine that the first portion of the multi-part response is complete. In response to the determination, a second user interface element 214 may be provided for display (e.g., as shown in FIG. 2C ). The second user interface element 214 may describe instructions for completing one or more actions on the second portion of the multi-part response. In some implementations, the augmented reality experience may provide a location of interest indicator 212, which may indicate a location of interest for the set of actions and / or may indicate a prompt type associated with the prompt 204.

[0100] 2D, additional handwritten text 216 has been provided and is being provided. The systems and methods may determine that a second portion of the multi-part response has been detected, and a third user interface element 218 may be provided for display. The third user interface element 218 may describe instructions for completing one or more actions on the third portion of the multi-part response.

[0101] 2E, the multi-part response has been executed and a final handwritten response has been identified 220. The augmented reality experience may provide a completion indicator 222 and / or one or more final user interface elements 224 that indicate that the set of actions has been completed.

[0102] The system and method may then be repeated for the next identified prompt. The augmented reality experience may include additional user interface elements to provide intuitive instructions. The augmented reality experience may be provided with one or more audio outputs.

[0103] 4 illustrates a block diagram of an exemplary augmented reality tutoring system 400, according to an exemplary embodiment of the present disclosure. In some implementations, the augmented reality tutoring system 400 can receive image data 402 and / or additional input data 404 describing an environment and / or a particular prompt, and can provide a user interface element 420 describing one or more classifications and / or instructions for completing a task as a result of receiving the image data 402 and the additional input data 404. Thus, in some implementations, the augmented reality tutoring system 400 can include an optical character recognition model 406 operable to recognize text in an image and a semantic understanding model 410 operable to generate a semantic output 412.

[0104] For example, image data 402 may be acquired by one or more image sensors. Image data 402 may describe one or more images depicting an environment. The environment may include a plurality of characters describing one or more prompts (e.g., one or more questions and / or one or more instructions).

[0105] Image data 402 may be processed by an optical character recognition model 406 to recognize one or more characters, one or more symbols, and / or one or more figures and generate text data 408. Text data 408 may include words, numbers, equations, figures, text structure, text layout, syntax, and / or symbols. Optical character recognition model 406 may be trained on printed text and / or may be trained specifically to determine handwritten characters.

[0106] In some implementations, additional input data 404 may be acquired. The additional input data 404 may be acquired by one or more additional sensors, which may include audio sensors and / or touch sensors. The additional input data 404 may be generated based on voice utterances and / or one or more selections made in a user interface.

[0107] The image data 402 and / or additional input data 404 may be processed by a semantic understanding model 410 to generate a semantic output 412. The semantic understanding model 410 may include one or more segmentation models, one or more augmentation models, one or more natural language processing models, one or more quantitative reasoning models, and / or one or more classification models. The semantic understanding model 410 may include one or more Transformer models, one or more convolutional neural networks, one or more genetic algorithm neural networks, one or more classifier models, and / or recurrent neural networks. The semantic understanding model 410 may be trained with a large-scale language training dataset, a quantitative reasoning training dataset, a textbook dataset, a flashcard training dataset, and / or a proof dataset. The semantic understanding model 410 may be trained to determine the semantic intent of the input data and perform one or more tasks based on the semantic intent. For example, the semantic understanding model 410 may be trained for multiple tasks, which may include input summarization, response tasks, completion tasks, diagnostic tasks, problem-solving tasks, and error detection tasks, classification tasks, and / or enhancement tasks.

[0108] Based on the semantic output 412, no further action 414 may be determined, which may lead to the process starting again. Alternatively and / or additionally, the augmented reality tutor interface 416 may be initiated based on the semantic output 412. For example, the augmented reality tutor interface 416 may be initiated based on the semantic output 412 describing an error (e.g., an inaccuracy in a response and / or a configuration problem in the environment). Alternatively and / or additionally, the augmented reality tutor interface 416 may be initiated based on the semantic output 412 describing a threshold time that occurs without an action occurring (e.g., a threshold time that occurs without new handwriting). In some implementations, the augmented reality tutor interface 416 may be initiated by a user input that triggers the acquisition of image data 402 and / or additional input data 404.

[0109] The information output 418 may be determined based on the semantic output 412. For example, the semantic output 412 may describe an error in the environment, and the information output may include instructions for a set of actions to correct the error (e.g., corrective actions). The set of actions may be determined based on machine-learned model outputs, based on one or more knowledge graphs, and / or based on one or more search results generated by querying a database. The instructions for the set of actions may be ordered sequentially.

[0110] One or more user interface elements 420 may be generated based on the information output 418. In some implementations, a user interface element 420 may be generated for each action in the set of actions. The user interface elements 420 may then be provided to the user. The user interface elements 420 may be provided sequentially. The process may then begin again.

[0111] 5 illustrates a block diagram of an exemplary augmented reality tutoring system 500, according to an exemplary embodiment of the present disclosure. The augmented reality tutoring system 500 is similar to the augmented reality tutoring system 400 of FIG. 4, except that the augmented reality tutoring system 500 further includes an augmented reality generation block 518.

[0112] The augmented reality tutoring system 500 can acquire input data, which can include words 502, numbers 504, equations 506, diagrams 508, structural data 510, and / or other data 512. The other data 512 can include time data, audio data, touch data, and / or contextual data. The input data can include multimodal data and / or can be conditioned or supplemented based on profile data and / or preference data.

[0113] The input data may be processed by the semantic understanding model 514 to generate or determine a prompt. The prompt may be based on a semantic understanding of the input data, which may include a semantic understanding of the environment. The prompt may include a problem to solve (e.g., a math problem, a reading comprehension problem, and / or a science problem), a writing prompt (e.g., an analysis prompt, an essay prompt, and / or a creative writing prompt), and / or a DIY project (e.g., furniture assembly, appliance repair, and / or vehicle maintenance).

[0114] The prompts may be processed in a response determination block 516 to generate a response. In some implementations, the response determination block 516 may be part of the semantic understanding model 514. The semantic understanding model 514 and / or the response determination block 516 may include one or more machine-learned models. In some implementations, the response determination block 516 may include determining a query based on the prompt and querying a database (e.g., a search engine and / or an academic database).

[0115] The response may include a multi-part response that includes a set of actions, which may be part of a larger corrective action to remedy the error. The response may be processed by the augmented reality generation block 518 to generate a plurality of augmented reality user interface elements that are provided to the user. The augmented reality user interface elements may describe instructions for performing the plurality of actions associated with the response.

[0116] The augmented reality user interface elements may include inline renderings 520 (e.g., text and / or symbols provided inline with text and / or objects in the environment), pop-up elements 522 (e.g., speech bubbles rendered in the augmented reality display), highlight elements 524 (e.g., brightening pixels and / or darkening pixels displaying the environment), animation elements 526 (e.g., animated images and / or animated text that change over the course of a presentation period), symbols (e.g., representative indicators and / or classification symbols), and / or other user interface output 530 (e.g., three-dimensional augmented renderings of objects in the scene).

[0117] FIG. 9 illustrates an example of an exemplary smart wearable 900 for acquiring image data and providing user interface elements according to an exemplary embodiment of the present disclosure. For example, the systems and methods may be implemented by the smart wearable 900. In some implementations, the smart wearable 900 may include smart glasses. The smart wearable 900 may include one or more image sensors 902, one or more computer component shells 904, one or more displays 906, and / or one or more lenses 908. The lenses may be prescription lenses, blue light lenses, tinted lenses, and / or clear non-prescription lenses. The one or more image sensors 902 may be positioned such that acquired image data describes the environment within a user's field of view. The one or more computer component shells 904 may store one or more processors, one or more communication components (e.g., Bluetooth receivers, ultra-wideband receivers, and / or WiFi receivers), one or more audio components (e.g., microphones and / or speakers), and / or one or more storage devices. The one or more displays 906 may be configured to display one or more user interface elements.

[0118] For example, one or more image sensors 902 can generate image data, which can be processed by one or more processors in one or more computer component shells 904. User interface elements can be selected and / or generated based on the image data, whereupon one or more user interface elements can be provided for display via one or more displays 906.

[0119] Exemplary Methods 6 shows a flowchart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. While FIG. 6 shows steps performed in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particularly shown order or configuration. Various steps of method 600 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0120] At 602, the computing system may acquire image data. The image data may describe one or more images. In some implementations, the one or more images may describe an environment. The environment may include one or more problems. For example, the environment may include questions for a user to respond to. Alternatively and / or additionally, the environment may include objects for completing a DIY project. The image data may be generated by one or more image sensors of a mobile computing device (e.g., a smartphone). In some implementations, the mobile computing device may be a smart wearable (e.g., smart glasses).

[0121] At 604, the computing system can process the image data to generate semantic data. The semantic data can describe the semantic understanding of at least a portion of one or more images. In some implementations, the image data can be processed with a semantic understanding model. The semantic understanding model can include one or more machine-learned models. The semantic understanding model can include a natural language processing model (e.g., one or more large-scale language models trained with multiple examples). In some implementations, the semantic understanding model can include a machine-learned model trained to understand equations and / or other quantitative expressions (e.g., a language model trained for quantitative reasoning such as those discussed in Dyer et al., Minerva: Solving Quantitative Reasoning Problems with Language Models, GOOGLE AI BLOG (June 30, 2022), https: / / ai.googleblog.com / 2022 / 06 / minerva-solving-quantitative-reasoning.html). Additionally and / or alternatively, the image data can be processed with an optical character recognition model to generate text data, which can then be processed with the semantic understanding model. The semantic data may be based on text content, recognized objects, data structure, data layout, information structure, one or more diagrams, received additional input data, the context of the image capture, the type of image capture device, user profile data, and / or one or more other contexts. The semantic data may include one or more queries that summarize a problem (e.g., a question) at the focal point of one or more images.

[0122] At 606, the computing system may determine an error in one or more images based at least in part on the semantic data. The error may include a discrepancy with semantic understanding. In some implementations, the error may include a deviation from a multi-part process. The multi-part process may be associated with the semantic data. For example, the multi-part process may include one or more actions to answer a question and / or solve a problem. The error may be determined based on heuristics, based on acquired data, and / or based on the output of a machine-learned model. The error may be determined based on handwritten text differing from the semantic intent of the printed text. For example, the semantic data may include the semantic intent of the printed text and the semantic understanding of the handwritten text. An error may be determined if the semantic understanding of the handwritten text does not relate to the semantic intent of the printed text.

[0123] In some implementations, determining errors in one or more images based at least in part on the semantic data may include obtaining a particular machine-learned model based on the semantic data and processing the image data with the particular machine-learned model to detect errors. For example, the semantic data may describe a particular problem type (e.g., a literature analysis problem type, a calculus problem, and / or an organic chemistry problem) and a problem-specific machine-learned model (e.g., a literature analysis model, a calculus model, and / or an organic chemistry model). Alternatively and / or additionally, a mathematical engine (e.g., a system of mathematical functions utilized to process problems utilizing one or more processors) may be obtained and utilized based on the semantic data.

[0124] Alternatively and / or additionally, determining a corrective action based on the semantic data and the error may include detecting a location of the error in the environment, determining an error-free data set associated with the semantic data and the one or more images, and determining replacement data from the error-free data set based on the location of the error in the environment.

[0125] In some embodiments, the error may be determined using an error detection model. The error detection model may generate the text data based on optical character recognition. The error detection model may parse the text data based on one or more features in the environment. In some embodiments, the error detection model may process each parsed segment of the plurality of parsed segments to determine the error. The error detection model may be trained with a plurality of mathematical proofs. Additionally and / or alternatively, the error detection model may include an optical character recognition model and a natural language processing model.

[0126] At 608, the computing system can determine a corrective action based on the semantic data and the error. The corrective action can describe at least one of replacing the error or an action to correct the error. In some implementations, the corrective action can include indicating a location of the error within the environment and one or more actions to correctly respond to a prompt identified within the environment.

[0127] At 610, the computing system can provide a user interface element for display based on the corrective action. The user interface element can include information data describing the corrective action. In some implementations, the user interface element can be provided for display by a mobile computing device. The user interface element can be provided through an augmented reality experience. The user interface element can include a prompt, an inline comment, a pop-up bubble, and / or one or more highlighting arrows.

[0128] 7 shows a flowchart of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. While FIG. 7 shows steps performed in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particularly shown order or configuration. Various steps of method 700 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0129] At 702, the computing system can acquire image data. The one or more images can describe one or more pages. In some implementations, the one or more pages can include one or more questions. The one or more pages can include printed text and handwritten text. The one or more questions can include mathematical equations, writing prompts, and / or science questions that include one or more diagrams.

[0130] At 704, the computing system can process the image data with an optical character recognition model to generate text data. The text data can describe text on one or more pages. The optical character recognition model can include one or more machine-learned models. The optical character recognition model can include a model trained specifically on handwritten text. The text data can include recognized printed text and / or recognized handwritten text.

[0131] At 706, the computing system can determine a prompt based on the text data and a multi-part response to the prompt. The prompt can describe a request for a response. The prompt can be determined based on a semantic understanding of the text on one or more pages. Alternatively and / or additionally, the prompt can be a query generated based on recognized text. The prompt can be determined based on text including one or more keywords associated with one or more prompts and / or one or more prompt types.

[0132] A multi-part response may include multiple individual responses associated with the prompt. The multi-part response may be determined based on the output of a machine-learned model, based on heuristics, based on one or more search results received from a search engine and / or one or more knowledge graphs. The multi-part response may be based on the output of a machine-learned model trained on one or more textbooks. For example, a machine-learned model may be trained to identify a particular type of problem based on one or more identified features, and the same or a separate model may be trained to generate a proof that describes how to solve the particular problem. The generated proof may be a multi-part response, with each line of the proof being part of the response.

[0133] At 708, the computing system may obtain additional image data. The additional image data may describe one or more additional images. The one or more additional images may describe one or more pages having user-generated text (e.g., additional handwritten text and / or user-typed data (e.g., user-generated code and / or user-generated equations)). The user-generated text may include user responses to one or more questions.

[0134] At 710, the computing system can process the additional image data with an optical character recognition model to generate additional text data. The additional text data can describe user-generated text on one or more pages. The user-generated text can describe a user's attempt to answer the prompt (e.g., answer a question).

[0135] At 712, the computing system can determine that the user-generated text deviates from a multi-part response and provide a notification. The deviation can be a deviation from a multi-part response such that the user-generated text is the opposite of the multi-part response. For example, the multi-part response can include taking a first action and then a second action, and the user-generated text can include taking the first action and then a third action that is not equivalent to the second action.

[0136] In some implementations, determining that the user-generated text deviates from the multi-part response may include determining that the user-generated text is inconsistent with the multi-part response, e.g., the user-generated text includes a semantic intent that is inconsistent with the semantic intent of one or more parts of the multi-part response.

[0137] Alternatively and / or additionally, determining that the user-generated text deviates from a multi-part response may include determining that the user-generated text lacks one or more particular characteristics of a multi-part response. For example, a multi-part response may include multiplying both sides of an equation by 2, while the user-generated text multiplies only one side by 2.

[0138] The system and method may provide a notification. The notification may describe user-generated text containing the error. The notification may be provided by an augmented reality experience that renders one or more user interface elements to provide the notification. The notification may describe where the error occurred and how to resolve the error.

[0139] In some implementations, the systems and methods may process the image data with a machine-learned model to determine prompts and multi-part responses. The machine-learned model may be a language model trained in quantitative reasoning. In some implementations, the machine-learned model may be trained on one or more subjects, particularly using academic materials (e.g., textbooks and / or academic papers).

[0140] In some implementations, the systems and methods may process the additional image data with a machine-learned model to determine that the user-generated text deviates from a multi-part response.

[0141] 8 shows a flowchart diagram of an exemplary method for performing in accordance with an exemplary embodiment of the present disclosure. While FIG. 8 shows steps performed in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particularly shown order or configuration. Various steps of method 800 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0142] At 802, the computing system may acquire video data. The image data may describe one or more images. The one or more images may describe one or more pages. In some implementations, the one or more pages may include multiple characters. The multiple characters may be part of a problem (e.g., a question, a writing prompt, and / or a task statement). The characters may include letters, numbers, and / or symbols. The one or more pages may include text, pictures, shapes, diagrams, and / or white space.

[0143] At 804, the computing system can process the video data to generate recognition data.

[0144] At 806, the computing system can process the recognition data to generate semantic data. The semantic data can describe a meaning understanding of at least a portion of the plurality of characters. In some implementations, the semantic data can be based on text, pictures, shapes, drawings, and / or white space.

[0145] At 808, the computing system may determine, based at least in part on the semantic data, that the plurality of characters contains an error. The error may describe text that is at least one of an inaccuracy or a contradiction in semantic understanding. The inaccuracy may be determined by processing the plurality of characters with one or more machine-learned models.

[0146] At 810, the computing system can determine a correction action based on the semantic data and the error. The correction action can describe at least one of replacing the error or an action to fix the error. The correction action can include a delete action (e.g., deleting a subset of multiple characters) and a write action (e.g., writing one or more new characters).

[0147] At 812, the computing system can provide a user interface element for display based on the corrective action. The user interface element can include informational data describing the corrective action. In some implementations, the user interface element can include one or more popup elements describing multiple sub-actions for performing the corrective action. The user interface element can include an inline overlay. The inline overlay can be utilized to augment at least one of the one or more images or the one or more additional images to generate one or more augmented images. The one or more augmented images can include an inline overlay superimposed on at least a portion of the one or more pages. In some implementations, the inline overlay can describe the corrective action. Additionally and / or alternatively, the user interface element can include augmenting one or more of the images to indicate the location of the error.

[0148] Additional Disclosures The technology described herein refers to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionality among components. For example, the processes described herein can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0149] While the present subject matter has been described in detail with reference to various specific exemplary embodiments thereof, each example is provided for purposes of illustration and not as a limitation of the present disclosure. Those skilled in the art, upon understanding the foregoing, may readily create modifications, variations, and equivalents to such embodiments. Accordingly, the present disclosure does not exclude the inclusion of such modifications, variations, and / or additions to the present subject matter that would be readily apparent to those skilled in the art. For example, features illustrated or described as part of one embodiment may be used with other embodiments to create yet another embodiment. Accordingly, the present disclosure is intended to cover such modifications, variations, and equivalents.

Claims

1. 1. A computing system comprising: one or more processors; one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations including: acquiring image data, the image data describing one or more images, the one or more images describing an environment; processing the image data to generate semantic data, the semantic data describing a semantic understanding of at least a portion of the one or more images; determining errors in the one or more images based at least in part on the semantic data; determining a corrective action based on the semantic data and the error, the corrective action describing at least one of a replacement of the error or an action to correct the error; providing a user interface element for display based on the corrective action, the user interface element including information data describing the corrective action; a computing system including:

2. Determining the errors in the one or more images based at least in part on the semantic data includes: Obtaining a specific machine-learned model based on the semantic data; and processing the image data with the particular machine-learned model to detect the error; The system of claim 1 , comprising:

3. The system of claim 1 or 2, wherein the error comprises a mismatch with the semantic understanding.

4. The system of claim 1 , wherein the error comprises a deviation from a multi-part process, the multi-part process being associated with the semantic data.

5. Determining the corrective action based on the semantic data and the error includes: Detecting the location of the error within the environment; determining an errorless dataset associated with the semantic data and the one or more images; determining replacement data from the error-free data set based on the location of the error within the environment; The system of claim 1 , comprising:

6. The error is determined using an error detection model, the error detection model comprising: Generate text data based on optical character recognition; Parsing the text data based on one or more features in the environment; The system of claim 1 , further comprising: processing each parsed segment of a plurality of parsed segments to determine the error.

7. The system of claim 6 , wherein the error detection model is trained with multiple mathematical proofs.

8. The system of claim 6 , wherein the error detection models include an optical character recognition model and a natural language processing model.

9. 9. The system of claim 1, wherein the image data is generated by one or more image sensors of a mobile computing device, and the user interface elements are provided for display by the mobile computing device.

10. The system of claim 9 , wherein the mobile computing device is a smart wearable.

11. 1. A computer-implemented method comprising: acquiring, by a computing system including one or more processors, image data describing one or more images, the one or more images describing one or more pages; processing, by the computing system, the image data with an optical character recognition model to generate text data, the text data describing text on the one or more pages; determining, by the computing system, a prompt based on the text data, the prompt describing a request for a response; determining, by the computing system, a multi-part response to the prompt, the multi-part response including a plurality of individual responses associated with the prompt; acquiring, by the computing system, additional image data describing one or more additional images, the one or more additional images describing the one or more pages having user-generated text; processing, by the computing system, the additional image data with the optical character recognition model to generate additional text data, the additional text data describing the user-generated text on the one or more pages; determining, by the computing system, that the user-generated text deviates from the multipart response; providing, by the computing system, a notification, the notification describing the user-generated text having an error; and A computer-implemented method comprising:

12. Determining, by the computing system, that the user-generated text deviates from the multipart response includes: The method of claim 11 , comprising determining that the user-generated text is inconsistent with the multi-part response.

13. Determining, by the computing system, that the user-generated text deviates from the multipart response includes:

13. The method of claim 11 or 12, comprising determining that the user-generated text lacks one or more particular characteristics of the multi-part response.

14. 14. The method of claim 11, wherein the one or more pages include one or more questions, and the user-generated text includes user responses to the one or more questions.

15. 15. The method of claim 11, further comprising processing, by the computing system, the image data with a machine-learned model to determine the prompt and the multi-part response.

16. 16. The method of claim 11, further comprising processing, by the computing system, the additional image data with a machine-learned model to determine that the user-generated text deviates from the multi-part response.

17. One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: acquiring image data, the image data describing one or more images, the one or more images describing one or more pages, the one or more pages including a plurality of characters; processing the image data to generate semantic data, the semantic data describing a meaning of at least a portion of the plurality of characters; determining, based at least in part on the semantic data, that the plurality of characters contain an error, the error describing text that is at least one of a contradiction or an inaccuracy of the semantic understanding; determining a corrective action based on the semantic data and the error, the corrective action describing at least one of a replacement of the error or an action to correct the error; providing a user interface element for display based on the corrective action, the user interface element including information data describing the corrective action; 1. One or more non-transitory computer-readable media, including:

18. 20. The one or more non-transitory computer-readable media of claim 17, wherein the user interface elements include one or more pop-up elements that describe multiple sub-actions for performing the corrective action.

19. 19. The one or more non-transitory computer-readable media of claim 17 or 18, wherein the user interface element includes an inline overlay, the inline overlay being utilized to augment at least one of the one or more images or the one or more additional images to generate one or more augmented images, the one or more augmented images including the inline overlay superimposed on at least a portion of the one or more pages, the inline overlay describing the corrective action.

20. 20. The one or more non-transitory computer-readable media of claim 17, wherein the user interface element includes expanding one or more of the images to indicate a location of the error.

Citation Information

Patent Citations

  • Plant facility maintenance support device and maintenance support system

    JP2018181232A

  • Information processing apparatus, information processing system, and program

    JP2019061590A

  • Fiber sensor, light receiving fiber, and diffusion member

    JP2019061889A

  • Systems and methods for UX-based automated content evaluation and delivery

    US20190272775A1