Dynamic adjustment instructions in augmented reality experiences

By processing image data to generate semantic data, identify and correct errors in the environment, the complex and inefficient problems of error detection and correction in the prior art are solved, and real-time and accurate error handling is achieved.

CN120112968APending Publication Date: 2025-06-06GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073816.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-08-29
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and correct errors in the environment, especially in text, where error detection and correction are complex and prone to further propagation of errors.

Method used

By obtaining image data, processing image data to generate semantic data, determining errors in image data, and determining correction actions based on semantic data and errors, user interface elements are provided to display correction actions.

Benefits of technology

Real-time error detection and correction are realized, reducing the time spent by users on errors and improving the accuracy and efficiency of error handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112968A_ABST
    Figure CN120112968A_ABST
Patent Text Reader

Abstract

Systems and methods for augmented reality coaching may utilize optical character recognition, natural language processing, and / or augmented reality rendering to provide real-time notifications for completing a determined task. The systems and methods may include utilizing one or more machine learning models trained for quantitative reasoning, and may include providing a plurality of different user interface elements at different times.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority claim

[0002] This application is based on and claims priority from U.S. application No. 17 / 969,303, filed on October 19, 2022, which is incorporated herein by reference. Technical Field

[0003] The present disclosure generally relates to error detection and corrective action communication via augmented reality experience. More specifically, the present disclosure relates to obtaining image data, processing the image data to determine that an error exists in the image data, determining a corrective action, and providing one or more user interface elements for display to indicate the corrective action. Background of the Invention

[0005] Identifying errors in an environment and figuring out how to correct the identified errors can be difficult. In particular, errors in text can be difficult to detect. Furthermore, if errors are not detected, the errors can lead to the propagation of further errors, which leads to further confusion. The lack of real-time error detection can cause users to spend time on problems without knowing when and where they made an error.

[0006] In addition, some errors and / or problems may include difficult and complex responses to resolve the problem or error. Such difficult and complex responses may easily confuse the user; therefore, further errors may be generated when attempting to resolve the error. Summary of the invention

[0007] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.

[0008] An example aspect of the present disclosure relates to a computing system. The system may include one or more processors and one or more non-transitory computer-readable media that store instructions together, which when executed by the one or more processors causes the computing system to perform operations. The operation may include obtaining image data. The image data may describe one or more images. In some implementations, one or more images may describe an environment. The operation may include processing the image data to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of the one or more images. The operation may include determining an error in one or more images based at least in part on the semantic data. The operation may include determining a corrective action based on the semantic data and the error. In some implementations, the corrective action may describe at least one of a replacement for the error or an action to repair the error. The operation may include providing a user interface element for display based on the corrective action. The user interface element may include information data describing the corrective action.

[0009] In some implementations, determining an error in one or more images based at least in part on semantic data may include: obtaining a specific machine learning model based on the semantic data, and processing the image data with the specific machine learning model to detect errors. The error may include an inconsistency with the semantic understanding. In some implementations, the error may include a deviation from a multi-part process. The multi-part process may be associated with the semantic data. In some implementations, determining a corrective action based on the semantic data and the error may include: detecting a location of the error within an environment, determining an error-free data set associated with the semantic data and the one or more images, and determining replacement data from the error-free data set based on the location of the error within the environment.

[0010] In some implementations, an error detection model may be used to determine errors. The error detection model may generate text data based on optical character recognition, may parse the text data based on one or more features in the environment, and may process each of a plurality of parsed segments to determine errors. In some implementations, the error detection model may be trained on a plurality of mathematical proofs. The error detection model may include an optical character recognition model and a natural language processing model.

[0011] In some implementations, the image data may be generated by one or more image sensors of the mobile computing device. A user interface element may be provided for display via the mobile computing device. The mobile computing device may be a smart wearable device.

[0012] Another example aspect of the present disclosure relates to a computer-implemented method. The method may include obtaining image data by a computing system including one or more processors. The image data may describe one or more images. In some implementations, the one or more images may describe one or more pages. The method may include processing the image data by the computing system with an optical character recognition model to generate text data. In some implementations, the text data may describe text on one or more pages. The method may include determining a prompt by the computing system based on the text data. The prompt may describe a request for a response. The method may include determining, by the computing system, a multi-part response to the prompt. The multi-part response may include multiple separate responses associated with the prompt. The method may include obtaining additional image data by the computing system. The additional image data may describe one or more additional images. In some implementations, the one or more additional images may describe one or more pages with user-generated text. The method may include processing the additional image data by the computing system with an optical character recognition model to generate additional text data. The additional text data may describe user-generated text on one or more pages. The method may include determining, by the computing system, that the user-generated text deviates from the multi-part response, and providing a notification by the computing system. The notification may describe that the user-generated text has an error.

[0013] In some implementations, determining by the computing system that the user-generated text deviates from the multi-part response may include determining that the user-generated text contradicts the multi-part response. Determining by the computing system that the user-generated text deviates from the multi-part response may include determining that the user-generated text lacks one or more specific features of the multi-part response. In some implementations, one or more pages may include one or more questions. The user-generated text may include user responses to one or more questions. The method may include processing, by the computing system, image data with a machine learning model to determine a prompt and a multi-part response. In some implementations, the method may include processing, by the computing system, additional image data with a machine learning model to determine that the user-generated text deviates from the multi-part response.

[0014] Another example aspect of the present disclosure relates to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause one or more computing devices to perform operations. The operation may include obtaining image data. The image data may describe one or more images. In some implementations, the one or more images may describe one or more pages. One or more pages may include multiple characters. The operation may include processing the image data to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of the multiple characters. The operation may include determining, based at least in part on the semantic data, that the multiple characters include an error. The error may describe text that is at least one of inconsistent with or inaccurate with the semantic understanding. The operation may include determining a corrective action based on the semantic data and the error. In some implementations, the corrective action may describe at least one of a replacement for the error or an action to repair the error. The operation may include providing a user interface element for display based on the corrective action. The user interface element may include information data describing the corrective action.

[0015] In some implementations, the user interface element may include one or more pop-up elements that describe multiple sub-actions for performing a corrective action. The user interface element may include an inline overlay. The inline overlay may be used to enhance at least one of the one or more images or the one or more additional images to generate one or more enhanced images. In some implementations, the one or more enhanced images may include an inline overlay superimposed on at least a portion of the one or more pages. The inline overlay may describe the corrective action. In some implementations, the user interface element may include one or more of the enhanced images to indicate the location of the error.

[0016] Other aspects of the disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.

[0017] These and other features, aspects and advantages of various embodiments of the present disclosure will be better understood with reference to the following description and appended claims.The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The specification sets forth a detailed discussion of embodiments for those of ordinary skill in the art with reference to the accompanying drawings, in which:

[0019] Figure 1A A block diagram of an example computing system performing augmented reality coaching is depicted according to an example embodiment of the present disclosure.

[0020] Figure 1BDepicted is a block diagram of an example computing device performing augmented reality coaching according to an example embodiment of the present disclosure.

[0021] Figure 1C Depicted is a block diagram of an example computing device performing augmented reality coaching according to an example embodiment of the present disclosure.

[0022] FIG. 2A to FIG. 2E Depicted are diagrams of example augmented reality experiences according to example embodiments of the present disclosure.

[0023] Figure 3 A block diagram of an example computing system performing augmented reality coaching is depicted according to an example embodiment of the present disclosure.

[0024] Figure 4 A block diagram of an example augmented reality tutoring system is depicted, according to an example embodiment of the present disclosure.

[0025] Figure 5 A block diagram of an example augmented reality tutoring system is depicted, according to an example embodiment of the present disclosure.

[0026] Figure 6 Depicted is a flow chart of an example method to perform augmented reality tutoring according to an example embodiment of the present disclosure.

[0027] Figure 7 Depicted is a flow chart of an example method to perform image data processing and corrective action determination according to an example embodiment of the present disclosure.

[0028] Figure 8 Depicted is a flow chart of an example method to perform video data processing and corrective action determination according to an example embodiment of the present disclosure.

[0029] Fig. 9 Depicted is an illustration of an example smart wearable device for obtaining image data and providing user interface elements according to an example embodiment of the present disclosure.

[0030] Reference numerals repeated in multiple figures are intended to identify like features in the various implementations. DETAILED DESCRIPTION

[0031] Overview

[0032] In general, the present disclosure relates to an augmented reality experience that provides multiple augmented reality assets based on processed image data. In particular, the systems and methods disclosed herein can utilize image processing (e.g., optical character recognition and / or object recognition), semantic understanding, and / or one or more user interface elements (e.g., inline overlays, pop-up bubbles, and / or highlights) to provide real-time instructions to correct errors and / or solve problems. In some implementations, the systems and methods can be used for augmented reality tutoring, do-it-yourself projects, and / or error detection and correction.

[0033] For example, the system and method may include obtaining image data. The image data may describe one or more images. In some implementations, one or more images may describe an environment. The system and method may include processing the image data to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of the one or more images. The system and method may include determining errors in one or more images based at least in part on the semantic data. In some implementations, the system and method may include determining a corrective action based on the semantic data and the error. The corrective action may describe at least one of a replacement for the error or an action to repair the error. The system and method may include providing a user interface element for display based on the corrective action. The user interface element may include information data describing the corrective action.

[0034] Image data may be obtained. The image data may describe one or more images. In some implementations, the one or more images may describe an environment. The environment may include one or more questions. For example, the environment may include questions for a user to answer. Alternatively and / or additionally, the environment may include objects for completing a do-it-yourself project. The image data may be generated by one or more image sensors of a mobile computing device (e.g., a smart phone). In some implementations, the mobile computing device may be a smart wearable device (e.g., smart glasses).

[0035] The image data may be processed to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of one or more images. In some implementations, the image data may be processed with a semantic understanding model. The semantic understanding model may include one or more machine learning models. The semantic understanding model may include a natural language processing model (e.g., one or more large language models trained on multiple examples). In some implementations, the semantic understanding model may include a machine learning model trained to understand equations and / or other quantitative representations (e.g., a language model trained to perform quantitative reasoning as discussed in the following article: Minerva: Solving Quantitative Reasoning Problems with Language Models by Dyer et al. (GOOGLE AI BLOG (June 30, 2022)), https: / / ai.googleblog.com / 2022 / 06 / minerva-solving-quantitative-reasoning.htm l )). Additionally and / or alternatively, the image data may be processed with an optical character recognition model to generate text data, which may then be processed with a semantic understanding model. The semantic data may be based on text content, recognized objects, data structures, data layouts, information structures, one or more diagrams, additional input data received, context of image capture, type of image capture device, user profile data, and / or one or more other contexts. The semantic data may include one or more queries that summarize issues (e.g., questions) in focus of one or more images.

[0036] The system and method can determine errors in one or more images based at least in part on semantic data. Errors can include inconsistencies with semantic understanding. In some implementations, errors can include deviations from a multi-part process. A multi-part process can be associated with semantic data. For example, a multi-part process can include one or more actions for responding to questions and / or solving problems. Errors can be determined based on heuristics, based on obtained data, and / or based on the output of a machine learning model. Errors can be determined based on the difference in semantic intent between handwritten text and printed text. For example, semantic data can include the semantic intent of printed text and the semantic understanding of handwritten text. If the semantic understanding of handwritten text is not associated with the semantic intent of printed text, it can be determined as an error.

[0037] In some implementations, determining errors in one or more images based at least in part on the semantic data can include obtaining a specific machine learning model based on the semantic data, and processing the image data with the specific machine learning model to detect errors. For example, the semantic data can describe a specific problem type (e.g., a literary analysis problem type, a calculus problem, and / or an organic chemistry problem) and a problem-specific machine learning model (e.g., a literary analysis model, a calculus model, and / or an organic chemistry model). Alternatively and / or additionally, a mathematical engine (e.g., a mathematical function system that uses one or more processors to process the problem) can be obtained and utilized based on the semantic data.

[0038] Alternatively and / or additionally, determining a corrective action based on the semantic data and the error may include: detecting a location of the error within the environment, determining an error-free data set associated with the semantic data and one or more images, and determining replacement data from the error-free data set based on the location of the error within the environment.

[0039] In some implementations, an error detection model may be used to determine errors. The error detection model may generate text data based on optical character recognition. The error detection model may parse text data based on one or more features in an environment. In some implementations, the error detection model may process each of a plurality of parsed segments to determine errors. The error detection model may be trained on a plurality of mathematical proofs. Additionally and / or alternatively, the error detection model may include an optical character recognition model and a natural language processing model.

[0040] A corrective action may be determined based on the semantic data and the error. The corrective action may describe at least one of a replacement for the error or an action to repair the error. In some implementations, the corrective action may include indicating a location of the error in the environment and one or more actions for correctly responding to a prompt identified in the environment.

[0041] The system and method can provide a user interface element for display based on the corrective action. The user interface element can include information data describing the corrective action. In some implementations, the user interface element can be provided to be displayed via a mobile computing device. The user interface element can be provided via an augmented reality experience. The user interface element can include a highlight prompt, an embedded annotation, a pop-up bubble, and / or one or more arrows.

[0042] Additionally and / or alternatively, the system and method may continuously process image data to determine and correct the user's actions in real time via one or more user interface elements provided in response to the determined error. For example, the system and method may include obtaining image data. The image data may describe one or more images. In some implementations, one or more images may describe one or more pages. The image data may be processed with an optical character recognition model to generate text data. The text data may describe the text on one or more pages. The system and method may include determining a prompt based on the text data. The prompt may describe a request for a response. The system and method may include determining a multi-part response to the prompt. The multi-part response may include multiple separate responses associated with the prompt. In some implementations, the system and method may include obtaining additional image data. The additional image data may describe one or more additional images. One or more additional images may describe one or more pages with user-generated text (e.g., additional handwritten text and / or user-entered data (e.g., user-generated code and / or user-generated equation)). The additional image data may be processed with an optical character recognition model to generate additional text data. In some implementations, the additional text data may describe user-generated text on one or more pages. The system and method may include determining that the user-generated text deviates from the multi-part response and providing a notification. The notification may describe that the user-generated text has an error.

[0043] The system and method may obtain image data, wherein the image data describes one or more images. The one or more images may describe one or more pages. In some implementations, the one or more pages may include one or more questions. The one or more pages may include printed text and handwritten text. The one or more questions may include a mathematical equation, a writing prompt, and / or a scientific question including one or more graphs.

[0044] The image data may be processed with an optical character recognition model to generate text data. The text data may describe the text on one or more pages. The optical character recognition model may include one or more machine learning models. The optical character recognition model may include a model specifically trained on handwritten text. The text data may include recognized printed text and / or recognized handwritten text.

[0045] The system and method may determine a prompt based on text data. The prompt may describe a request for a response. The prompt may be determined based on a semantic understanding of text on one or more pages. Alternatively and / or additionally, the prompt may be a query generated based on the recognized text. The prompt may be determined based on text that includes one or more keywords associated with one or more prompts and / or one or more prompt types.

[0046] The system and method can determine a multi-part response to a prompt. The multi-part response can include multiple separate responses associated with the prompt. The multi-part response can be determined based on the output of a machine learning model, based on a heuristic method, based on one or more search results received from a search engine, and / or one or more knowledge graphs. The multi-part response can be based on the output of a machine learning model trained on one or more textbooks. For example, a machine learning model can be trained to recognize a specific type of problem based on one or more identified features, and the same or a separate model can be trained to generate a proof showing how to solve a specific problem. The generated proof can be a multi-part response, where each line of the proof is a part of the response.

[0047] Additional image data may be obtained. The additional image data may describe one or more additional images. The one or more additional images may describe one or more pages with user-generated text (e.g., additional handwritten text and / or user-entered data (e.g., user-generated code and / or user-generated equations)). The user-generated text may include a user response to one or more queries.

[0048] The additional image data may be processed with an optical character recognition model to generate additional text data. The additional text data may describe user-generated text on one or more pages and / or a computer screen. The user-generated text may describe a user's attempt to answer a prompt (e.g., answer a question).

[0049] It may be determined that the user-generated text deviates from the multi-part response. The deviation may be a deviation from the multi-part response such that the user-generated text deviates from the multi-part response. For example, the multi-part response may include taking a first action and then taking a second action, and the user-generated text may include taking the first action and then taking a third action that is not equivalent to the second action.

[0050] In some implementations, determining that the user-generated text deviates from the multi-part response may include determining that the user-generated text contradicts the multi-part response. For example, the user-generated text includes a semantic intent that contradicts the semantic intent of one or more parts of the multi-part response.

[0051] Alternatively and / or additionally, determining that the user-generated text deviates from the multi-part response may include determining that the user-generated text lacks one or more specific features of the multi-part response. For example, the multi-part response may include multiplying both sides of the equation by 2x, while the user-generated text only multiplies one side by 2x.

[0052] The system and method may provide a notification. The notification may describe that the user-generated text has an error. The notification may be provided via an augmented reality experience that renders one or more user interface elements to provide the notification. The notification may describe where the error occurred and how to resolve the error.

[0053] In some implementations, the systems and methods can process the image data with a machine learning model to determine the prompt and multi-part response. The machine learning model can be a language model trained on quantitative reasoning. In some implementations, the machine learning model can be specifically trained on one or more topics using academic materials (e.g., textbooks and / or academic articles).

[0054] In some implementations, the systems and methods can process the additional image data with a machine learning model to determine whether the user-generated text deviates from a multi-part response.

[0055] The system and method may include obtaining image data. The image data may describe one or more images. In some implementations, the one or more images may describe one or more pages. The one or more pages may include multiple characters. The image data may be processed to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of the multiple characters. The system and method may include determining, based at least in part on the semantic data, that the multiple characters include an error. The error may describe text that is at least one of inconsistent with or inaccurate with the semantic understanding. In some implementations, the system and method may include determining a corrective action based on the semantic data and the error. The corrective action may describe at least one of a replacement for the error or an action to repair the error. The system and method may include providing a user interface element for display based on the corrective action. The user interface element may include information data describing the corrective action.

[0056] The system and method may obtain image data. The image data may describe one or more images. The one or more images may describe one or more pages. In some implementations, the one or more pages may include a plurality of characters. The plurality of characters may be part of a question (e.g., a question, a writing prompt, and / or a problem statement). The characters may include letters, numbers, and / or symbols. The one or more pages may include text, pictures, shapes, charts, and / or blank spaces.

[0057] The image data may be processed to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of the plurality of characters. In some implementations, the semantic data may be based on text, pictures, shapes, charts, and / or whitespace.

[0058] The system and method can determine that a plurality of characters include errors based at least in part on the semantic data. The errors can describe text that is at least one of inconsistent with the semantic understanding or inaccurate. The inaccuracy can be determined by processing the plurality of characters with one or more machine learning models.

[0059] The system and method can determine a corrective action based on the semantic data and the error. The corrective action can describe at least one of a replacement for the error or an action to repair the error. The corrective action can include a delete action (e.g., delete a subset of multiple characters) and a write action (e.g., write one or more new characters).

[0060] The system and method can provide a user interface element for display based on the corrective action. The user interface element can include information data describing the corrective action. In some implementations, the user interface element can include one or more pop-up elements that describe multiple sub-actions for performing the corrective action. The user interface element can include an embedded overlay. The embedded overlay can be used to enhance at least one of the one or more images or one or more additional images to generate one or more enhanced images. The one or more enhanced images can include an embedded overlay superimposed on at least a portion of one or more pages. In some implementations, the embedded overlay can describe the corrective action. Additionally and / or alternatively, the user interface element can include enhancing one or more images to indicate the location of the error.

[0061] In some implementations, the systems and methods may include continuous ingestion of image data, continuous image data processing, continuous error diagnosis, and / or continuous generation of user interface elements for correcting errors. For example, the systems and methods may utilize streaming optical character recognition. User interface elements (e.g., user interface elements for notifications) may be provided in a dialogic manner such that a multi-part response may be provided in stages as the user proceeds through different parts of problem solving.

[0062] In some implementations, one or more machine learning models may be trained on one or more textbooks, real-world flashcards (e.g., flashcards for a foreign language or flashcards for scientific names of structures, elements, or compounds), architectural drawings, knowledge graphs, and / or study guides. Additionally and / or alternatively, one or more machine learning models may be trained on proofs. Training may include training that conforms to the output of a rule engine. Alternatively and / or additionally, training may include black box optimization. The systems and methods may be used for writing tasks, language learning tasks, mathematical problem solving tasks (e.g., algebra, calculus, and / or discrete mathematics), scientific problem solving tasks (e.g., physics, organic chemistry, biology, and / or chemistry), architectural design (e.g., tracing lines and following measurements), and / or surgical procedures. The systems and methods may determine that a prompt includes multiple requested criteria, and may process the user's responses to determine which criteria have been met and which criteria have not been met.

[0063] The systems and methods disclosed herein can be initiated based on one or more inputs (e.g., voice commands, button selections, and / or visual queues).

[0064] In some implementations, the system and method may include parsing one or more images and / or parsing text data generated via optical character recognition. Parsing may be based on lines, paragraphs, grammar, page structure, image regions, and / or one or more other features.

[0065] In some implementations, a software development kit may be provided to one or more instructors to customize augmented reality tutoring for their specific class, course, and / or teaching style.

[0066] The systems and methods disclosed herein provide many technical effects and benefits. As an example, the systems and methods can provide an augmented reality tutoring experience. In particular, the systems and methods disclosed herein can utilize optical character recognition, natural language processing, and augmented reality rendering to provide an interactive experience for identifying errors and providing multi-part responses.

[0067] Another technical benefit of the disclosed systems and methods is the ability to utilize one or more machine learning models to understand an environment and provide a step-by-step process for completing a task. For example, the systems and methods can determine the semantics of the environment, can determine a prompt associated with the environment, can determine a multi-part response associated with the prompt, and can continuously collect data to ensure that the user completes an action associated with the multi-part response.

[0068] Another example of technical effects and benefits relates to improved computing efficiency and improved computing system functionality. For example, the systems and methods disclosed herein can utilize the storage of determined prompts and multi-part responses to continuously compare additionally obtained data with the multi-part responses without having to continuously redetermine the semantics of the environment for error detection.

[0069] Referring now to the drawings, example embodiments of the present disclosure will be discussed in further detail.

[0070] Example Apparatus and Systems

[0071] Figure 1A A block diagram of an example computing system 100 for executing dynamic adjustment instructions in an augmented reality experience according to an example embodiment of the present disclosure is depicted. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.

[0072] The user computing device 102 may be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smartphone or tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0073] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0074] In some implementations, the user computing device 102 may store or include one or more semantic understanding models 120. For example, the semantic understanding model 120 may be or may otherwise include various machine learning models, such as a neural network (e.g., a deep neural network) or other types of machine learning models, including nonlinear models and / or linear models. The neural network may include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Figures 2A to 5 An example semantic understanding model 120 is discussed.

[0075] In some implementations, one or more semantic understanding models 120 may be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single semantic understanding model 120 (e.g., to perform parallel image data processing across multiple image instances of an environment).

[0076] More specifically, the system and method can utilize one or more machine learning models, which can include one or more semantic understanding models 120, which can process image data, text data, and / or audio data to generate semantic data associated with the environment. The semantic data can be used to detect errors in the environment and generate corrective actions for correcting the errors.

[0077] Additionally or alternatively, one or more semantic understanding models 140 may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the semantic understanding models 140 may be implemented by the server computing system 140 as part of a web service (e.g., a tutoring service). Thus, one or more models 120 may be stored and implemented at the user computing device 102, and / or one or more models 140 may be stored and implemented at the server computing system 130.

[0078] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices by which a user can provide user input.

[0079] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or a plurality of processors operatively connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0080] In some implementations, server computing system 130 includes or is otherwise implemented by one or more server computing devices. In instances where server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0081] As described above, the server computing system 130 may store or otherwise include one or more machine learning semantic understanding models 140. For example, the model 140 may be or may otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Figures 2A to 5 Discuss example model 140 .

[0082] User computing device 102 and / or server computing system 130 may train models 120 and / or 140 via interaction with training computing system 150 communicatively coupled via network 180. Training computing system 150 may be separate from server computing system 130 or may be part of server computing system 130.

[0083] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or multiple processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0084] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques such as, for example, error back propagation. For example, a loss function may be back propagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters in multiple training iterations.

[0085] In some implementations, performing error backpropagation may include performing truncated backpropagation through time.The model trainer 160 may perform a variety of generalization techniques (eg, weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0086] Specifically, the model trainer 160 can train the semantic understanding models 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, textbooks, flashcards, academic articles, equations, natural language, books, proofs, and / or homework answers.

[0087] In some implementations, if the user has provided consent, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some cases, this process may be referred to as personalizing the model.

[0088] Model trainer 160 includes computer logic for providing the desired functionality. Model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general purpose processor. For example, in some implementations, model trainer 160 includes a program file stored on a storage device, loaded into a memory, and executed by one or more processors. In other implementations, model trainer 160 includes one or more sets of computer executable instructions stored in a tangible computer readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.

[0089] The network 180 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications over the network 180 may be conducted via any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0090] The machine learning models described in this specification can be used for a variety of tasks, applications, and / or use cases.

[0091] In some implementations, the input to the machine learning model of the present disclosure may be image data. The machine learning model may process the image data to generate an output. As an example, the machine learning model may process the image data to generate an image recognition output (e.g., recognition of image data, potential embedding of image data, encoded representation of image data, hash of image data, etc.). As another example, the machine learning model may process the image data to generate an image segmentation output. As another example, the machine learning model may process the image data to generate an image classification output. As another example, the machine learning model may process the image data to generate an image data modification output (e.g., a change in image data, etc.). As another example, the machine learning model may process the image data to generate an encoded image data output (e.g., an encoded representation and / or a compressed representation of image data, etc.). As another example, the machine learning model may process the image data to generate an amplified image data output. As another example, the machine learning model may process the image data to generate a prediction output.

[0092] In some implementations, the input to the machine learning model of the present disclosure may be text or natural language data. The machine learning model may process the text or natural language data to generate an output. As an example, the machine learning model may process the natural language data to generate a language encoding output. As another example, the machine learning model may process the text or natural language data to generate a potential text embedding output. As another example, the machine learning model may process the text or natural language data to generate a translation output. As another example, the machine learning model may process the text or natural language data to generate a classification output. As another example, the machine learning model may process the text or natural language data to generate a text segmentation output. As another example, the machine learning model may process the text or natural language data to generate a semantic intent output. As another example, the machine learning model may process the text or natural language data to generate an amplified text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language, etc.). As another example, the machine learning model may process the text or natural language data to generate a predicted output.

[0093] In some implementations, the input to the machine learning model of the present disclosure may be speech data. The machine learning model may process the speech data to generate an output. As an example, the machine learning model may process the speech data to generate a speech recognition output. As another example, the machine learning model may process the speech data to generate a speech translation output. As another example, the machine learning model may process the speech data to generate a potential embedding output. As another example, the machine learning model may process the speech data to generate an encoded speech output (e.g., an encoded representation and / or a compressed representation of the speech data, etc.). As another example, the machine learning model may process the speech data to generate an amplified speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine learning model may process the speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine learning model may process the speech data to generate a predicted output.

[0094] In some implementations, the input to the machine learning model of the present disclosure may be latent coded data (e.g., a latent space representation of an input, etc.). The machine learning model may process the latent coded data to generate an output. As an example, the machine learning model may process the latent coded data to generate a recognition output. As another example, the machine learning model may process the latent coded data to generate a reconstruction output. As another example, the machine learning model may process the latent coded data to generate a search output. As another example, the machine learning model may process the latent coded data to generate a re-clustering output. As another example, the machine learning model may process the latent coded data to generate a prediction output.

[0095] In some implementations, the input to the machine learning model of the present disclosure may be statistical data. The machine learning model may process the statistical data to generate an output. As an example, the machine learning model may process the statistical data to generate an identification output. As another example, the machine learning model may process the statistical data to generate a prediction output. As another example, the machine learning model may process the statistical data to generate a classification output. As another example, the machine learning model may process the statistical data to generate a segmentation output. As another example, the machine learning model may process the statistical data to generate a segmentation output. As another example, the machine learning model may process the statistical data to generate a visualization output. As another example, the machine learning model may process the statistical data to generate a diagnostic output.

[0096] In some implementations, the input to the machine learning model of the present disclosure may be sensor data. The machine learning model may process the sensor data to generate an output. As an example, the machine learning model may process the sensor data to generate a recognition output. As another example, the machine learning model may process the sensor data to generate a prediction output. As another example, the machine learning model may process the sensor data to generate a classification output. As another example, the machine learning model may process the sensor data to generate a segmentation output. As another example, the machine learning model may process the sensor data to generate a segmentation output. As another example, the machine learning model may process the sensor data to generate a visualization output. As another example, the machine learning model may process the sensor data to generate a diagnostic output. As another example, the machine learning model may process the sensor data to generate a detection output.

[0097] In some cases, the input includes visual data, and the task is a computer vision task. In some cases, the input includes pixel data of one or more images, and the task is an image processing task. For example, the image processing task may be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the possibility that the one or more images depict an object belonging to an object class. The image processing task may be object detection, where the image processing output identifies one or more regions in the one or more images, and for each region, identifies the possibility that the region depicts an object of interest. As another example, the image processing task may be image segmentation, where the image processing output defines the corresponding possibility of each category in a set of predetermined categories for each pixel in the one or more images. For example, the group of categories may be foreground and background. As another example, the group of categories may be object classes. As another example, the image processing task may be depth estimation, where the image processing output defines a corresponding depth value for each pixel in the one or more images. As another example, the image processing task may be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted at the pixel between images in the network input for each pixel of one of the input images.

[0098] In some cases, the input includes audio data representing a spoken utterance, and the task is a speech recognition task. The output may include a text output mapped to the spoken utterance.

[0099] Figure 1AAn example computing system that can be used to implement the present disclosure is shown. Other computing systems may also be used. For example, in some implementations, the user computing device 102 may include a model trainer 160 and a training data set 162. In such implementations, the model 120 may be both trained and used locally at the user computing device 102. In some of such implementations, the user computing device 102 may implement the model trainer 160 to personalize the model 120 based on user-specific data.

[0100] Figure 1B Depicted is a block diagram of an example computing device 10 performing in accordance with an example embodiment of the present disclosure. Computing device 10 may be a user computing device or a server computing device.

[0101] Computing device 10 includes multiple applications (e.g., application 1 to application N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.

[0102] like Figure 1B As shown, each application can communicate with multiple other components of the computing device (e.g., such as one or more sensors, context managers, device state components, and / or additional components). In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to the application.

[0103] Figure 1C Depicted is a block diagram of an example computing device 50 performing in accordance with an example embodiment of the present disclosure. Computing device 50 may be a user computing device or a server computing device.

[0104] The computing device 50 includes a plurality of applications (e.g., Application 1 to Application N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a public API across all applications).

[0105] The central intelligence layer includes multiple machine learning models. For example, Figure 1CAs shown, a corresponding machine learning model (e.g., model) can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model (e.g., a single model) for all applications. In some implementations, the central intelligence layer is included in the operating system of the computing device 50 or is otherwise implemented by the operating system.

[0106] The central intelligence layer may communicate with the central device data layer. The central device data layer may be a centralized repository for data of the computing device 50. Figure 1C As shown, the central device data layer can communicate with multiple other components of the computing device (such as, for example, one or more sensors, context managers, device state components, and / or additional components). In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0107] Figure 3 A block diagram of an example computing system 300 for performing augmented reality coaching according to an example embodiment of the present disclosure is depicted. For example, the systems and methods disclosed herein may include one or more computing devices communicatively connected via a network 302. In some implementations, the computing system 300 may include one or more image sensors, one or more visual displays, one or more audio sensors, one or more audio output components, one or more storage devices, and / or one or more processors. The computing system 300 may include a smart device 304 (e.g., a smart phone) and / or a smart wearable device 306 (e.g., smart glasses). In some implementations, the smart device 304 and the smart wearable device 306 may be communicatively connected via the network 302, a Bluetooth connection, and / or via another communication medium. For example, the display, sensor, and / or processor of the smart device 304 may be utilized together with the smart wearable device 306. In some implementations, the smart device 304 and / or the smart wearable device 306 may exchange data with one or more server computing systems 308 to perform the systems and methods disclosed herein.

[0108] Example system layout and example model layout

[0109] FIG. 2A to FIG. 2E An illustration of an example augmented reality experience according to an example embodiment of the present disclosure is depicted. The example augmented reality experience can be initiated based on one or more inputs (e.g., a button press, a touch input to a touch screen, object detection or classification, and / or audio input). In response to the augmented reality experience being initiated, one or more images 202 can be obtained. In some implementations, audio input can be obtained.

[0110] One or more images 202 may be processed to determine a cue 204 in environmental focus (e.g., Figure 2A ). The focus may be determined based on a location in the environment, based on the user's orientation, based on detected objects, based on the user's gaze, and / or based on one or more machine learning parameters. Hint 204 may be highlighted, underlined, and / or indicated via one or more other techniques.

[0111] The audio input may be processed to determine the spoken utterance. The spoken utterance may be provided for display via a caption user interface element 206. The augmented reality experience may provide a user interface element to indicate that data is being obtained via an image sensor and / or an audio sensor.

[0112] The prompt 204 may be processed to determine a multi-part response for a particular prompt 204. The multi-part response may be determined based on one or more machine learning parameters, based on a knowledge graph, and / or based on data obtained from a database. A plurality of user interface elements may be generated based on a set of actions associated with the multi-part response.

[0113] exist Figure 2B In the example embodiment, prompt 204 remains highlighted and first user interface element 208 is provided for display. First user interface element 208 can describe instructions for one or more actions for completing a first part of a multi-part response.

[0114] Additional image data may then be received. The additional image data may be processed to identify new text data. The new text data may include handwritten text 210. The new text data may be processed to determine that the first part of the multi-part response has been completed. In response to this determination, a second user interface element 214 may be provided for display (e.g., such as Figure 2C ). Second user interface element 214 can describe instructions for completing one or more actions of the second part of the multi-part response. In some implementations, the augmented reality experience can provide an interesting location indicator 212, which can indicate an interesting location for a set of actions and / or can indicate a prompt type associated with prompt 204.

[0115] exist Figure 2D , further handwritten text 216 has been and is being provided. The systems and methods may have determined that a second part of a multi-part response has been detected, and may provide a third user interface element 218 for display. The third user interface element 218 may describe instructions for one or more actions for completing the third part of the multi-part response.

[0116] exist Figure 2EIn the embodiment of the present invention, a multi-part response has been performed, wherein a final handwritten response is recognized 220. The augmented reality experience may provide a completion indicator 222 and / or one or more final user interface elements 224 indicating that the set of actions has been completed.

[0117] The system and method may then be repeated for the next identified prompt.The augmented reality experience may include additional user interface elements for providing intuitive instructions.The augmented reality experience may be provided with one or more audio outputs.

[0118] Figure 4 A block diagram of an example augmented reality tutoring system 400 is depicted in accordance with an example embodiment of the present disclosure. In some implementations, the augmented reality tutoring system 400 can receive image data 402 and / or additional input data 404 describing an environment and / or a particular prompt, and provide a user interface element 420 describing one or more classifications and / or instructions for completing a task as a result of receiving the image data 402 and the additional input data 404. Thus, in some implementations, the augmented reality tutoring system 400 can include an optical character recognition model 406 operable to recognize text in an image and a semantic understanding model 410 operable to generate a semantic output 412.

[0119] For example, image data 402 may be obtained via one or more image sensors. Image data 402 may describe one or more images depicting an environment. The environment may include a plurality of characters describing one or more prompts (eg, one or more questions and / or one or more instructions).

[0120] The image data 402 may be processed with an optical character recognition model 406 to recognize one or more characters, one or more symbols, and / or one or more diagrams to generate text data 408. The text data 408 may include words, numbers, equations, diagrams, text structure, text layout, grammar, and / or symbols. The optical character recognition model 406 may be trained for printed text and / or may be specifically trained for determining handwritten characters.

[0121] In some implementations, additional input data 404 may be obtained. The additional input data 404 may be obtained via one or more additional sensors, which may include an audio sensor and / or a touch sensor. The additional input data 404 may be generated based on spoken utterances and / or one or more selections made in a user interface.

[0122] The image data 402 and / or the additional input data 404 can be processed by the semantic understanding model 410 to generate a semantic output 412. The semantic understanding model 410 may include one or more segmentation models, one or more enhancement models, one or more natural language processing models, one or more quantitative reasoning models, and / or one or more classification models. The semantic understanding model 410 may include one or more Transformer models, one or more convolutional neural networks, one or more genetic algorithm neural networks, one or more discriminator models, and / or recurrent neural networks. The semantic understanding model 410 can be trained on a large language training data set, a quantitative reasoning training data set, a textbook data set, a flashcard training data set, and / or a proof data set. The semantic understanding model 410 can be trained to determine the semantic intent of the input data and perform one or more tasks based on the semantic intent. For example, the semantic understanding model 410 can be trained for multiple tasks, which may include input summarization, response tasks, completion tasks, diagnostic tasks, problem solving tasks, and error detection tasks, classification tasks, and / or enhancement tasks.

[0123] Based on the semantic output 412, no further action 414 can be determined, which may cause the process to start again. Alternatively and / or additionally, the augmented reality coaching interface 416 can be initiated based on the semantic output 412. For example, the augmented reality coaching interface 416 can be initiated based on the semantic output 412 describing an error (e.g., an inaccuracy in the response and / or a problem with the configuration in the environment). Alternatively and / or additionally, the augmented reality coaching interface 416 can be initiated based on the semantic output 412 describing a threshold amount of time that occurs without action occurring (e.g., a threshold amount of time that occurs without new handwriting). In some implementations, the augmented reality coaching interface 416 can be initiated by user input that triggers the capture of the image data 402 and / or the additional input data 404.

[0124] The information output 418 may be determined based on the semantic output 412. For example, the semantic output 412 may describe an error in the environment, and the information output may include instructions for a set of actions (e.g., corrective actions) for correcting the error. The set of actions may be determined based on the machine learning model output, based on one or more knowledge graphs, and / or based on one or more search results generated by querying a database. The instructions for the set of actions may be arranged in a sequence.

[0125] One or more user interface elements 420 may be generated based on information output 418. In some implementations, user interface element 420 may be generated for each action in the set of actions. User interface element 420 may then be provided to the user. User interface element 420 may be provided in a sequential manner. The process may then begin again.

[0126] Figure 5 A block diagram of an example augmented reality tutoring system 500 is depicted according to an example embodiment of the present disclosure. The augmented reality tutoring system 500 is similar to Figure 4 The augmented reality tutoring system 400 , in addition, the augmented reality tutoring system 500 further includes an augmented reality generation block 518 .

[0127] The augmented reality tutoring system 500 can obtain input data, which can include words 502, numbers 504, equations 506, charts 508, structural data 510, and / or other data 512. Other data 512 can include time data, audio data, touch data, and / or context data. The input data can include multimodal data and / or can be adjusted or supplemented based on profile data and / or preference data.

[0128] The input data may be processed with a semantic understanding model 514 to generate or determine prompts. The prompts may be based on a semantic understanding of the input data, which may include a semantic understanding of the environment. The prompts may include problems to be solved (e.g., math problems, reading comprehension problems, and / or science problems), writing prompts (e.g., analysis prompts, essay prompts, and / or new literary creation prompts), and / or do-it-yourself projects (e.g., building furniture, repairing appliances, and / or maintaining vehicles).

[0129] The prompt may be processed with a response determination block 516 to generate a response. In some implementations, the response determination block 516 may be part of the semantic understanding model 514. The semantic understanding model 514 and / or the response determination block 516 may include one or more machine learning models. In some implementations, the response determination block 516 may include determining a query based on the prompt and querying a database (e.g., a search engine and / or an academic database).

[0130] The response may include a multi-part response including a set of actions. The set of actions may be part of a larger corrective action to correct an error. The response may be processed by augmented reality generation block 518 to generate a plurality of augmented reality user interface elements to be provided to the user. The plurality of augmented reality user interface elements may describe instructions for performing a plurality of actions associated with the response.

[0131] Multiple augmented reality user interface elements may include inline rendering 520 (e.g., text and / or symbols provided inline with text and / or objects in the environment), pop-up elements 522 (e.g., speech bubbles rendered in the augmented reality display), highlight elements 524 (e.g., brightening of multiple pixels and / or darkening of multiple pixels of the display environment), animation elements 526 (e.g., animated images and / or animated text that change during the passage of a presentation period), symbols (e.g., representative indicators and / or classification symbols), and / or other user interface outputs 530 (e.g., three-dimensional augmented renderings of objects in the scene).

[0132] Fig. 9 A diagram of an example smart wearable device 900 for obtaining image data and providing user interface elements according to an example embodiment of the present disclosure is depicted. For example, the system and method can be implemented via a smart wearable device 900. In some implementations, the smart wearable device 900 may include smart glasses. The smart wearable device 900 may include one or more image sensors 902, one or more computer component housings 904, one or more displays 906, and / or one or more lenses 908. The lenses may be prescription lenses, blue light blocking lenses, tinted lenses, and / or transparent non-prescription lenses. One or more image sensors 902 may be located at a position where the image data obtained describes the environment in the user's field of view. One or more computer component housings 904 may store one or more processors, one or more communication components (e.g., a Bluetooth receiver, an ultra-wideband receiver, and / or a WiFi receiver), one or more audio components (e.g., a microphone and / or a speaker), and / or one or more storage devices. One or more displays 906 may be configured to display one or more user interface elements.

[0133] For example, one or more image sensors 902 can generate image data, which can be processed by one or more processors in one or more computer component housings 904. User interface elements can be selected and / or generated based on the image data. Then, one or more user interface elements can be provided for display via one or more displays 906.

[0134] Example Method

[0135] Figure 6 Depicted is a flowchart of an example method performed in accordance with an example embodiment of the present disclosure. Figure 6 The steps performed in a specific order are depicted for the purpose of illustration and discussion, but the method of the present disclosure is not limited to the order or arrangement specifically shown. The various steps of method 600 may be omitted, rearranged in various ways, combined and / or adjusted without departing from the scope of the present disclosure.

[0136] At 602, a computing system may obtain image data. The image data may describe one or more images. In some implementations, the one or more images may describe an environment. The environment may include one or more questions. For example, the environment may include questions for a user to answer. Alternatively and / or additionally, the environment may include objects for completing a do-it-yourself project. The image data may be generated by one or more image sensors of a mobile computing device (e.g., a smart phone). In some implementations, the mobile computing device may be a smart wearable device (e.g., smart glasses).

[0137] At 604, the computing system may process the image data to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of one or more images. In some implementations, the image data may be processed with a semantic understanding model. The semantic understanding model may include one or more machine learning models. The semantic understanding model may include a natural language processing model (e.g., one or more large language models trained on multiple examples). In some implementations, the semantic understanding model may include a machine learning model trained to understand equations and / or other quantitative representations (e.g., a language model trained to perform quantitative reasoning as discussed in Dyer et al., Minerva: Solving Quantitative Reasoning Problems with Language Models (GOOGLE AI BLOG (June 30, 2022)), https: / / ai.googleblog.com / 2022 / 06 / minerva-solving-quan titative- reasoning.html )). Additionally and / or alternatively, the image data may be processed with an optical character recognition model to generate text data, which may then be processed with a semantic understanding model. The semantic data may be based on text content, recognized objects, data structures, data layouts, information structures, one or more diagrams, additional input data received, context of image capture, type of image capture device, user profile data, and / or one or more other contexts. The semantic data may include one or more queries that summarize issues (e.g., questions) in focus of one or more images.

[0138] At 606, the computing system may determine errors in one or more images based at least in part on the semantic data. Errors may include inconsistencies with semantic understanding. In some implementations, errors may include deviations from a multi-part process. A multi-part process may be associated with semantic data. For example, a multi-part process may include one or more actions for responding to questions and / or solving problems. Errors may be determined based on heuristics, based on acquired data, and / or based on the output of a machine learning model. Errors may be determined based on the semantic intent of handwritten text being different from that of printed text. For example, semantic data may include the semantic intent of printed text and the semantic understanding of handwritten text. If the semantic understanding of handwritten text is not associated with the semantic intent of printed text, it may be determined to be an error.

[0139] In some implementations, determining errors in one or more images based at least in part on the semantic data can include obtaining a specific machine learning model based on the semantic data, and processing the image data with the specific machine learning model to detect errors. For example, the semantic data can describe a specific problem type (e.g., a literary analysis problem type, a calculus problem, and / or an organic chemistry problem) and a problem-specific machine learning model (e.g., a literary analysis model, a calculus model, and / or an organic chemistry model). Alternatively and / or additionally, a mathematical engine (e.g., a mathematical function system that uses one or more processors to process the problem) can be obtained and utilized based on the semantic data.

[0140] Alternatively and / or additionally, determining a corrective action based on the semantic data and the error may include: detecting a location of the error within the environment, determining an error-free data set associated with the semantic data and one or more images, and determining replacement data from the error-free data set based on the location of the error within the environment.

[0141] In some implementations, an error detection model may be used to determine errors. The error detection model may generate text data based on optical character recognition. The error detection model may parse text data based on one or more features in an environment. In some implementations, the error detection model may process each of a plurality of parsed segments to determine errors. The error detection model may be trained on a plurality of mathematical proofs. Additionally and / or alternatively, the error detection model may include an optical character recognition model and a natural language processing model.

[0142] At 608, the computing system may determine a corrective action based on the semantic data and the error. The corrective action may describe at least one of a replacement for the error or an action to repair the error. In some implementations, the corrective action may include indicating a location of the error in the environment and one or more actions for correctly responding to a prompt identified in the environment.

[0143] At 610, the computing system may provide a user interface element for display based on the corrective action. The user interface element may include information data describing the corrective action. In some implementations, the user interface element may be provided to be displayed via a mobile computing device. The user interface element may be provided via an augmented reality experience. The user interface element may include a highlight prompt, an inline annotation, a pop-up bubble, and / or one or more arrows.

[0144] Figure 7 A flowchart is depicted for an example method performed according to an example embodiment of the present disclosure. Figure 7 The steps performed in a specific order are depicted for the purpose of illustration and discussion, but the method of the present disclosure is not limited to the order or arrangement specifically shown. The various steps of method 700 may be omitted, rearranged in various ways, combined and / or adjusted without departing from the scope of the present disclosure.

[0145] At 702, a computing system may obtain image data. The one or more images may describe one or more pages. In some implementations, the one or more pages may include one or more questions. The one or more pages may include printed text and handwritten text. The one or more questions may include a math equation, a writing prompt, and / or a science question including one or more graphs.

[0146] At 704, the computing system may process the image data with an optical character recognition model to generate text data. The text data may describe text on one or more pages. The optical character recognition model may include one or more machine learning models. The optical character recognition model may include a model specifically trained on handwritten text. The text data may include recognized printed text and / or recognized handwritten text.

[0147] At 706, the computing system may determine a prompt based on the text data and determine a multi-part response to the prompt. The prompt may describe a request for a response. The prompt may be determined based on a semantic understanding of the text on one or more pages. Alternatively and / or additionally, the prompt may be a query generated based on the recognized text. The prompt may be determined based on text including one or more keywords associated with one or more prompts and / or one or more prompt types.

[0148] A multi-part response may include multiple separate responses associated with a prompt. A multi-part response may be determined based on the output of a machine learning model, based on a heuristic, based on one or more search results received from a search engine, and / or one or more knowledge graphs. A multi-part response may be based on the output of a machine learning model trained on one or more textbooks. For example, a machine learning model may be trained to recognize a particular type of problem based on one or more identified features, and the same or a separate model may be trained to generate a proof showing how to solve a particular problem. The generated proof may be a multi-part response in which each line of the proof is a part of the response.

[0149] At 708, the computing system may obtain additional image data. The additional image data may describe one or more additional images. The one or more additional images may describe one or more pages with user-generated text (e.g., additional handwritten text and / or user-entered data (e.g., user-generated code and / or user-generated equations)). The user-generated text may include user responses to one or more queries.

[0150] At 710, the computing system may process the additional image data with an optical character recognition model to generate additional text data. The additional text data may describe user-generated text on one or more pages. The user-generated text may describe a user's attempt to answer a prompt (e.g., answer a question).

[0151] At 712, the computing system may determine that the user-generated text deviates from the multi-part response and provide a notification. The deviation may be a deviation from the multi-part response such that the user-generated text deviates from the multi-part response. For example, the multi-part response may include taking a first action and then taking a second action, and the user-generated text may include taking the first action and then taking a third action that is not equivalent to the second action.

[0152] In some implementations, determining that the user-generated text deviates from the multi-part response may include determining that the user-generated text contradicts the multi-part response. For example, the user-generated text includes a semantic intent that contradicts the semantic intent of one or more parts of the multi-part response.

[0153] Alternatively and / or additionally, determining that the user-generated text deviates from the multi-part response may include determining that the user-generated text lacks one or more specific features of the multi-part response. For example, the multi-part response may include multiplying both sides of the equation by 2x, while the user-generated text only multiplies one side by 2x.

[0154] The system and method may provide a notification. The notification may describe that the user-generated text has an error. The notification may be provided via an augmented reality experience that renders one or more user interface elements to provide the notification. The notification may describe where the error occurred and how to resolve the error.

[0155] In some implementations, the systems and methods can process the image data with a machine learning model to determine the prompt and multi-part response. The machine learning model can be a language model trained on quantitative reasoning. In some implementations, the machine learning model can be specifically trained on one or more topics using academic materials (e.g., textbooks and / or academic articles).

[0156] In some implementations, the systems and methods can process the additional image data with a machine learning model to determine whether the user-generated text deviates from a multi-part response.

[0157] Figure 8 A flowchart is depicted for an example method performed according to an example embodiment of the present disclosure. Figure 8 The steps performed in a specific order are depicted for the purpose of illustration and discussion, but the method of the present disclosure is not limited to the order or arrangement specifically shown. The various steps of method 800 may be omitted, rearranged in various ways, combined and / or adjusted without departing from the scope of the present disclosure.

[0158] At 802, a computing system may obtain video data. The image data may describe one or more images. The one or more images may describe one or more pages. In some implementations, the one or more pages may include a plurality of characters. The plurality of characters may be part of a question (e.g., a question, a writing prompt, and / or a problem statement). The characters may include letters, numbers, and / or symbols. The one or more pages may include text, pictures, shapes, charts, and / or blank spaces.

[0159] At 804 , the computing system may process the video data to generate recognition data.

[0160] At 806, the computing system may process the recognition data to generate semantic data. The semantic data may describe a semantic understanding of at least a portion of the plurality of characters. In some implementations, the semantic data may be based on text, pictures, shapes, charts, and / or whitespace.

[0161] At 808, the computing system may determine that the plurality of characters include an error based at least in part on the semantic data. The error may describe text that is at least one of inconsistent with the semantic understanding or inaccurate. The inaccuracy may be determined by processing the plurality of characters with one or more machine learning models.

[0162] At 810, the computing system may determine a corrective action based on the semantic data and the error. The corrective action may describe at least one of a replacement for the error or an action to repair the error. The corrective action may include a delete action (e.g., deleting a subset of a plurality of characters) and a write action (e.g., writing one or more new characters).

[0163] At 812, the computing system may provide a user interface element for display based on the corrective action. The user interface element may include information data describing the corrective action. In some implementations, the user interface element may include one or more pop-up elements that describe multiple sub-actions for performing the corrective action. The user interface element may include an embedded overlay. The embedded overlay may be utilized to enhance at least one of the one or more images or one or more additional images to generate one or more enhanced images. The one or more enhanced images may include an embedded overlay superimposed on at least a portion of one or more pages. In some implementations, the embedded overlay may describe the corrective action. Additionally and / or alternatively, the user interface element may include enhancing one or more images to indicate the location of the error.

[0164] Additional public content

[0165] The technology discussed herein relates to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for multiple possible configurations, combinations, and partitioning of tasks and functions between and among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system, or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0166] Although the present subject matter has been described in detail with respect to various specific example embodiments of the present subject matter, each example is provided by way of explanation rather than limitation of the present disclosure. Those skilled in the art may easily produce changes, modifications, and equivalents to such embodiments after understanding the foregoing. Therefore, the present disclosure does not exclude such modifications, variations, and / or additions to the present subject matter that would be readily apparent to those of ordinary skill in the art. For example, a feature shown or described as part of one embodiment may be used together with another embodiment to produce yet another embodiment. Therefore, the present disclosure is intended to encompass such changes, variations, and equivalents.

Claims

1. A computing system, the system include: one or more processors; as well as one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations comprising: obtaining image data, wherein the image data describes one or more images, wherein the one or more images describe an environment; processing the image data to generate semantic data, wherein the semantic data describes a semantic understanding of at least a portion of the one or more images; determining errors in the one or more images based at least in part on the semantic data; determining a corrective action based on the semantic data and the error, wherein the corrective action describes at least one of a replacement for the error or an action to repair the error; and A user interface element is provided for display based on the corrective action, wherein the user interface element includes informational data describing the corrective action.

2. The system according to claim 1, in, Determining the error in the one or more images based at least in part on the semantic data comprises: Obtaining a specific machine learning model based on the semantic data; and The image data is processed with the particular machine learning model to detect the errors.

3. A system according to any preceding claim, in, The errors include inconsistencies with the semantic understanding.

4. A system according to any preceding claim, in, The error comprises a deviation from a multi-part procedure, wherein the multi-part procedure is associated with the semantic data.

5. A system according to any preceding claim, in, Determining the corrective action based on the semantic data and the error comprises: detecting a location of the error within the environment; determining an error-free data set associated with the semantic data and the one or more images; and Replacement data is determined from the error-free data set based on the location of the error within the environment.

6. A system according to any preceding claim, in, The error is determined by an error detection model, wherein the error detection model: generating text data based on optical character recognition; parsing the text data based on one or more features in the environment; and Each of the plurality of parsed segments is processed to determine the error.

7. The system according to claim 6, in, The error detection model is trained on multiple mathematical proofs.

8. The system according to claim 6, in, The error detection model includes an optical character recognition model and a natural language processing model.

9. A system according to any preceding claim, in, The image data is generated by one or more image sensors of a mobile computing device, and wherein the user interface element is provided for display via the mobile computing device.

10. The system according to claim 9, in, The mobile computing device is a smart wearable device.

11. A computer-implemented method, the method include: Obtaining, by a computing system including one or more processors, image data, wherein the image data describes one or more images, wherein the one or more images describe one or more pages; processing the image data using an optical character recognition model by the computing system to generate text data, wherein the text data describes text on the one or more pages; determining, by the computing system, a prompt based on the textual data, wherein the prompt describes a request for a response; determining, by the computing system, a multi-part response to the prompt, wherein the multi-part response comprises a plurality of individual responses associated with the prompt; obtaining, by the computing system, additional image data, wherein the additional image data describes one or more additional images, wherein the one or more additional images describe the one or more pages having user-generated text; processing, by the computing system, the additional image data with the optical character recognition model to generate additional text data, wherein the additional text data describes the user-generated text on the one or more pages; determining, by the computing system, that the user-generated text deviates from the multi-part response; and A notification is provided by the computing system, wherein the notification describes the user-generated text having an error.

12. The method according to claim 11, in, Determining, by the computing system, that the user-generated text deviates from the multi-part response includes: It is determined that the user-generated text contradicts the multi-part response.

13. A method according to any preceding claim, in, Determining, by the computing system, that the user-generated text deviates from the multi-part response includes: It is determined that the user-generated text lacks one or more specific features of the multi-part response.

14. A method according to any preceding claim, in, The one or more pages include one or more questions, and wherein the user-generated text includes user responses to the one or more questions.

15. A method according to any preceding claim, further comprising: include: The image data is processed by the computing system using a machine learning model to determine the prompt and the multi-part response.

16. A method according to any preceding claim, further comprising: include: The additional image data is processed, by the computing system, with a machine learning model to determine that the user-generated text deviates from the multi-part response.

17. One or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations include: obtaining image data, wherein the image data describes one or more images, wherein the one or more images describe one or more pages, wherein the one or more pages include a plurality of characters; processing the image data to generate semantic data, wherein the semantic data describes a semantic understanding of at least a portion of the plurality of characters; determining, based at least in part on the semantic data, that the plurality of characters include an error, wherein the error describes text that is at least one inconsistent with or inaccurate with the semantic understanding; determining a corrective action based on the semantic data and the error, wherein the corrective action describes at least one of a replacement for the error or an action to repair the error; and A user interface element is provided for display based on the corrective action, wherein the user interface element includes informational data describing the corrective action.

18. One or more non-transitory computer readable media according to claim 17, in, The user interface element includes one or more pop-up elements describing a plurality of sub-actions for performing the corrective action.

19. One or more non-transitory computer readable media according to any preceding claim, in, The user interface element comprises an inline overlay, wherein the inline overlay is used to enhance at least one of the one or more images or the one or more additional images to generate one or more enhanced images, wherein the one or more enhanced images comprise the inline overlay superimposed on at least a portion of the one or more pages, and wherein the inline overlay describes the corrective action.

20. One or more non-transitory computer readable media according to any preceding claim, in, The user interface elements include enhancing one or more of the images to indicate the location of the error.