System and method for automated grading and quality assessment of educational content

The system uses LLMs to automate grading by analyzing educational content, addressing the inefficiencies and inconsistencies of traditional methods, ensuring accurate and consistent evaluation through OCR and noisy data elimination, thereby enhancing the grading process.

WO2025206954A1PCT designated stage Publication Date: 2025-10-02THE TEACHER

Patent Information

Application Number
PCT/NL2025/050151
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-03-28
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Traditional grading systems require significant time and effort from educators, and are prone to subjective inconsistencies due to their manual nature, necessitating a more efficient and accurate automated grading solution.

Method used

A computer-implemented system using large language models (LLMs) to analyze and quantify the quality of educational content by comparing user inputs to predefined standards, incorporating a graphical user interface for data upload and output, and employing optical character recognition (OCR) to extract textual information from various document formats, with accuracy verification and noisy data elimination mechanisms.

Benefits of technology

This approach streamlines grading, reduces preparation time, enhances accuracy, and ensures consistency by providing precise, unbiased analysis of textual data, while maintaining transparency and flexibility across different educational contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure NL2025050151_02102025_PF_FP_ABST
    Figure NL2025050151_02102025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method for analyzing and quantifying the quality of a first user input in relation to a second user input by comparison to a third user input in a computer system, wherein the second user input defines at least one problem statement, the third user input defines at least one solution statement to the at least one problem statement, the method comprising: - providing a graphical user interface for receiving said first, second and third user inputs; - generating at least one prompt structure based on the first, second and third user inputs, such as a prompt structure for each of the first, second and third user inputs, wherein the at least one prompt structure comprises instructions to quantify the quality of a first user input, wherein the quantified quality at least in part reflects the degree to which the first user input corresponds to the at least one solution statement as defined by the third user, such as in its completeness, and wherein the quantified quality further at least in part reflects its correctness in relation to the at least one problem statement as defined by the second user input; - processing the at least one prompt structure using at least one large language model; and - outputting the quantified quality of the first user input from the at least one large language model, such that it is accessible to a user within the graphical user interface.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] System and Method for Automated Grading and Quality Assessment of Educational Content Field of the invention: The present invention relates to the field of educational technology, particularly to a system and method for automated grading and quality assessment of educational content using artificial intelligence (AI) and machine learning models. Background of the invention: Traditional grading systems require significant time and effort from educators to assess and grade student submissions accurately. Moreover, the subjective nature of grading can introduce inconsistencies in evaluation. With the advancement in AI and machine learning, there is a pressing need for a system that can automate the grading process while maintaining high accuracy and consistency. Summary of the invention: The invention provides a computer-implemented method and system for analyzing, grading, and quantifying the quality of educational content, such as exam answers, using large language models (LLMs). The method involves comparing user inputs, including problem statements, solution statements, and responses, to generate quantified quality assessments. These assessments reflect the correctness and completeness of the responses in relation to predefined answers. The system incorporates a graphical user interface (GUI) for receiving inputs and displaying the quantified quality assessments. 1. Core Methodology The invention initiates with a computer-implemented method to analyze and quantify the quality of user inputs by comparison within a computational system. This system uses a graphical user interface (GUI) to receive user inputs that define problem statements, solution statements, and corresponding responses. The process includes generating a prompt structure for each input, processed through an LLM, to output a quantified quality assessment accessible via the GUI. This methodology beneficially streamlines the evaluation process, allowing for a scalable, precise, and unbiased analysis of textual data. It supports educators and analysts in assessing the completeness and correctness of responses relative to predefined standards. More accurately, the core methodology pertains to: A computer- implemented method for analyzing and quantifying the quality of a first user input in relation to a second user input (2) by comparison to a third user input in a computer system, wherein the second user input defines at least one problem statement, the third user input defines at least one solution statement to the at least one problem statement, the method comprising: - providing a graphical user interface for receiving said first, second and third user inputs; - generating at least one prompt structure based on the first, second and third user inputs, such as a prompt structure for each of the first, second and third user inputs, wherein the at least one prompt structure comprises instructions to quantify the quality of a first user input, wherein the quantified quality at least in part reflects the degree to which the first user input corresponds to the at least one solution statement as defined by the third user, such as in its completeness, and wherein the quantified quality further at least in part reflects its correctness in relation to the at least one problem statement as defined by the second user input; - processing the at least one prompt structure using at least one large language model; and - outputting the quantified quality of the first user input from the at least one large language model, such that it is accessible to a user within the graphical user interface. 2. Digital Document Integration The GUI may be designed to facilitate the uploading of digital documents, enabling the extraction of textual information, including complex elements like mathematical formulas, via optical character recognition (OCR). This feature democratizes the analysis process, allowing for a broad range of document formats to be directly uploaded and evaluated. It significantly reduces the preparation time by automating the extraction of textual data from physical or scanned documents. 3. Accuracy Assurance A further improvement may see the method comprise steps to verify the accuracy of the OCR-extracted information. If accuracy falls below a predetermined threshold, the system alerts the user through the GUI. This step ensures the reliability of the data fed into the LLM, thereby enhancing the overall accuracy of the analysis. It acts as a quality control mechanism, prompting users to provide clearer documents when necessary. Accuracy may be determined to fall below a predetermined threshold. OCR software can express a level of confidence it has in the characters that it has recognized. Usually, such confidence can be expressed in a percentage. 98% would for example mean that the OCR software believes that 980 out of 1000 characters are confidently recognized, whereas 20 out of 1000 are doubtful. A threshold can be selected at 95-99.9%, preferably 98% or higher, because inaccurate character recognition may introduce noisy data. Alternatively, a confidence can be expressed on a character by character basis and subsequently averaged. This is particularly relevant for mathematics exercises in which alternative symbols, such as Greek symbols, are used together with numbers. In such a case, the average confidence in the accurate recognition of a detected formula, can be used as a threshold. Whatever the situation may be 'noisy data’ can be corrected by an LLM in light of contextual information. This allows the LLM to restore text to what was most likely intended. For example, “vebcity” could be corrected to “velocity” when the exam question pertained to a physics question related to the calculation of the speed of an object. 4. Prompt Generation Control Building on the previously mentioned it is useful to prevent prompt generation until the OCR accuracy meets the required standards. Users may be ‘prompted’, that is to say request by via the GUI, to replace or adjust their uploads to satisfy this threshold. This control mechanism beneficially reinforces the system's commitment to accuracy, ensuring that only high-quality data informs the analysis process. It further prevents the propagation of errors through the subsequent steps of evaluation. 5. Noisy Data Elimination Applicable to all embodiments, this option involves using the LLM to eliminate noisy or irrelevant data from the prompts or user inputs. Cleaning the data before analysis drastically improves the quality and relevance of the feedback provided. This step thus ensures that evaluations are based solely on meaningful content, enhancing the utility of the output. In a preferred embodiment the prompt generation prevention threshold is chosen such that, when the noisy data elimination is applied, the outputted quantified quality does not deviate more than 5-10% of the maximum obtainable quantity, such as a grade, with respect to a quantified quality, such as a grade, as generated by a human benchmark provided with the same first, second and third user inputs. A human benchmark provides for an objective measure for the quantified quality by a human. A human benchmark can be obtained for cases in which there are a minimum of 20 mutually different first user inputs that pass the threshold, for which at least 95%, are assigned an averaged quantified quality which concurs with the allowed deviation of 5-10% of the maximum obtainable quantity, and wherein the assigned quantified quality of a human is the average of two humans per first user input. In this example of benchmarking the second and third user inputs do not vary. In other words, the threshold is chosen such that for 95% of cases passing the threshold the computer generated grade differs no more than 0,5-1 points on a decimal grading scale as one would have gotten had a team of two co-correctors performed the evaluation. In a more preferable example computer generated grades, a form of quantified quality, are flagged that are within allowed deviation from a failing threshold. Optionally, the method may comprise a step of requesting a user to correct the outputted quantified quality, wherein the system is programmed to use the correction to adjust the threshold. In one example, the threshold is adjusted to be more exclusive of a first user input when the user’s corrected quantified quality exceeds the allowed deviation of 5-10%, and wherein the threshold is optionally adjusted to be less exclusive of a first user input when the user’s corrected quantified quality is within 0-4% of the outputted quantified quality. 6. Motivation Output Optionally, the system outputs not just the quantified quality of the first user input but also provides a motivation output, explaining the LLM's rationale behind the quantified quality. This transparency in the evaluation process aids users in understanding the basis of their feedback, facilitating learning and improvement. It offers a deeper insight into the analysis, allowing for more informed decision-making and adjustments. 7. Sequential LLM Processing and Optimization Optionally the at least one large language model comprises a plurality of large language models in sequence, wherein the output of a first model of the plurality becomes the input for a second model of the plurality to enrich the context or information available for the second model in the sequence. This layered approach enriches the analysis, allowing for nuanced understanding and evaluation of inputs. It enables the system to tailor the evaluation process to the specific requirements of each task, enhancing both accuracy and relevance. 8. Sequential LLM Processing and Optimization - expanded Further optionally the prompt structure may be generated by the first large language model, wherein the prompt structure is processed by the second large language model to output the quantified quality, and wherein preferably the first and second large language models are more suitable, such as more optimized, than the other for their mutually different task within the sequence. It should be noted though, also separately from this that the first large language model and second large language model may simply be the same large language model. Further entirely separate from these features it can be understood that text recognition may be performed either by an LLM or by known OCR-software. A structure is possible where the such an LLM is a separate third LLM for extracting text, or the same LLM as the first and the second LLM. Currently, a very suitable single LLM structure is obtained when using Open AI’s - GPT-4-1106, or GPT- 4-Turbo for such LLM. The person skilled in the art will understand that there are many different alternatives. 8 bis. Sub LLM It is possible for student writing to be intelligible in varying degrees. While OCR itself may already provide a suitable output for assessment, assignment context and subject may allow for OCR data to be cleaned further with the context – such as the problem statement, also known as the second user input - in mind. The first LLM can comprise at least one customized sub-LLM, wherein such customization derives from being trained on a specific dataset, such as a school subject like math, physics, history or any other. For example: A sub-LLM trained on physics – such as a dataset of teacher problem statements and solutions in math - may derive from context that a “p” in a first user input should have probably been interpreted to be a Greek symbol “ρ”. Thereby, the sub-LLM may correct for erroneously interpreted written text. It is noted that contextual correction is beneficial whether OCR – read a third LLM - is used or not. The first LLM, comprising at least one sub-LLM, may simply be arranged to present data to the second LLM in a prompt that is suitably digestible. The at least one sub-LLM may be selected from a series of sub- LLMs each trained on a mutually different dataset. Alternatively and / or additionally, the at least one sub-LLM comprises multiple sub-LLMs that process user inputs in parallel. Such parallel proofreading adds to consistency in grading. It helps reduce bias, catch errors and improve consistency in grading. 8 ter. Main LLM The second large language model may be modularly interchangeable by a main large language model, such as by user selection such as in the GUI. Further optionally, the main large language model may be selectable by a user for a differently optimized task and / or overarching subject such as comprising a subject of the sub large language model. Optionally, the second large language model may receive sanitized text and / or structured data, such as JSON, from the first large language model, and may use ensemble logic, such as weighted voting or heuristics for final grading. 8 quarter: sub LLM and Main LLM architecture The at least one sub large language model may comprise a pre- processing step of the data, wherein a cleaning and / or a tokenization step is performed in order to remove personal identifiable data, such as names, page references, brand names, before it is send to the main large language model so as to maintain privacy. Optionally, the cleaning and / or tokenization step may comprise chunking the text into one or more sections for efficient training. Optionally, the cleaning and / or tokenization step may comprise stopping and / or removing the to the personal identifiable data corresponding container and data volume. Optionally, the first sub large language model may comprising the cleaning step, and the second large language model may comprise the tokenization step. Advantageously, the first large language model may comprise an overarching subject, such as STEM, and first large language model may comprise a relatively more specific subject compared to STEM, such as physics or chemistry. Further optionally, the main large language model may comprise a fault tolerance, wherein if one sub large language model is offline, the main large language model may still operate using the remaining sub large language models. 9. LLM Selection Further optionally, the first large language model is used to select the second large language model from a list of large language models which it determines to be most suitable for the processing of the prompt structure. This may allow a user to run the prompt structure in a more cost-effective manner without sacrificing quality. It is widely known that LLM’s differ in operational costs. 10. Task-Type Customization It is possible to expand the flexibility of the method by allowing users to specify the type of task for analysis, enabling the system to optimize the selection and sequence of LLMs accordingly. This beneficially customizes the evaluation process to suit different subjects or types of analysis, from mathematical exercises to textual analysis, thereby increasing the system's versatility and effectiveness in various educational or analytical contexts. 11. Fault Tolerance In one example a first LLM comprises at least one sub-LLM trained on a specific dataset, wherein the method comprises skipping the first LLM when the at least one sub-LLM becomes unavailable, such as because of rate limits, downtime, or any other communication disruption, and feeding data directly into the second LLM so operate without the first LLM. Alternatively, the at least one sub-LLM preferably comprises a plurality of sub-LLMs trained on mutually different datasets, and wherein preferably, when one sub-LLM of the plurality of sub-LLMs is becomes unavailable, such as through a communication disruption, the second LLM operates using at least one of the remaining sub-LLMs. The above renders the system robust under some communication disruption within the system. The person skilled in the art will understand that first and second LLMs may communicate over APIs as well as a third LLM, such as for OCR. All communication to and from these LLMs or services may mediated by APIs allowing the system to be decentralized. 12. System Implementation The invention may be provided as a system comprising a computer, which itself may in turn comprise a human-machine interface. The system, in particular the computer, may be preprogrammed to perform the method according to any one of the mentioned method steps. 13-14. System Implementation by server In order to reduce the system requirements on the user’s device, the human interface may be provided as a desktop, laptop, tablet or smartphone, programmed to provide the GUI to the user, wherein the computer is actually a distant server that is communicatively connected with the human-machine interface via the internet for processing the at least one prompt structure and optionally for generating the at least one prompt structure. It stands to reason that the server would return the output to the human-machine interface for display via the GUI, or even issue a warning for display via the GUI, if for example any of the uploaded document comprise noisy data above a predefined threshold. The above outlines the implementation of the method in a system comprising a computer with a human-machine interface, potentially hosted on a distant server, and accessible via various devices. This systematization ensures that the method is widely accessible, user-friendly, and capable of handling large volumes of data remotely. It democratizes access to advanced analytical tools, making them available to a broader audience across different platforms. 15. App store It is foreseeable that a user will practically wish to download an application from a digital library which when installed on a device, such as a laptop, desktop, tablet or smartphone, will allows a user to perform the method as described herein. The person skilled in the art will understand that method steps can be executed on any computer or group of computers that communicate with each other, wherein the device in one such computer of a group of computers providing the opportunity for receiving the input and conveying the output, such as via a graphical user interface. 16. Flow enhancement The graphical user interface comprises two separate sub- interfaces, at least one student interface and at least one teacher interface. The graphical user interface may be designed to provide the at least one teacher interface with a test environment for quality control, wherein the environment is designed for a teacher to upload first, second and third user inputs to verify the quality of the computer-implemented method of assessment. Separately from the above, but compatible therewith, the graphical user interface may be designed to provide the at least one teacher interface with an environment in which the teacher is presented with the grading as performed by the computer- implemented method. The teacher may make adjustments to any output or outputs and subsequently accept and publish the data such that at leas the grade becomes available for viewing via the at least one student interface. Optionally, the graphical user interface is designed so that a teacher is allowed to enter feedback for a student, which feedback is made available in the at least one student interface. Optionally, the at least one student interface comprises a plurality of student interfaces dedicated to mutually different users, wherein grades pertaining to any one student are, preferably exclusively, made available that student in the corresponding student interface associated with said student as well as the at least one teacher interface. To this end, the first user input may comprise an user identifier, such as an anonymous one. 17. Dispute tracking The method, also separately from this example and compatible with all embodiments, may comprises receiving the user inputs from a web application defining a user interface, wherein the user interface is provided with student and teacher specific user environments, and wherein the method provides for receiving and processing a request from the student user environment pertaining to the quantified quality of the first user input, and wherein the method provides for receiving and processing a response to said request from the teacher specific user environment. The beneficially allows for dispute tracking between teacher about the grading. Optionally, the request is a second opinion request. 18. Quality control The method, according to item 17, wherein the method comprises issuing a request to check the quantified quality of the first user input, wherein issuing such a request is randomized. This provides for a quality control on the grading method. 19. Training of Sub-LLMs In the event that a first LLM and a second LLM are provided in sequence, a first and second LLM may be run on separate servers connected via an API. In some instances the second LLM and API may be a third party, whereas the first LLM is internally hosted, such as by a school on a school server. In such a situation it is preferable to prevent data protection issues. The first LLM may comprise sub-LLMs trained on publisher datasets, wherein all training of the sub-LLMs occurs exclusively on a first server, and wherein only anonymized data is shared with the second LLM on a second server. Detailed invention: Figure 1 shows the method architecture according to the invention in which the method steps outside of that which is provided as user input is implemented by a computer. Figure 2 shows the flow of data according to the invention. Figure 3 shows another flow of data according to the invention. With reference to Figure 1, the architecture shows that a first step involves providing a first user input 1UI, a second user input 2UI and a third user input 3UI. In this example these inputs are each PDF documents uploaded via a Graphical User Interface for processing by the computer. Such PDF documents are here digital representations of a student answer, a question paper, and a model answer respectively. It is explained that exams in particular are still sat and written down on physical documents, which can be converted into PDF-format by scanning them. However, exams may also be sat digitally, allowing for the exam documents to be provided in alternative formats, such as .doc, .docx, .txt, .rtf, html, or any other file in which text is available without the need for optical character recognition. In so far that an input is a scanned document, an optical character recognition step OCR may be performed on said document. Optionally, the document is converted to an image file, such as .png, .jpg or other such extensions prior to the OCR step in a PDF^IMG step. The optional steps are provided with dashed lines --- in Figure 1. The person skilled in the art will understand that these steps depend on the format in which the inputs 1UI, 2UI and 3UI are provided, which may further mutually differ. In one example, the question paper and / or model answer, representing 2UI and 3UI respectively can be provided as .doc, .docx, .txt, .rtf, html, of .pdf in which text is accessible, leaving only the student document to be a scanned document. In the OCR step the amount or ratio of noisy data is determined for at least the student answer, 1UI. Based on this amount or ratio the method involves returning a warning about the low readability of the document and / or a request to replace the user input. This is the NON step in Figure 1. Since a traditional grading process is usually performed by a teacher on multiple students using the same model answer and the same question paper it stands to reason that the readability of 2UI and 3UI are absolutely critical to the analysis. Accordingly, 2UI and 3UI are preferably free of noisy data, and uploaded as .doc, .docx, .txt, .rtf, html, of .pdf with available digital text rather than scanned. However, in case the user uploads a scanned 2UI and / or 3UI, and also separately from this example, a lower threshold for either the amount or ratio of noisy data may be applied for 2UI and / or 3UI with respect to 1UI, the student answer. A large language model (LLM), such as Open AI’s GPT-4-1106 in the version it was available on 26.3.2024, may be used to construct a prompt structure based on the user inputs, in which the prompt structure is substantially free of any noisy data. This is method step PROMPT. The LLM may here also be used to attempt to interpret any noisy data favorably in the given context of question paper, but preferably not given the context of the model answer. This avoids that noisy data is reconstructed to meet the desired answer. In other words, poorly intelligible answers mustn’t derive an advantage for being poorly intelligible. It is possible for the method step to comprise generating a prompt for each of the user inputs separately first, PROMPT1 for 1UI, PROMPT2 for 2UI, and PROMPT3 for 3UI. The prompt or prompts, if there are multiple, are subsequently merged into a so called CO-STAR framework. This framework is characterized by comprising the following prompt structures: 1) Context 2) Objective 3) Style 4) Tone 5) Audience This visualized in the step CO-STAR. The framework is arranged such that the processing of the prompt or prompts using said CO- STAR framework results in a first output OUTPUT1 and a second output OUTPUT2, wherein OUTPUT1 comprises giving a marks compared to total marks in table format, optionally broken down per sub-question or per part, and wherein OUTPUT2 comprises explains what is erroneous about a question and how this could have been improved upon to obtain full marks per question, and optionally sub-question or part. Figure 2 shows a low resolution representation of a data flow 1000 to elucidate the interfacing of separate software-hardware components in the implementation of the invention. User inputs UI1, UI2 and U3, as shown in Figure 1, may be provided to a front-End User Interface (FE UI) Web Application 1001 that may be run on a server and accessed via a webpage. The Web Application comprises a graphical user interface GUI with separate dashboards for teacher and student. It allows for the uploading of PDFs, namely the user inputs, and viewing of grading. The Web App 1001 connects to a file extraction module 1002, for PDF^ IMG and OCR, as shown in Figure 1 via an application programming interface (API). This allows the App to upload the user inputs to the extraction module. The file extraction module 1002 can use libraries like ‘pfdminer’ for text extraction and ‘pdf2image’ for converting pdf pages to images. The module 1002 may programmed to organize the extracted data into structured formats, i.e. text. Module 1002 connects to a grading module 1003 for feeding it the structured formats. The Web App 1001 also connects to a grading module 1003. The grading module 1003 is a set of scripts that call external services 1004 via API’s, such as Mathpix for OCR for extracting text from images, and LLMs, such as OpenAI’s ‘ChatGPT’, such as via ‘ LangChain’, for receiving grading logic and answer interpretation. The grading module also builds a prompt to feed to an external service 1004 LLM. The grading module 1003 receives output from the external service LLM, and converts the output from the external services, which is generally raw text, into graded or feedback-ready output, such as via CO-STAR. In this example all data is processed on external servers grading and other outputs are fed back into the user interface and available in their corresponding dashboards. Figure 3 shows a more advanced representation of a data flow 2000 to elucidate the interfacing of separate software-hardware components in the implementation of the invention. In this example, which differs from that of Figure 2, an FE UI Web App 2001 is similarly represented. The App 2001 also provides teacher and student dashboards, and allows for first, second and third user inputs to be received. The App feeds the data into a file extraction module 2002 so as to obtain text from either PDF or images. The text is fed into a Sub-LLM 2003 which anonymizes the text and pre-processes the text for ingestion by a Main LLM 2004. The Sub-LLM may comprises multiple sub-LLMs either arranged to process input data in parallel or in series, or both, each trained on a specific publisher’s dataset. The publisher here may be the teacher. Publisher hereinafter, and also separately of this particular section to be understood as an entity responsible for creating, managing, or distributing tests. This is generally a teacher. The Main LLM 2004 provides grading results 2005 which are fed back to the App 2001 to be reflected in the relevant interfaces. In this example the first or second LLM [represented by 2003 as a sub-LLM, and 2004 as main LLM] may each run on a dedicated server without need to connect with each other, or 2001 or 2002 via an API. It is believed that rate limits, downtime and policy changes from external APIs can disrupt grading. Accordingly, also separate from this example - and entirely compatible with any embodiment comprising a first and second LLM - a computer implemented method defining a first and second LLM may be such that a first LLM and a second LLM communicate via an API, and wherein the first LLM assumes the task of the second LLM for grading when communication over the API indicates a rate limit or downtime of the second. Optionally, the first LLM may comprise a plurality of sub-LLMs, wherein one sub-LLM in the plurality assumes the task of the second LLM. While the above pertains to high-level overview the following detailed specifications can be focused on. Sub-LLMs: Integration Steps Data Ingestion & Preprocessing 1. Data Acquisition a. Securely upload or sync publisher-provided materials to a dedicated storage area (e.g. AWS S3 buckets). 2. Cleaning & Tokenization a. Remove PII; standardize text, possibly chunk into smaller sections for efficient training. 3. Labeling (If Needed) a. For QA tasks, label question-answer pairs to guide the sub-LLM. Model Training 1. Base Model a. Start with an open-source LLM (e.g., DeepSeek). b. Initialize a containerized environment for each publisher to avoid cross-data contamination. 2. Test with RAG a. Test to use RAG for to see with publisher’s data 3. If required, fine-tuning and validation a. Run training scripts using the publisher’s data. b. Monitor training metrics (loss, accuracy) with tools like TensorBoard. c. Use a validation set or cross-validation to ensure accuracy, identify overfitting, etc. Deployment & Containerization 1. Isolated Hosting a. Each sub-LLM runs as endpoint, exposing an internal endpoint for inference requests. 2. Publisher Opt-Out a. Removing a publisher’s data is as simple as stopping and removing the corresponding container and data volume. Main LLM Orchestration 1. Anonymization Pipeline a. Before sub-LLM outputs reach the main LLM, identify and remove publisherspecific identifiers (e.g., page references, brand names) to maintain privacy. 2. Data Aggregation a. The main LLM receives sanitized text or structured data (like JSON) from multiple sub-LLMs. b. It may use ensemble logic (weighted voting or heuristics) for final grading. 3. Fault Tolerance a. If one sub-LLM is offline, the main LLM can still operate using the remaining sub-LLMs. Website Flow Enhancements Student Flow 1. Dashboard & Authentication a. Dashboard shows all enrolled subjects and completed tests. 2. Test Review a. Split-view layout: Left = student answers, Right = main LLM feedback. b. Split-view layout will describe reasons and questions the student got wrong c. “Request Review” button for disputes. 3. Yearly Reset a. Scheduled process clears older test records from the student’s dashboard. Teacher Flow 1. Dashboard & Class Management a. A consolidated view of classes, tests, and student performance. 2. Test Upload & Automated Grading a. Single upload point for PDFs (questions, answer keys, student submissions). b. System invokes PDF extraction, sub-LLM processing, and final main LLM grading. 3. Exception Handling & Quality Control a. Flags grading inconsistencies (e.g., incomplete or ambiguous answers). b. Teachers can manually override or request a second opinion (anonymously). 4. Finalization a. After teacher approval, results push to the school’s Student Information System Data Schematics A consistent data model is important for tracking publishers (teachers), tests, students, and grade results. These might be subject to change depending on what is required / not necessary and can be removed. Publisher Schematic Tests / Exams Schematic Students Schematic Grading Results Schematic

[0002] 6. Test Cases for Each User Flow 6.1 Student Flow 6.1.1 Viewing Graded Tests - Scenario: A student checks their dashboard to view a recent exam. - Preconditions: 1. Student is enrolled and has valid login credentials. 2. The exam has been uploaded and graded (no pending steps). - Steps: 1. Student logs in; system loads student’s data from Students schema. 2. Dashboard lists tests from Tests / Exams linked to the student’s classes. 3. Students click on a particular test_id. 4. The system retrieves graded data from Grading Results and displays it in a splitview layout. - Expected Result: 1. Correctly displays the final score, sub-question feedback from the main LLM. 2. “Request Review” button visible for each question. 6.1.2 Disputing a Grading Decision - Scenario: Student disputes question #3 on the exam. - Preconditions: 1. Exam data is in the Grading Results table. 2. Student is allowed to file disputes (the test is not yet “finalized”). - Steps: 1. Student clicks “Request Review” next to question #3 2. System logs the dispute in grade details (e.g., a disputes key in JSON). 3. Teacher’s dashboard updates to show the pending dispute. - Expected Result: - The dispute is tracked; teacher can see it in their to-do list. 6.2 Teacher Flow 6.2.1 Uploading a New Test - Scenario: A teacher uploads a Math exam PDF with an answer key. - Preconditions: 1. Teacher has valid login and the correct permissions. 2. publisher_id is available (the test is associated with a known publisher). - Steps: 1. Teacher navigates to “Upload Test.” 2. Uploads PDF(s) (exam questions, model answers, or student submissions). 3. Tests / Exams entry is created with test_id. 4. Extraction scripts parse the PDF, and sub-LLMs + main LLM handle initial grading. - Expected Result: - A new entry in the Tests / Exams table. - The system eventually populates Grading Results for each student submission. 6.2.2 Handling Exceptions & Quality Control - Scenario: The system flags 2 ambiguous questions after automated grading. - Preconditions: 1. Automated grading has run. 2. The sub-LLM or main LLM produced partial or uncertain feedback. - Steps: 1. Teacher sees an alert on the dashboard: “2 items need manual review.” 2. Teacher opens the list of flagged items, reviews the sub-question logic, and modifies or confirms grades. 3. Teacher finalizes the grading. - Expected Result: - Grading Results is updated with the teacher’s confirmed marks. - Students can see final feedback, unless further disputes arise. 6.2.3 Removing a Publisher’s Data - Scenario: Publisher A requests data deletion. - Preconditions: 1. Publisher A has an active sub-LLM and corresponding container. 2. System is designed for quick opt-out. - Steps: 1. Admin triggers the “Remove Publisher” process in a management interface. 2. The system stops subllm_publisherA container and deletes associated volumes. 3. Only anonymized data remains in the main LLM. - Expected Result: - publisher_id A is set to INACTIVE. - The main LLM remains fully operational without direct references to Publisher A’s dataset. Definitions: Prompt Structure: An AI prompt structure comprises context, task, and constraints for a large language model (LLM) to interpret. It enables the model to analyze text accurately within specified parameters Noisy Data: Data that is corrupt, unstructured, or otherwise unreadable by machines. This encompasses data that has been improperly saved, received, or modified, making it unreadable for the originating program. In other words, noisy data is any data that’s been received, saved, or modified in such a way that it’s prevents a computer program to read or use. Optical Recognition Software will run into the problem that a text string derived from an image, such as a visual representation of a document, may contain a string of letters and numbers that together fail to form a coherent word or sentence. Such data is noisy data.

Claims

CLAIMS 1. A computer-implemented method for analyzing and quantifying the quality of a first user input (1UI) in relation to a second user input (2UI) by comparison to a third user input (3UI) in a computer system, wherein the second user input defines at least one problem statement, the third user input defines at least one solution statement to the at least one problem statement, the method comprising: - providing a graphical user interface for receiving said first, second and third user inputs; - generating at least one prompt structure based on the first, second and third user inputs, such as a prompt structure for each of the first, second and third user inputs, wherein the at least one prompt structure comprises instructions to quantify the quality of a first user input, wherein the quantified quality at least in part reflects the degree to which the first user input corresponds to the at least one solution statement as defined by the third user, such as in its completeness, and wherein the quantified quality further at least in part reflects its correctness in relation to the at least one problem statement as defined by the second user input; - processing the at least one prompt structure using at least one large language model; and - outputting the quantified quality of the first user input from the at least one large language model, such that it is accessible to a user within the graphical user interface.

2. The method according to claim 1, wherein the graphical user interface is designed for uploading digital representations of physical documents thereby receiving at least the first user input, preferably the first, second and third user inputs, wherein the method comprises the step of:- performing optical character recognition on the uploaded documents to obtain textual information, such as mathematical formulas, for generating the at least one prompt structure.

3. The method according to claim 2, comprising the step of: - determining the accuracy with which the textual information is obtained through optical character recognition for each uploaded document; and - generating a warning within the graphical user interface in response to determining that the accuracy is below a predefined threshold.

4. The method according to claim 3, comprising the step of: - preventing the user from generating the at least one prompt in response to determining that the accuracy is below a predefined threshold until the user inputs are substituted, such as by replacement uploads, and determined to satisfy the predefined threshold.

5. The method according to any one of claims 1-4, comprising eliminating noisy data from the at least one prompt or from the first, second and third user inputs, such as using the at least one large language model.

6. The method according to any one of claims 1-5, wherein the outputting of the quantified quality of the first user input also comprises a motivation output for the assigned quantified quality, wherein the motivation output articulates the motivation of the at least one large language model for arriving at the quantified quality.

7. The method according to any one of claims 1-6, wherein the at least one large language model comprises a plurality of large language models in sequence, wherein the output of a first model of the plurality becomes the input for a second model of the plurality to enrich the context or information available for the second model in the sequence.

8. The method according to claim 7, wherein the prompt structure is generated by the first large language model and wherein the prompt structure is processed by the second large language model to output the quantified quality, and wherein preferably the first and second large language models are more suitable, such as more optimized, than the other for their mutually different task within the sequence.

9. The method according to claim 7 or 8, wherein the first large language model is used to select the second large language model from a list of large language models which it determines to be most suitable for the processing of the prompt structure.

10. The method according to claim 7 or 8, wherein the graphical user interface allows a user to add a fourth user input pertaining to a task-type, such as by selection from a list of different tasks, such as more specifically selecting between grading a mathematical exercise and grading a non- mathematical exercise, and wherein the at least one large language model is configured, such as from a sequence of a suitable large language models, in response to the fourth user input, preferably so as to be optimized for the selected task- type.

11. The method according to any one of claims 7-10, wherein first LLM comprises a plurality of sub-LLMs trained on mutually different datasets, and wherein preferably, when one sub-LLM of the plurality of sub-LLMs is becomes unavailable, the second LLM operates using at least one of the remaining sub- LLMs.

12. A system comprising a computer, comprising a human- machine interface, that is preprogrammed to perform the method according to any one of claims 1-11.

13. The system according to claim 11, wherein the computer comprises a distant server for processing the at least one prompt structure and optionally for generating the at least one prompt structure.

14. The system according to claim 13, wherein the human-machine interface is provided as a desktop, laptop, tablet or smartphone, and is preprogrammed to provide the graphical user interface, and wherein the human-machine interface and distant server are designed and programmed for being mutually communicatively connectable.

15. An application library comprising a downloadable program, which when installed on a digital device allows a user to perform the method according to any one of claims 1-11 on said device or on a combination of server and device.

Citation Information

Patent Citations

  • Semi-automated evaluation of long answer exams

    US20220383767A1

  • Automated Evaluation of Free-Form Answers and Generation of Actionable Feedback to Multidimensional Reasoning Questions

    US20240054909A1

Cited By

  • System and method for optimizing rules using a machine learning model

    US20260080328A1