Automatic evaluation of free form answers to multi-dimensional inference questions and generation of actionable feedback

Through the tokenization technology of machine learning classifiers and parse tree formats, the problem of insufficient consistency and feedback of computerized learning platforms when evaluating free-form text answers is solved, and more accurate evaluation and personalized learning resource recommendations are achieved.

CN120390951APending Publication Date: 2025-07-29BRAINPOP IP LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073143.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-15
Filing Date
2023-08-15
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

When evaluating students' free-form text answers, existing computerized learning platforms have problems such as having a large difference from the results of manual evaluation, difficulty in providing timely and actionable feedback, and accurate recommendation of learning resources.

Method used

A machine learning classifier is used to identify and evaluate the logical connections of various parts of student answers, and the free-form text answers to multidimensional inference problems are automatically evaluated through training data transformation and tokenization formats, combined with parse trees and n-member grammatical analysis.

Benefits of technology

It realizes more accurate answer evaluation, provides timely feedback, and can recommend personalized learning resources based on the evaluation results, improving the consistency and effectiveness of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390951A_ABST
    Figure CN120390951A_ABST
Patent Text Reader

Abstract

In an illustrative embodiment, a method and system for automatically evaluating free-form text answer content, a method includes obtaining textual portions that answer a multi-dimensional inference question, analyzing portion content of each portion using AI model (s), analyzing logical connections between the textual portions, and calculating score (s) corresponding to a free form answer based on the analysis.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 397,971, filed on Aug. 15, 2022, entitled “Automated Evaluation of Free-Form Answers to Multidimensional Reasoning Questions”. The entire content of the above application is hereby incorporated by reference in its entirety. Background Art

[0003] Students typically interact with computerized learning platforms by typing text answers into response fields in a graphical user interface. To evaluate a student's response, various platforms may automatically apply a series of scoring criteria to the input text. However, the scores or grades assigned by computerized learning platforms based on such scoring criteria may vary significantly from those assigned manually (e.g., by a teacher). For example, depending on the ability level of the child learner, answers are typically “low-quality” in terms of format (e.g., typographical errors, irregular use of punctuation, unorthodox grammar, misuse of words with related meanings, etc.). In addition, scoring criteria typically assess the use of language (e.g., correctness and sufficient complexity of grammar or vocabulary), while for grading or scoring answers related to various topics, the appropriate emphasis is on requiring the answer to be relevant to the appropriate topic and to present a correct logical structure.

[0004] To overcome the above limitations, the inventors recognized the need for an alternative method for analyzing via scoring criteria and providing timely and actionable feedback to students and teachers. In addition, the inventors recognized the need for a new evaluation mechanism to accurately apply automated evaluation to recommend further learning resources (e.g., personalized learning) and / or cluster learners based on proficiency. Summary of the Invention

[0005] In one aspect, the present disclosure relates to systems and methods for applying machine learning to automatically evaluate the content of student answers to multidimensional reasoning questions formatted in a formal response architecture, such as the claim-evidence-reasoning (CER) structure in the scientific field, mathematical reasoning, or argumentative writing in English language arts. Machine learning classifiers can be trained to identify the relative strength / weakness of the consistency of each part of a student answer (e.g., claim, evidence, and reasoning parts). In addition, machine learning classifiers can be trained to evaluate the logical connections between the claim part and the evidence part, and between the evidence part and the reasoning part. For example, an answer can be constructed as a “mini-essay” involving one or more sentences in each part of the CER, mathematical reasoning, or argumentative writing structure.

[0006] In one aspect, the present disclosure relates to systems and methods for developing training data for training a machine learning classifier to automatically evaluate the content of free-form text answers to multi-dimensional reasoning problems. Converting the text answers into training data can include converting free-form text answers into tokens that represent at least a portion of the content of the free-form text answers. For example, certain tokens, such as declarations and / or punctuation, can be discarded from the analysis. The tokens can be arranged in a parse tree format. Attributes or characteristics of individual tokens, such as word attributes of individual tokens and / or dependencies between tokens, can be added to enhance certain individual tokens. Additionally, the enhanced set of tokens can be converted from the parse tree format into one or more syntactic n-gram forms. The training data can also include metrics representing aspects of the original tokens, enhanced tokens, and / or parse tree format.

[0007] In one aspect, the present disclosure relates to systems and methods for training a machine learning classifier to automatically evaluate the content of free-form text answers to multi-dimensional reasoning problems. The ground truth annotation data can include example answers designed and scored by professionals according to standard scoring criteria. The training can be supplemented by automatically scoring the students' answers, identifying those answers that receive very high and / or very low scores during the automated machine learning process, and queuing those free-form answers for manual scoring. The manually scored answers can be provided as additional training data to optimize the training of the machine learning classifier. For example, very high and / or very low scores can include all free-form answers that are assigned perfect scores and zero scores by the automated machine learning-based evaluation. To avoid incorrect answers (e.g., incomplete answers) from being evaluated as zero scores and then fed into manual evaluation, a preliminary evaluation of their completeness can be performed by a first automated process before each free-form answer is submitted to the automated machine learning evaluation.

[0008] The foregoing general description of the illustrative embodiments and the following detailed description are merely exemplary aspects of the teachings of the present disclosure and are not restrictive. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings incorporated in and constituting a part of this specification illustrate one or more embodiments and, together with the specification, explain these embodiments. The drawings are not necessarily to scale. The dimensions of any values illustrated in the drawings and diagrams are for illustrative purposes only and may or may not represent actual or preferred values or dimensions. Where applicable, some or all features may not be illustrated to assist in describing the underlying features. In the drawings:

[0010] Figure 1A and Figure 1B is a block diagram of an example platform for automatically evaluating free-form answers to multi-dimensional reasoning problems;

[0011] Figure 2 shows an example parse tree that parses a sentence into word relationships;

[0012] Figure 3 , Figure 4A , Figure 4B , Figure 4C , Figure 5A , Figure 5B , Figure 6 , Figure 7A and Figure 7B shows a flowchart of an example method for automatically evaluating free-form answers to multi-dimensional reasoning questions;

[0013] Figure 8A and Figure 8B shows a screenshot of an example user interface display for entering a text response to a multi-dimensional reasoning question;

[0014] Figure 9 is a flowchart of an example process for training a machine learning model to automatically evaluate vectorized free-form answers to multi-dimensional reasoning questions; and

[0015] Figure 10 is a flowchart of an example process for tuning an artificial intelligence model to automatically evaluate free-form answers to multi-dimensional reasoning questions. Detailed Description

[0016] The following description set forth in conjunction with the accompanying drawings is intended to describe various illustrative embodiments of the disclosed subject matter. Specific features and functions are described in conjunction with each illustrative embodiment; however, it will be apparent to those skilled in the art that the disclosed embodiments can be practiced even without each of those specific features and functions.

[0017] Reference throughout the specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosed subject matter. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout the specification are not necessarily all referring to the same embodiment. Moreover, in one or more embodiments, the particular features, structures, or characteristics may be combined in any suitable manner. Additionally, embodiments of the disclosed subject matter are intended to cover modifications and variations thereof.

[0018] It should be noted that, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" used in this specification and the appended claims include plural referents. That is, unless otherwise clearly specified, as used herein, the words "a", "an", "the", etc. have the meaning of "one or more". In addition, it should be understood that terms such as "left", "right", "top", "bottom", "front", "rear", "side", "height", "length", "width", "upper", "lower", "inner", "outer", "inside", "outside", etc. that may be used herein only describe reference points and do not necessarily limit the embodiments of the present disclosure to any particular orientation or configuration. In addition, terms such as "first", "second", "third", etc. only identify one of the multiple parts, components, steps, operations, functions, and / or reference points disclosed herein, and likewise do not necessarily limit the embodiments of the present disclosure to any particular configuration or orientation.

[0019] In addition, terms such as "about", "approximately", "close to", "slight variation", and similar terms generally refer to a range that, in certain embodiments, includes the indicated value with an error margin of within 20%, within 10%, or in certain preferred embodiments within 5%, and includes any value between these error margins.

[0020] All functions described in connection with one embodiment are intended to be applicable to the additional embodiments described below, unless expressly stated or the feature or function is incompatible with the additional embodiments. For example, in the case where a given feature or function is clearly described in connection with one embodiment but not expressly mentioned in connection with an alternative embodiment, it should be understood that the inventor intends that the feature or function can be deployed, utilized, or implemented in connection with the alternative embodiment, unless the feature or function is incompatible with the alternative embodiment.

[0021] Figure 1A is a block diagram of an example platform 100 that includes an automated assessment system 102a for automatically assessing free-form answers to multi-dimensional reasoning questions. The free-form answers can be submitted by students interacting with an online learning platform via a student device 104.

[0022] In some embodiments, the automated assessment system 102a includes a student graphical user interface (GUI) engine 112 for providing learning resources to the student device 104. The student GUI engine 112 can prepare instructions for presenting an interactive student GUI view at the display of the student device. The interactive GUI can include multi-dimensional reasoning questions and / or essay topic descriptions that require answers involving multi-dimensional reasoning, and one or more text input fields for accepting free-form answers. For example, an interactive student GUI view can be developed to encourage the student to think about the learning content provided by the online learning platform. In addition to the text input fields for accepting free-form answers, the GUI view can also include other student input fields, such as, in some examples, multiple-choice questions, matching exercises, word or phrase input fields for simple questions, and / or binary (e.g., yes / no) answer selection fields. The student GUI engine 112 can provide the free-form answers to the answer vectorization engine 114 to convert the text into a format for assessment.

[0023] In some embodiments, the automated grading system 102a includes an answer vectorization engine 114 for formatting free-form answers into one or more vector formats for machine learning analysis. In some examples, the vector formats can include various n-gram formats, which are described in more detail below. In particular, free-form answers can be parsed into a set of directed trees, also known as a directed graph. In a directed graph G, each weakly connected component is a directed tree, where no two vertices are connected by more than one path. Each tree (each weakly connected component of the directed graph G) of the directed graph G represents a sentence. The graph vertices V(G) are words and punctuation marks, and the edges E(G) of the graph are grammatical relations that can be of various types. Thus, each edge has an attribute "dependency type".

[0024] Moving on to Figure 2 , the example parse tree 200 shows the mapping of the sentence 202 to word relationships. Figure 2 The parse tree 200 in

[0025] The most basic way to analyze a sentence (such as sentence 202) is to treat each word or punctuation mark as a token, and the entire text as an unordered list of tokens (commonly referred to as a "bag of words"). Additionally, as part of the automatic cleaning of Chinese text, punctuation marks are typically discarded. Thus, in the "bag of words" approach, the tokens of sentence 202 include the representations of the following words: "is" (or the lemma "is"), "you", "certain", "say" (or the lemma "say"), "rabbit", "even more", "surprised".

[0026] In a method commonly referred to as generating n-grams, the tokens can be converted into sequences or groupings of tokens. In some embodiments, an integer n represents the number of "bags of words" of simple (singular) tokens, and n tokens are converted into a sequence. For example, for n = 2 (bigrams), the bigram tokens of sentence 202 include ["is", "you"], ["you", "certain"], ["certain", "say"], etc. In some embodiments, the token groupings are non-sequential and are called skip-grams, where the n-gram tokens are formed by skipping a specified number of intermediate words between the words in the text (e.g., sentence 202). In one example, the skip-gram form of bigrams can be generated by selecting every other word (e.g., ["is", "certain"], ["you", "say"], ["certain", "rabbit"], etc.).

[0027] The classical n-grams in the above examples can be represented as vectors. For example, if an n-gram is a sequence of words w 1 , w 2 , … w n , then this n-gram is represented by the concatenated vector v = [V(w 1 ), V(w 2 ), … V(w n )]. As shown, n-grams can provide insights into the grammatical structure of a sentence. However, classical n-grams rely on word order. Even if the grammar is perfect, neighboring words are not necessarily words with a direct grammatical relationship. Additionally, the set of word groupings in n-gram tokens changes with the sentence structure. In the example, sentence 202 can be constructed in a different structure but with the same meaning by writing it as "Rabbit, even more surprised, said, 'Are you certain?'".

[0028] Therefore, in some embodiments, the analysis does not only use classical n-grams to analyze responses, or in addition to applying n-grams, it also includes using a parse tree format (e.g., Figure 2The n-grams generated from the parse tree 200 shown, rather than analyzing according to the word order of the sentence 202. The inventor refers to this type of tokenized format as "syntactic n-grams". Syntactic n-grams are represented by tokens formed by any two vertices of G connected by an edge. The dependency types between the simple tokens of syntactic n-grams are included as part or section of the syntactic n-grams. For example, the sentence 202 includes the following syntactic bigram tokens: ["is" subject relation "you"], ["is" complement relation "sure"], ["said" complement clause relation "is"], etc. In addition, syntactic trigrams can be formed by any subgraph consisting of three vertices and the paths connecting them, such as ["rabbit" modifier relation "surprised" modifier adverb relation "also"]. Similarly, higher-order n-grams (such as 4-grams and 5-grams) can be constructed using this general subgraph format.

[0029] In addition to generating more descriptive tokens (syntactic n-gram tokens) from the text, the parse tree format also provides other insights into the language structure. Graphical metrics can be derived from the parse tree format and incorporated as part of the text analysis. In some examples, the graphical metrics can include the number of leaf nodes (i.e., vertices with no outgoing edges), the average number of outgoing edges per vertex (branching factor), and the height of the tree (the length of the longest path). In addition, more complex metrics can be derived from graph theory and these metrics can also be used as variables for analysis.

[0030] Returning to Figure 1A , in some embodiments, the answer vectorization engine 114, when formatting the free-form answer into a vector format (one or more), may adjust the received text for consistency, such as spelling consistency (e.g., spelling error correction, converting British / American spellings to a single style, etc.), verb tense consistency, and / or number format (e.g., written language vs. integer form), to allow recognition by the trained machine learning model. The answer vectorization engine 114 may also remove parts of the answer, such as punctuation and / or determiners (e.g., "the", "this", etc.). The answer vectorization engine 114 may provide the vector format (one or more) of the free-form answer to the automatic evaluation engine 116a for content evaluation. In some embodiments, the answer vectorization engine 114 stores the vector format (one or more) of the free-form answer in a non-volatile storage area as the vectorized answer 148.

[0031] In some embodiments, the answer metric engine 130 calculates a set of answer metrics 152 associated with the vectorized answer 148 generated by the answer vectorization engine 114. In cases where the answer vectorization engine 114 stores multiple formats of free-form answers for automated evaluation, the answer metrics 152 can correspond to one or more forms of the vectorized answer. An example of a metric is given in conjunction with Figure 4A the method 320a in

[0032] In some embodiments, the automated evaluation engine 116a applies the trained machine learning model 108 to evaluate the vectorized answer 148 generated by the answer vectorization engine 114. At least a portion of the vectorized answer 148 can be evaluated based on the answer metrics 152 calculated by the answer metric engine 130.

[0033] In some embodiments, the automated evaluation engine 116a applies one or more machine learning models 108 to one or more syntactic n-gram forms of the vectorized answer on a part-by-part basis to evaluate the content of a given part. For example, when evaluating part content using syntactic n-gram forms, sentences provided by a student related to a particular part can be analyzed according to the goal or purpose of each part (such as claims, evidence, or reasoning, etc.) to automatically evaluate the content of that part based on the gist of its text, rather than simply focusing on the correctness of the answer and / or the correctness of the writing style (such as spelling, grammar, etc.). For example, the portion of the machine learning model 108 that is intended to evaluate each part can be trained (e.g., by the machine learning model training engine 126) using parts from the historical graded answer set 140 on a part-type-by-part-type basis. This set of historical graded answers contains breakdown scores, including one or more part content scores.

[0034] In some embodiments, the automated evaluation engine 116a applies one or more machine learning models 108 to one or more syntactic n-gram forms of the vectorized answer as a complete set of parts to evaluate the logical connections between the parts of the free-form answer. For example, the portion of the machine learning model 108 that is used to evaluate logical connections can be trained (e.g., by the machine learning model training engine 126) using a portion of the historical scored answer set 140 that has score breakdowns including at least one logical connection score.

[0035] In some embodiments, the automatic evaluation engine 116a applies one or more machine learning (“ML”) models 108 to one or more classical n-gram forms of the vectorized answer to evaluate the style and / or grammatical quality of the content of the free-form answer. For example, the portion of the machine learning model 108 that is designed to evaluate the style and / or grammatical quality of the content can be trained (e.g., by the machine learning model training engine 126) using a portion of the historical scored answer set 140 that has a scoring breakdown including style and / or grammatical scores.

[0036] In some embodiments, the automatic evaluation engine 116a evaluates at least a portion of the vectorized answer 148 according to one or more evaluation rules 154. In some examples, the evaluation rules 154 can relate to the type of topic, the student's level of learning (e.g., based on the student statistics 144), and / or the particular question being answered. In one example, younger students and / or students with a lower level of learning can be evaluated based on less stringent evaluation rules 154 (e.g., content and logical flow, but not style and grammatical analysis), while older students and / or students with a higher level of learning can be evaluated based on more stringent evaluation rules 154. In some embodiments, the answer vectorization engine 114 can use the evaluation rules 154 to generate vectors suitable for a particular level of analysis (e.g., classical n-gram versus non-classical n-gram analysis) to be applied by the automatic evaluation engine 116a. In some embodiments, the output of the ML model 108 is stored as the ML analysis result 156.

[0037] In some embodiments, the score calculation engine 118 obtains the ML analysis result 156 output from the trained machine learning model 108 and grades the student's answer. The score calculation engine 118 can generate one or more scores. For example, in some examples, a total score, a content score for each part representing the match between the expected gist of a particular part and the text provided by the learner, at least one topic correctness score representing whether the text provided by the learner contains the correct answer and / or the topic answer, a logical connection score representing the logical flow between parts of a free-form answer, and / or a style and grammar score. In some examples, the scores can include a graded score (e.g., A+ to F), a percentage score (e.g., 0% to 100%), and / or a relative performance rating (e.g., excellent, satisfactory, unsatisfactory, incomplete, etc.). In some embodiments, the score calculation engine 118 calculates one or more scores based on the automated scoring rules 142. For example, the automated scoring rules 142 can include weights applied when combining the output of the trained machine learning model 108, the output types of each type of score combined to generate one or more scores for each answer, and / or rules for normalizing or adjusting the scores based on group performance and / or expected group performance (e.g., grading on a curve). The score calculation engine 118 can provide the score(s) to the student GUI engine 112 to present the results to the student and / or the teacher GUI engine 128 to present the results to the student's teacher. In another example, the score calculation engine 118 can provide the score(s) to the learning resource recommendation engine 134 for recommending the next learning material.

[0038] In some embodiments, depending on the score, free-form answers can be presented for manual scoring. For example, the manual scoring rules 146 can be applied to the score(s) calculated by the score calculation engine 118, and for scores above an upper threshold and / or below a lower threshold, a manual review may be submitted. For example, the manual scoring GUI engine 122 can present the student free-form answer to a teacher or other learning professional for manual evaluation and scoring. Additionally, the manually scored answers can be provided to the machine learning model training engine 126 as additional examples of very good (or very bad) examples of student free-form answers for updating one or more of the trained ML models 108.

[0039] In some embodiments, a portion of the scores generated by the score calculation engine 118 is used by the learning resource recommendation engine 134 to determine the next learning material to present to a student. For example, based on the student's proficiency in a topic (as demonstrated by free-form answers), the learning resource recommendation engine 134 may recommend a more in-depth analysis of the topic or a transition to a new topic. Conversely, based on the student's lack of proficiency in a topic (as demonstrated by free-form answers), the learning resource recommendation engine 134 may recommend one or more learning resources aimed at strengthening the understanding of the current topic. In some examples, the recommended learning resources may include learning resources that the student has not yet seen, learning resources on which the student has spent little time, and / or learning resources in a format that the student likes.

[0040] In some embodiments, the student clustering engine 120 analyzes the scores of students on free-form answers and / or other interactions with the learning platform to cluster groups of students by ability. For example, the student clustering engine 120 may generate groupings of students based on a student clustering rule set 150. In some examples, the student clustering rules may identify the total number of groupings, the score or level cutoffs for each grouping, and / or the learning levels / proficiencies associated with each grouping. In an illustrative example, the student clustering engine 120 may cluster a population of students into a set of advanced students, a set of proficient students, and a set of novice students related to a particular subject and / or learning topic area, at least in part based on scores generated by the score calculation engine 118 with respect to one or more free-form answers submitted by each student in the population of students. In some examples, student clustering can be used to generate comparison metrics between populations of students (e.g., different schools, different geographic regions, etc.), provide an overall assessment of class performance to a teacher (e.g., via the teacher GUI engine 128), and / or provide further input to the learning resource recommendation engine 134 regarding appropriate materials to present to different groups of students.

[0041] In some embodiments, the machine learning model training engine 126 trains one or more machine learning models to analyze each form of various forms of vectorized free-form answers, including forms of at least one syntactic n-gram generated by the vectorization engine 114, to automatically evaluate the quality of at least the subject matter of the free-form answers. For example, the machine learning model training engine 126 can obtain the historical scored answer set 140 from the storage area 110a for analysis. In some embodiments, the historical scored answers 140 represent various subject areas, student levels (e.g., age, grade, and / or proficiency), and scoring results. In some embodiments, a portion of the historical scored answers 140 is manually scored by subject matter experts or other professionals trained in applying a scoring criterion (e.g., via the manual scoring GUI engine 122), the scoring criterion including at least one criterion associated with the multi-part form used for each historical scored answer. In some embodiments, a portion of the historical scored answers 140 is example answers generated by subject matter experts or other professionals and trained as excellent free-form answers with an appropriate scoring criterion. In training the ML model 108, the ML model training engine 126 can incorporate one or more answer metrics associated with a particular vectorized form of the free-form answer, such as the answer metric 152 generated by the answer metric engine 130. The ML model training engine 126 can store the trained ML model 108 for use by the automatic evaluation engine 116a.

[0042] Go to Figure 1B , the block diagram of the example platform 160 includes an automatic evaluation system 102b, similar to Figure 1A the automatic evaluation system 102a in Figure 1A for automatically evaluating free-form answers to multi-dimensional reasoning problems. Different from the platform 100 of

[0043] In some embodiments, to evaluate free-form answers using one or more AI models 170, at least a portion of the free-form answers can be formatted using an answer formatting engine 164. For example, the answer formatting engine 164 can tokenize each part of the free-form answer. In another example, the answer formatting engine 164 can adjust the received text for consistency, such as consistency in spelling (e.g., spelling error correction, converting British / American spellings to a single style, etc.), verb tense, and / or formatting of numbers (e.g., written language versus integer form), to allow for consistent processing by one or more AI models 170. In another example, the answer formatting engine 164 can remove a portion of the answer, such as certain punctuation marks and / or special characters. For example, the formatted answer can be stored as one or more formatted answers 168.

[0044] In some embodiments, the formatting portion depends on the type of AI model 170 used for a particular answer. For example, a model selection engine 178 can select a particular AI model 170 for a student's free-form answer based on student statistics 144 and / or the context of the student's answer. In some examples, the context can include the topic, level, question style, and / or grading rules (e.g., auto-grading rules 142) applicable to the student's free-form answer.

[0045] In some embodiments, one or more formatted answers 168 are provided by an auto-evaluation engine 116b to one or more AI models 170. The auto-evaluation engine 116b can communicate with one or more AI models 170 via an engineered model prompt 162 to instruct the one or more AI models 170 to analyze the one or more formatted answers 168. For example, the auto-evaluation engine 116b can provide an engineered model prompt 162 corresponding to the desired grading rules 142, evaluation rules 154, and / or other grading and / or evaluation criteria. Additionally, the auto-evaluation engine 116b can provide an engineered model prompt 162 corresponding to specific student statistics 144, such as, in some examples, student age, student grade level, and / or student skill level.

[0046] In some embodiments, a prompt selection engine 180 selects a particular set of engineered model prompts 162 for requesting an evaluation of a student's free-form answer (or its formatted version) from one of the AI models 170. For example, the prompt selection engine 180 can match student statistics 144, auto-grading rules 142, and / or auto-evaluation rules 154 with certain engineered model prompts 162 to obtain an evaluation consistent with the context of the student's answer.

[0047] In some embodiments, the engineering model prompt 162 is generated at least in part by the prompt engineering engine 176. For example, the prompt engineering engine 176 can iteratively adjust the prompt language to predictably achieve an automated score that is consistent with a human score using a set of pre-scored training examples (e.g., historical scored answers 140). In an illustrative example, the prompt engineering engine can insert the pre-scored training examples into the prompt for scoring guidance.

[0048] In some embodiments, the automated evaluation engine 116b selects certain AI models 170 to perform the evaluation. In some examples, the selection can be at least partially based on the training of the (one or more) specific AI models 170 and / or the fine-tuning of the (one or more) specific AI models 170. For example, certain AI models 170 can be trained and / or fine-tuned such that these AI models are adapted to evaluate free-form answers and / or formatted answers 168 according to desired scoring and / or evaluation rules. In other examples, certain AI models 170 can be trained and / or fine-tuned to perform evaluations at different ability levels (e.g., skill level, grade level, etc.), in different languages, and / or taking into account different subject areas (e.g., science, history, etc.).

[0049] In some embodiments, certain base AI models 170 are fine-tuned using the model tuning engine 174 to perform the evaluation of free-form answers. For example, the model tuning engine 174 can provide the historical scored answers 140 to certain base AI models 170 to adjust the training of the base model for analyzing the style of free-form answers submitted by students to the automated evaluation system 102b (e.g., as described with respect to Figure 8A ).

[0050] In some implementations, the automated evaluation engine 116b receives the artificial intelligence (AI) analysis results 166 including at least one score from the selected AI models 170. For example, the at least one score can be provided in the format specified by the automated scoring rules 142 (refer to Figure 1A for description). In some embodiments, a portion of the AI analysis results 166 includes text-based feedback regarding the reasoning behind the automated scores generated by the selected AI model(s). In some examples, the feedback portion can include tips, suggestions, pointers to specific parts of the student's answer without revealing the actual correct / desired answer, and / or a step-by-step explanation of the teaching content behind the answer. For example, the feedback portion can be provided by the automated feedback engine 172 for presenting feedback to the student and / or the instructor (e.g., via the student device 104 and / or the teacher device 106). For example, the feedback provided to the student can include actionable suggestions and / or tips for improving the final score of the free-form answer.

[0051] Figure 3 FIG. 300 is a flow chart showing an example method for automatically evaluating free-form answers to multi-dimensional reasoning questions. For example, portions of method 300 may be performed by answer vectorization engine 114 and / or Figure 1A automated evaluation engine 116a in Figure 1B or one of automated evaluation engines 116b in

[0052] In some embodiments, method 300 begins by obtaining a free-form answer to a multi-dimensional reasoning question submitted by a student (302). For example, the free-form answer may include one or more sentences. In some embodiments, the free-form answer is formatted in a particular structure, such as a claim, evidence, reasoning (CER) framework or a claim, evidence, reasoning, summary (CERC) framework based on the scientific method, a restate, explain, example (REX) model for answering short answer questions, an introduction, body, conclusion framework, an answer, proof, explanation (APE) framework, and / or a topic sentence, specific details, comment, ending / conclusion sentence (TS / CD / CM / CS) framework. The free-form answer may be obtained electronically from an answer submission form presented to the student.

[0053] In some embodiments, the sufficiency of the student's free-form answer is evaluated (301). For example, this evaluation may be performed by answer sufficiency engine 132. Turning to Figure 6 method 301 of Figure 3 in some embodiments, a student free-form answer to a multi-dimensional reasoning question is obtained (303). For example, the answer may be obtained from

[0054] In some embodiments, if the free-form answer has not been divided into structured parts (305), then a partial structure corresponding to the multi-dimensional reasoning question is identified (307). As previously described, these parts may be structured in various formats, such as the CER format for scientific method style answers or the introduction, body, conclusion standard framework.

[0055] In some embodiments, the answer is submitted in a graphical user interface having text input fields separated by parts in a structured format. For example, turning to Figure 8A a first example screenshot 400 of a graphical user interface presented on a display 404 of device 410 includes a set of text input fields 402 that are separated into a claim text input field 402a, an evidence text input field 402b, and a reasoning text input field 402c. The student may prepare a free-form answer by entering text into text input fields 402. After answering is complete, the student may select a submit control 406 to submit the contents of text input fields 402 as an answer prepared in CER structure format.

[0056] Return to Figure 6 ,if the answer (305) has not been received in a form that has already been divided into structured parts, then in some embodiments, identify partial structures (307) corresponding to the multi-dimensional reasoning problem. In some examples, information associated with the learning unit related to the free-form answer itself, the student who submitted the free-form answer, and / or the question being answered in the free-form answer can be analyzed to determine the specific structured format to be applied. In one example, the question associated with the free-form answer can be associated with a specific type of structured format. In another example, the student's learning level and / or age (based on Figure 1A and Figure 1B student statistics 144) can be associated with a specific type of structured format. In a third example, the subject area (e.g., history, science, literature, geography, etc.) can be associated with a specific type of structured format.

[0057] In some embodiments, use the identified partial structures to analyze the free-form answer to automatically divide the answer into the respective parts of the partial structures (309).

[0058] In some embodiments, if parts of the free-form answer are missing or incomplete (311), then present feedback to the student to improve the free-form answer (313). For example, each part can be reviewed to see if it contains some text. Additionally, the spelling and / or grammar of each part can be reviewed ( Figure 8A the last sentence in the evidence part 402b is a sentence fragment without punctuation). In some embodiments, each part is reviewed to confirm that it contains the language representative of that part (e.g., as Figure 8A shown, the sentence in claim part 402a should contain a claim, not an instruction).

[0059] In some embodiments, go to Figure 8B and provide general feedback to the user, giving the student the opportunity to continue to improve the free-form answer. As shown, the pop-up window 426 that overrays the graphical user interface 404 displays "We think what you submitted could be improved" and provides the student with two control options: the "Submit as is" button 424, which, when selected, continues to submit the free-form answer for evaluation, and the "Continue to improve" button 422, which, when selected, returns the student to the graphical user interface 404 to continue to improve one or more parts 402 of the free-form answer. In other embodiments, the student can be directed to specific parts (e.g., the identification of parts that do not contain text, or a single sentence without punctuation as shown in claim part 402a).

[0060] Return to Figure 6, in some embodiments, if the analysis does not identify any missing and / or incomplete parts (311), a free-form answer is provided for evaluation (315). For example, the evaluation can be performed by the Figure 1A automatic evaluation engine 116a of Figure 1B or the automatic evaluation engine 116b of

[0061] Although presented as a specific series of operations, in other embodiments, method 301 includes more or fewer operations. For example, in some embodiments, the free-form answer is further analyzed for spelling and / or grammar. In some embodiments, certain operations of method 301 are performed simultaneously and / or in a different order. For example, before analyzing the free-form answer to divide it into multiple parts (309), an analysis of the sufficiency of the length, the use of punctuation at the end of the submission (e.g., reconfirming whether the student inadvertently selected the submission during the drafting process), or other analysis of the free-form answer can be performed, and feedback related to the perceived insufficiency of the answer can be presented to the student (313). In some embodiments, at least a part of method 301 is performed by one or more AI models. For example, feedback (313) can be provided to the student through prompt engineering to give feasible suggestions on how to improve the sufficiency of the response. Other modifications of method 301 are possible.

[0062] Returning to Figure 3 , in some embodiments, if it is determined that the answer is insufficient (305), the student is provided with an opportunity to continue to improve the answer, and method 300 returns to waiting to obtain an updated answer (302).

[0063] In some embodiments, if the free-form answer is determined to be sufficient (305), a spelling and / or grammar review algorithm is applied to the free-form answer to correct the student's errors and make the answer conform to a standardized format (304). Misspelled words can be corrected. Words with multiple alternative spellings (e.g., cancel / cancelled) can be converted to the standard format. The format of numbers (e.g., written language vs. integer form) can be adjusted to the standard format. When applying the spelling and / or grammar review algorithm, in some embodiments, the text of the free-form answer is adjusted to more closely match the answers used to train the machine learning model, so that higher-quality results can be obtained from the automated machine learning analysis.

[0064] In some embodiments, automated analysis is applied to the free-form answer to evaluate its content (320). For example, the Figures 4A to 4C method 320a of Figure 5A and Figure 5B method 320b of

[0065] Go toFigure 4A , in some embodiments, the text (303) of each part of a free-form answer to a multi-dimensional reasoning question is obtained. For example, it can be obtained from the Figure 3 method 300.

[0066] In some embodiments, the free-form answer is converted to a parse tree format (306). For example, this conversion can be performed as described with respect to the answer vectorization engine 114. In some examples, a parse graph G is obtained, where each primitive token (e.g., word, punctuation, etc.) is a vertex. If two vertices are connected by an edge, the source vertex can be referred to as the head of the target vertex. Due to the tree format of the graph, no vertex will have multiple heads. Vertices without a head are possible - such vertices can be referred to as roots (e.g., vertices with an in-degree of 0). Vertices with an out-degree of 0 can be referred to as leaves.

[0067] In some embodiments, at least a portion of the primitive tokens are augmented by one or more attributes (308). For each edge, an attribute of "dependency type" can be assigned, such as the dependency types 206a through 206h described with respect to Figure 2 . For each vertex, multiple attributes can be assigned, which in some examples include the wordform presented in the text, the lemma of the word (e.g., as described by the lemma 208 with respect to Figure 2 ), the part of speech of the word (e.g., verb, adjective, adverb, noun, etc., as described by the part of speech 210 with respect to Figure 2 ), and / or morphological information of the wordform. In an example, the primitive token "putting" can be assigned the following attributes: wordform "putting"; lemma "put"; part of speech "verb"; morphological information "aspect: progressive, tense: present, verb form: participle". The attributes can be formatted as vectors.

[0068] In some embodiments, punctuation and / or determiners in the parse tree format are discarded (310). The tree of the parse graph G can be pruned to remove any vertex including ignorable parts of speech, such as punctuation and determiners (e.g., words such as "the" or "this"). The edges connected to the pruned vertices are also removed. Additionally, in some examples, edges with ignorable dependency types (e.g., dependency "punct", indicating punctuation) can be removed.

[0069] In some embodiments, the parse tree format is vectorized into at least one syntactic n-gram form (312). In some examples, the parse tree format is vectorized into syntactic bigrams, where each edge of the graph G becomes a new token. For example, the pre-trained vector embeddings of English words V(word) can be applied to the source and target vertices of each edge: V1 and V2. The encoded vector V1(e) Dependency types of edges that can be used to connect source vertices and target vertices. Thus, the complete vector representation of the new token is a vector formed by concatenating v = [V1 (e) , V1, V2]. Similarly, instead of or in addition to syntactic bigrams, in embodiments using other integer-valued syntactic n-grams, any syntactic n-gram of order n can be represented as v = [V1 (e) , V2 (e) , … V n-1 (e) , V1, V2, … V n .

[0070] In some embodiments, for computational convenience, the number of tokens in each text is formatted to the same number of tokens (e.g., K tokens). In some examples, formatting can be achieved by removing extra vectors and / or padding empty vectors (e.g., all-zero vectors). When using a unified token numbering, any text can be represented by a matrix of dimension K by dimension v. For example, to determine the order of tokens, the number of ancestors of a syntactic n-gram is set to the number of ancestors of its source vertex (e.g., the minimum number of ancestors among the vertices of the syntactic n-gram). Further with respect to this example, syntactic n-grams can be sorted by the number of ancestors, breaking ties by the order inherited from the word order in the text, to discard redundant tokens, thereby formatting the syntactic n-grams into a unified matrix.

[0071] In some embodiments, one or more graph metrics (314) of the transformed free-form answer are calculated. For example, one or more graph metrics can be calculated by Figure 1A the answer metric engine 130. In some embodiments, the version of the parse tree format before any pruning is used to calculate at least a portion of the metrics. In some examples, the total number of primitive tokens, the number of declarative tokens (e.g., "this", "the"), the number of "stop word" tokens (e.g., "a", "the", "is", "for", etc.), and / or the number of punctuation tokens in the text of the free-form answer can be calculated before pruning. In some embodiments, the version of any pruned parse tree format is used to calculate at least a portion of the metrics. In some examples, after applying pruning to the parse tree format, the remaining number of primitive tokens, the number of noun chunks (e.g., entities in the text), the number of roots (e.g., the number of sentences), the number of leaves, and / or the average out-degree of non-leaf vertices can be calculated.

[0072] In some embodiments, the free-form answer is vectorized into at least one classical n-gram form (316). Regarding Figure 2 classical vectorization is described.

[0073] Go toFigure 4B , in some embodiments, one or more scoring criteria (326) associated with a question, a student, and / or a learning unit are identified. As described with respect to Figure 1A , for example, the criteria may be stored as one or more automated assessment rules 154.

[0074] If an assessment criterion (328) is found, then in some embodiments, one or more machine learning models (330a) applicable to the syntactic n-gram form(s) and the scoring criterion are identified. Otherwise, one or more machine learning models (330b) applicable to the syntactic n-gram form(s) are identified. The machine learning model(s) may be identified based on further information (e.g., in some examples, the availability of graphical metrics used by certain machine learning models, the learning unit, the question topic, the student level, and / or the student age).

[0075] In some embodiments, the machine learning model(s) is / are applied to one or more syntactic N-gram forms of each part of the vectorized answer and any corresponding graphical metrics to evaluate the content of a given part of the free-form answer (332). For example, the content of a given part may be evaluated to determine how closely the content of the given part matches the goal of the given part. In an example, referring to Figure 8A , the text submitted via the claim text input field 402a may be evaluated as containing language that conforms to the established claim. Similarly, the text submitted via the evidence text input field 402b may be evaluated as containing language that matches the established evidence, and the text submitted via the reasoning text input field 402c may be evaluated as containing and presenting language that matches the reasoning related to the claim and the evidence.

[0076] Returning to Figure 4B , in some embodiments, one or more machine learning models are applied to one or more syntactic N-gram forms of the vectorized answer as a whole and any corresponding graphical metrics to evaluate the logical connection between the parts of the free-form answer (334). For example, the logical connection may represent the flow of the presentation of topics and concepts between the texts of each part of the free-form answer. For example, referring to Figure 8A , the logical connection may include the use of "dense / density" and "material" between the claim text input field 402a and the evidence text input field 402b. Similarly, the logical connection may include the use of "mass", "volume", and "density" between the evidence text input field 402b and the reasoning text input field 402c.

[0077] Referring to Figure 4C, in some implementations that provide the classical n-gram form (336), one or more machine learning models applicable to the classical n-gram vectorized answers are identified according to a scoring criterion (340a) or without a scoring criterion (340b), depending on whether an evaluation rule (338) related to the classical n-gram is found. As described with respect to Figure 1A , for example, the criteria can be stored as one or more automated evaluation rules 154.

[0078] In some implementations, one or more machine learning models are applied to one or more classical n-gram forms of the vectorized answers, along with any corresponding graphical metrics associated with the machine learning models, to evaluate the style and / or grammatical quality (342) of the content. For example, classical n-grams can be used to evaluate the relationship between text submitted as a free-form answer and its literary content (e.g., form), which is contrary to the syntactic n-gram evaluation that aims at substance rather than form as described above. When applying classical n-gram analysis, method 320a can evaluate the writing ability of more complex learners.

[0079] In some implementations, the output of the machine learning models is compiled and provided for scoring analysis (344). The output of the machine learning models can be associated with the learner, the question, the type of structured answer format, the topic, the learning unit, and / or other information related to converting the machine learning analysis output into one or more scores. For example, the machine learning model output can be stored as ML analysis results 156 by the automated evaluation engine 116a of Figure 1A .

[0080] Although presented as a specific series of operations, in other embodiments, method 320a includes more or fewer operations. For example, in some implementations, instead of or in addition to identifying the machine learning models according to the evaluation criteria, a part of the machine learning models can be identified according to the availability of graphical metrics corresponding to the types of trained machine learning models. In another example, the graphical metrics can be calculated (314) before vectorizing the parse tree format into syntactic n-gram form (312). In some embodiments, certain operations of method 320a are performed simultaneously and / or in a different order. For example, the machine learning models can be executed simultaneously to evaluate the syntactic n-gram form (332, 334) and the classical n-gram form (342). Other modifications of method 320a are possible.

[0081] Returning to Figure 3 , in some implementations, automated analysis is applied to free-form answers to evaluate their content, as described by method 320b of Figure 5A and Figure 5B .

[0082] Moving on toFigure 5A In some embodiments, text (350) is obtained for each part of the answer to a multi-dimensional reasoning problem. For example, it can be obtained from Figure 3 method 300.

[0083] In some embodiments, one or more scoring criteria (352) associated with the answer context are identified. In some examples, the answer context can include an identification of the question, an identification of the topic of the question, an identification of the student and / or student statistics, and / or an identification of the learning unit. For example, the scoring criteria can be identified from Figure 1A and Figure 1B the automated scoring rules 142 and / or the automated evaluation rules 154. For example, the scoring criteria can be identified by Figure 1A the automated evaluation engine 116a of Figure 1B or the automated evaluation engine 116b of

[0084] In some embodiments, if one or more scoring criteria applicable to the answer context are available (354), then one or more AI models (356) applicable to the answer context and the scoring criteria are identified. For example, the AI models can be identified as having been trained or adjusted to evaluate answers based on a specific scoring criterion among multiple potential scoring criteria. For example, Figure 1B the model selection engine 178 of

[0085] can select certain AI models 170 based in part on the appropriate automated scoring rules 142 and / or the automated evaluation rules 154.

[0086] Conversely, if there is only one scoring criterion applicable to the system, and / or if no specific scoring criterion is identified as applicable to the answer, then in some embodiments, one or more AI models (358) applicable to the answer context are identified. Similar to the answer context part of operation 356, for example, the AI models can be identified as having been trained or adjusted to evaluate answers based on certain context factors (e.g., identified relative to operation 352).

[0086] In some embodiments, the text input format compatible with each identified AI model is identified (360). For example, the input format may be different among different AI models 170. In such a case, for example, the format suitable for each identified AI model can be identified by Figure 1B the automated evaluation engine 116b of Figure 1B or the answer formatting engine 164. In some examples, the input format can include the format described by

[0087] In some embodiments, the text of each part of the free-form answer is converted to a format compatible with each identified AI model (362). For example, the text can be in terms ofFigure 1B formatted in one or more ways described by the answer formatting engine 164.

[0088] Go to Figure 5A and Figure 5B In some embodiments, where context - specific and / or scoring - specific engineering model prompts may be used for one or more of the identified AI models (362), engineering model prompts (366) suitable for the answer context and / or scoring criteria are selected.

[0089] In some embodiments, at least one of the selected AI models (s) is applied to each part of the student answer (e.g., raw or formatted) to evaluate the part content (368). For example, applying the selected AI model(s) may include submitting the student answer to each of at least one AI model using one or more engineering model prompts. For example, the engineering model prompts may be suitable for that particular model and / or particular task (e.g., evaluation of a single part). For example, Figure 1B the automatic evaluation engine 116b may apply at least one AI model to each part of the student answer. The application may include submitting each part to a specific AI model on a part - by - part basis, with each part corresponding to a different engineering model prompt.

[0090] In some embodiments, at least one of the selected AI models (s) is applied to each part of the student answer (e.g., raw or formatted) to evaluate the logical connection between parts of the student answer (370). For example, applying the selected AI model(s) may include submitting the student answer to each of at least one AI model using one or more engineering model prompts. For example, the engineering model prompts may be suitable for that particular model and / or particular task (e.g., evaluating the logical connection between answer parts). For example, Figure 1B the automatic evaluation engine 116b may apply at least one AI model to the student answer.

[0091] In some embodiments, at least one of the selected AI model(s) is applied to each part of the student answer (e.g., original or formatted) to evaluate the style and / or grammatical quality of the student answer content (372). For example, applying the selected AI model(s) may include submitting the student answer to each of at least one AI model using one or more engineering model prompts. For example, the engineering model prompts may be suitable for that particular model and / or particular task (e.g., evaluating the style and / or grammatical elements of the student answer). Due to the evaluation of style and / or grammar, unlike previous automated evaluations using AI models, the original student answer before formatting for spelling and / or grammar correction / consistency can be used for this particular evaluation so that various typographical errors can be identified by the selected AI model(s). For example, Figure 1B the automated evaluation engine 116b can apply at least one AI model to the student answer.

[0092] In some embodiments, the output received from the AI model(s) is compiled for scoring (374). For example, scores from various evaluation techniques and / or corresponding to each part of the student answer can be assembled for generating one or more final scores corresponding to the student answer. Referring Figure 1A to the described score calculation engine 118, the compiled scores can be obtained from method 320b for generating one or more final scores.

[0093] In some embodiments, if one or more of the selected AI models provide a feedback portion (376), the evaluation reasoning of the feedback portion is compiled to present feedback to the student and / or instructor (378). For example, the feedback portion(s) can be obtained by an automated feedback engine for converting the feedback into a component of a report or user interface for review by the student and / or instructor.

[0094] Returning Figure 3 to, in some embodiments, the answer evaluation results are applied to enhance the user experience (370). In some examples, the user experience can be enhanced through real-time feedback on free-form answers, scoring of free-form answers, and / or selection of additional learning materials partially based on the evaluation results. For example, using the evaluation results to enhance the user experience will be elaborated in more detail in conjunction with Figure 7A and Figure 7B the method 370 in.

[0095] Figure 7A and Figure 7B show a flowchart of an example method 370 for evaluating machine learning analysis results of an n-gram vectorized form applied to free-form answers. Portions of method 370 can be performed by Figure 1A and / or Figure 1Bthe score calculation engine 118, Figure 1A and / or Figure 1B the student clustering engine 120, and / or Figure 1A and / or Figure 1B the learning resource recommendation engine 134 to execute.

[0096] Go to Figure 7A , in some embodiments, method 370 begins with receiving a machine learning model output and free-form answer information (372). In some examples, the free-form answer information can identify the learner, the question, the learning unit, and / or the learner's age / learning level. It can be received from Figures 4A to 4C method 320a of Figure 5A and Figure 5B method 320b of Figure 1A the automatic evaluation engine 116a of Figure 1B or the automatic evaluation engine 116b of

[0097] In some embodiments, if the evaluation is used for scoring (374), then partial evaluations are aggregated to obtain an answer score (376). For example, the machine learning output can be converted into an overall score or rating, such as a grade, a percentage from 0 to 100, or other scoring forms as described, for example, with respect to Figure 1A and / or Figure 1B the score calculation engine 118. In some embodiments, a direct aggregation is performed (e.g., all contributions of the ML models are equally weighted). In other embodiments, a scoring algorithm that applies weights to certain ML model contributions is used to combine the ML model contributions. In some embodiments, the ML model contributions are combined using the automatic scoring rules 142, as described with respect to Figure 1A and / or Figure 1B the score calculation engine 118.

[0098] In some embodiments, if the answer score meets the human scoring rules (378), then the student's free-form answer is queued for human scoring (382). The human scoring rule(s) can include high-scoring and low-scoring thresholds. In a particular example, a perfect score can be manually verified. For example, as described with respect to Figure 1A and / or Figure 1B the human scoring GUI engine 122, the student's free-form answer can be queued for double review by a teacher or other learning professional. The student's free-form answer can be scored according to the same or similar scoring rules as those followed by the automatic scoring rules 142 (referenced Figure 1A as described).

[0099] Go to Figure 7A and Figure 7B In some embodiments, if the answer score does not meet the human scoring rules (378), if the assessment reasoning is available (379), then the assessment reasoning is compiled to present feedback to the student and / or teacher (380). For example, the assessment reasoning can be derived from the ML analysis result 156 of Figure 1A and / or the AI analysis result 166 of Figure 1B . For example, the automatic feedback engine 172 of Figure 1B can assemble feedback for presentation.

[0100] In some embodiments, the answer score (and optionally, the assessment reasoning) is provided to the teacher for review and / or the student for review (381). For example, the answer score and / or the assessment reasoning can be presented by the student GUI engine 112 of Figure 1A and / or the teacher GUI engine 128 of Figure 1B .

[0101] Back to Figure 7A In some embodiments, instead of and / or in addition to the assessment for scoring (374), partial assessments are aggregated to obtain a topic proficiency assessment (384). For example, scores and / or feedback can be collected for assessing proficiency in a topic.

[0102] In some embodiments, the assessment is used to recommend the next learning activity (386). The scoring should provide an assessment of the learner's comfort level with the topic. Thus, the score(s) and / or the machine learning assessment can be provided to a recommendation process for recommending additional learning materials (388). For example, the recommendation process can be performed by the learning resource recommendation engine 134 of Figure 1A and / or the learning resource recommendation engine 134 of Figure 1B .

[0103] In some embodiments, instead of and / or in addition to using the assessment to recommend the next learning activity, the score and / or the machine learning assessment can be provided to a student clustering process (390) to group students by proficiency. For example, the student clustering process can group students based on their proficiency in one or more learning areas, e.g., to assist in presenting appropriate materials to them and / or generating comparison metrics relevant to each group. For example, the student clustering engine 120 of Figure 1A and / or the student clustering engine 120 of Figure 1B can cluster students based in part on the student clustering rules 150.

[0104] Although presented as a specific sequence of operations, in other embodiments, method 370 includes more or fewer operations. For example, in some embodiments, if the assessment process is used for student clustering (390) and / or for recommending additional learning activities (388), the scoring criteria may be different. For example, while the assessment provided to students may be presented in letter grade format, percentage points or other mathematical level assessments may be used for student clustering and / or recommendation purposes. In some embodiments, certain operations of method 370 are performed simultaneously and / or in a different order. For example, the scores may be presented to the teacher and / or student for review (380), while the free-form answers may also be queued for manual scoring (382). Other modifications to method 370 are possible.

[0105] Return to Figure 3 , although presented as a specific sequence of operations, in other embodiments, method 300 includes more or fewer operations. For example, in addition to applying the spelling / grammar algorithm, method 300 may also include scoring the grammar / spelling portion of the answer. For example, depending on the age and / or ability of the student, the scoring may be optional. Other modifications to method 300 are possible.

[0106] Figure 9 is a flowchart of an example process 500 for training a machine learning model to automatically score vectorized free-form answers to multi-dimensional reasoning questions. For example, process 500 may be executed by the automatic assessment system 102a of FIG. 1.

[0107] In some implementations, in the first round of training, a set of sample answers 140a is provided to the answer vectorization engine 114 for generating one or more vectorized forms 504 of each sample answer 140a. Additionally, for each sample answer 140a, the answer metric engine 130 may cooperate with the answer vectorization engine 114 to generate an answer metric 506 associated with the one or more vectorized forms of the sample answer 140a generated by the answer vectorization engine 114. Additionally, the answer metric engine 130 may generate one or more metrics associated with each sample answer 140a prior to vectorization (e.g., token counting, etc.).

[0108] In some embodiments, the machine learning model training engine 126 accesses the vectorized form 504 of the sample answers 140 and the corresponding answer metrics 506 to train one or more models. For example, the machine learning model training engine 126 can feed the vectorized answers 504, the corresponding answer metrics 506, and a set of sample answer scores 140b corresponding to the sample answer 140a to one or more tree-based machine learning classifiers. In some embodiments, the type(s) of tree-based machine learning classifier(s) used can be selected by the ML model training engine 126, at least in part, based on the evaluation rule set 142. The evaluation rule set 142 can also specify combinations of the vectorized answers 504, such as a first combination consisting of the vectorized form of the claim portion of the sample answer and the vectorized form of the evidence portion of the sample answer, and a second combination consisting of the vectorized form of the evidence portion of the sample answer and the vectorized form of the reasoning portion of the sample answer. The ML model training engine 126 generates a set of trained models 508 from the answer metrics 506, the vectorized answers 504, and the sample answer scores 140b for storage as the trained machine learning model 108.

[0109] When the trained machine learning model 108 is applied to automatically evaluate free-form answers formatted in a multi-part answer architecture, in some embodiments, a set of manually re-scored answers 502a is collected. For example, the manually re-scored answers 502a can be generated from automatically identified free-form answers that are evaluated by the trained ML model 108 as having a score that matches the automated scoring rules 142, as described Figure 1A above. For example, the manually re-scored answers 502a can correspond to free-form answers that are scored by the trained ML model 108 as perfect (e.g., assigned a score of 100%) and / or very poor (e.g., assigned a score of at most 10%, at most 5%, or 0). When manually re-scoring answers corresponding to the ML model evaluation outputs on the edges of the scoring spectrum, a small number of training updates can be used to refine the trained ML model 108 in a targeted manner.

[0110] In some embodiments, the manually re-scored answers 502a are provided to the answer vectorization engine 114 and the answer metrics engine 130 to generate the vectorized answers 504 and the answer metrics 506. In some examples, the manually re-scored answers 502a can be used to re-train the trained ML model 108 whenever the manually re-scored answers 502a are available, whenever a threshold number (e.g., 5, 10, 20, etc.) of manually re-scored answers are available, and / or periodically.

[0111] In some embodiments, the vectorized answer 504 and answer metric 506 generated from the manually re-scored answer 502a, along with any trained ML model 108 corresponding to the manually re-scored answer 502a (e.g., same question, same subject area, same answer part format, and / or same learning unit, etc.) and the manual score 502b corresponding to the manually re-scored answer 502a, are provided to the ML model training engine 126 to update the corresponding trained ML model 108 to a trained model 508.

[0112] Go to Figure 10 , Figure 10 FIG. 600 is a flowchart of an example process for tuning an artificial intelligence model to automatically score free-form answers to multi-dimensional reasoning questions. For example, process 600 can be performed by Figure 1B the automated evaluation system 102b.

[0113] In some embodiments, in a first round of tuning, a set of sample answers 140a is provided to the answer formatting engine 164 to generate one or more formatted versions 604 of each sample answer 140a.

[0114] In some embodiments, the formatted answers 604 are accessed by Figure 1B the model tuning engine 174 to tune one or more base models 170. For example, the model tuning engine 174 can feed the formatted answers 604 and a set of sample answer scores 140b corresponding to the sample answers 140a into one or more base models 170a to tune the base model(s) 170a for performing automated analysis based on the sample answers 140a. The AI model tuning engine 174 causes a functional adjustment of one or more base models 170, converting them into adjusted AI model(s) 170b.

[0115] In some embodiments, when querying the adjusted AI model(s) 170b to automatically evaluate free-form answers formatted in a multi-part answer architecture, a set of manually re-scored answers 602a is collected. For example, the manually re-scored answers 602a can be generated from automatically identified free-form answers that are evaluated by the adjusted AI model 170b as having a score that matches the automated scoring rule 142, as described with respect to Figure 1AAs described. The manually re-scored answer 602a can correspond to, for example, a free-form answer that is scored as perfect (e.g., assigned a score of 100%) and / or a very poor score (e.g., assigned at most 10% of the score, assigned at most 5% of the score, or assigned a score of 0) by the adjusted AI model 170b. When manually re-scoring an answer corresponding to the AI model evaluation output on the edge of the scoring spectrum, a small number of training updates can be used to improve the adjusted AI model 170b in a targeted manner.

[0116] In some embodiments, the manually re-scored answer 602a is provided to the answer formatting engine 164 to generate a further formatted answer 604. In some examples, whenever the manually re-scored answer 602a is available, whenever a threshold number (e.g., 5, 10, 20, etc.) of manually re-scored answers are available, and / or periodically, the manually re-scored answer 602a can be used to improve the tuning of the tuned AI model 170b.

[0117] In some embodiments, the formatted answer 604, along with any adjusted AI model 170b corresponding to the manually re-scored answer 602a (e.g., the same question, the same subject area, the same answer part format, and / or the same learning unit, etc.) and the manual score 602b corresponding to the manually re-scored answer 602a are provided to the AI model tuning engine 174 to optimize the tuning of the corresponding adjusted AI model 170b.

[0118] Reference has been made to the drawings that represent methods and systems according to embodiments of the present disclosure. Aspects thereof can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device and / or a distributed processing system having a processing circuit, such that the instructions executed via the processor of the computer or other programmable data processing device create means for performing the functions / operations specified in the drawings.

[0119] One or more processors can be utilized to implement the various functions and / or algorithms described herein. Additionally, any of the functions and / or algorithms described herein can be executed on one or more virtual processors. For example, a virtual processor can be part of one or more physical computing systems (such as a computer cluster or a cloud drive).

[0120] Aspects of the present disclosure may be implemented by software logic, including machine-readable instructions or commands for execution via a processing circuit. In some examples, software logic may also be referred to as machine-readable code, software code, or programming instructions. In certain embodiments, software logic may be encoded as runtime executable commands and / or compiled into a machine-executable program or file. Software logic may be programmed and / or compiled into various coding languages or formats.

[0121] Aspects of the present disclosure may be implemented by hardware logic (wherein the hardware logic naturally also includes any necessary signal wiring, memory elements, etc.), which, except for initial system configuration and any subsequent system reconfiguration (e.g., for different object mode dimensions), is capable of operating without the involvement of active software. The hardware logic may be synthesized on a reprogrammable computing chip such as a field-programmable gate array (FPGA) or other reconfigurable logic device. Additionally, the hardware logic may be hard-coded onto a custom microchip, such as an application-specific integrated circuit (ASIC). In other embodiments, software stored as instructions on a non-transitory computer-readable medium, such as a memory device, on-chip integrated memory cell, or other non-transitory computer-readable memory, may be used to perform at least a portion of the functions described herein.

[0122] Aspects of the embodiments disclosed herein are executed on one or more computing devices, such as a laptop computer, a tablet computer, a mobile phone, or other handheld computing device, or one or more servers. Such computing devices include processing circuitry instantiated in one or more processors or logic chips, such as a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). Additionally, the processing circuitry may be implemented as multiple processors that work together (e.g., in parallel) to execute the instructions of the inventive methods described above.

[0123] The process data and instructions for performing the various methods and algorithms derived herein may be stored in a non-transitory (i.e., non-volatile) computer-readable medium or memory. The claimed technological advancements are not limited to the form of the computer-readable medium storing the instructions of the inventive method. For example, the instructions may be stored on a CD, DVD, flash memory, RAM, ROM, PROM, EPROM, EEPROM, a hard disk, or any other information processing device with which the computing device communicates, such as a server or a computer. In some examples, the processing circuitry and the stored instructions may enable the computing device to execute Figure 3 method 300, Figures 4A to 4C method 320a, Figure 5A and Figure 5B method 320b, Figure 6 method 360, Figure 7Aand Figure 7B method 370 and / or Figure 9 process 500.

[0124] These computer program instructions can direct a computing device or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in a computer-readable medium result in an article of manufacture including an instruction device that performs the specified functions / operations in the illustrated process flow.

[0125] Embodiments of this specification rely on network communication. It can be understood that the network can be a public network, such as the Internet, or a private network, such as a local area network (LAN) or wide area network (WAN), or any combination thereof, and can also include a PSTN or ISDN subnet. The network can also be wired, such as an Ethernet network, and / or can be wireless, such as a cellular network including EDGE, 3G, 4G, and 5G wireless cellular systems. The wireless network can also include or other wireless forms of communication. For example, the network can support Figure 1A and Figure 1B communications between the automated assessment systems 102a, 102b and the student device 104 and / or the teacher device 106.

[0126] In some embodiments, the computing device further includes a display controller for interacting with a display, such as a built-in display or an LCD monitor. The general I / O interface of the computing device can be connected and interact with a keyboard, a manually manipulated motion tracking I / O device (e.g., a mouse, a virtual reality glove, a trackball, a joystick, etc.), and / or a touch screen panel or a touchpad on or separate from the display. In some examples, the display controller and the display can be presented as Figure 8A and Figure 8B the screenshots 400 and 420 shown in

[0127] Furthermore, the present disclosure is not limited to the specific circuit elements described herein, nor is it limited to the specific dimensions and classifications of these elements. For example, those skilled in the art will understand that the circuits described herein can be adapted based on changes in battery size and chemistry or based on the requirements of the expected standby load to be powered.

[0128] The functions and features described herein can also be performed by various distributed components of the system. For example, one or more processors can perform these system functions, where the processors are distributed across multiple components communicating over a network. In addition to various human-machine interfaces and communication devices (e.g., display monitors, smartphones, tablets, personal digital assistants (PDAs)), the distributed components can include one or more client and server machines that can share processing. The network can be a private network (e.g., LAN or WAN), or can be a public network (e.g., the Internet). In some examples, inputs to the system can be received via direct user input and / or remotely in real-time or as a batch process.

[0129] Although provided in the context of this document, in other embodiments, the methods and logic flows described herein can be performed on modules or hardware different from those described. Accordingly, other embodiments are also within the scope of what can be claimed.

[0130] In some embodiments, cloud computing environments such as Google Cloud Platform TM or Amazon TM Web Services (AWS TM ) can be used to perform at least a portion of the above-described methods or algorithms. The processes associated with the methods described herein can be performed on computing processors in a data center. For example, the data center can also include an application processor that can be used as an interface to the systems described herein to receive data and output corresponding information. The cloud computing environment can also include one or more databases or other data stores, such as cloud storage and query databases. In some embodiments, cloud storage databases such as Google TM Cloud Storage or Amazon TM Elastic File System (EFS TM ) can store processed and unprocessed data provided by the systems described herein. For example, Figure 1A the contents of data store 110a, Figure 1B the contents of data store 110b, Figure 9 the sample answers 140a and sample answer scores 140b, and / or Figure 9 the manually re-scored answers 502a and corresponding manually assigned scores 502b can be maintained in a database structure.

[0131] The systems described herein can communicate with the cloud computing environment via a security gateway. In some embodiments, the security gateway includes a database query interface, such as Google BigQuery TMPlatform or Amazon RDS TM For example, the data query interface can support the automatic evaluation system 102a to access Figure 1A at least part of the data stored in the data storage 110a.

[0132] The systems described herein may include one or more artificial intelligence (AI) networks (e.g., neural networks) for natural language processing (NLP) of text input. In some examples, the AI network may include a synaptic neural network, a deep neural network, a transformer neural network, and / or a generative adversarial network (GAN). One or more machine learning techniques and / or classifiers may be used to train the AI network, e.g., in some examples, anomaly detection, clustering, and / or supervised and / or association. In one example, the AI network may be developed by Google in Mountain View, California and / or based on the Bidirectional Encoder Representations from Transformers (BERT) model.

[0133] The systems described herein may communicate with one or more base model systems (e.g., artificial intelligence neural networks). In some examples, the base model system(s) may be developed, trained, adjusted, fine-tuned, and / or prompt-engineered to evaluate text input, such as Figure 1B the student answer 168. In some examples, the base model system may include or be based on a Generative Pretrained Transformer (GPT) model (e.g., GPT-3, GPT-3.5, and / or GPT-4) obtained through the OpenAI platform of OpenAI in San Francisco, California and / or a generative AI model (e.g., PaLM 2) obtained through Azure OpenAI or Vertex AI of Google in Mountain View, California. Multiple base model systems may be applied according to the audience (e.g., students, teachers, student levels, learning topics, evaluation criteria, etc.). In another example, a single base model system may be dynamically adjusted based on user statistics and / or topic context. In the illustration, different engineering prompts based on statistics and / or topic context may be used to query a single large language model (LLM) for natural language processing (NLP) of student text.

[0134] Although certain embodiments have been described, these embodiments are presented by way of example only and are not intended to limit the scope of the disclosure. In fact, the novel methods, devices, and systems described herein may be instantiated in a variety of other forms; furthermore, various omissions, substitutions, and changes may be made to the forms of the methods, devices, and systems described herein without departing from the spirit of the disclosure. The appended claims and their equivalents are intended to cover forms or modifications that fall within the scope and spirit of the disclosure.

Claims

1. A system for automatically evaluating the content of free-form answers, the system comprising: A non-volatile data store including a plurality of multi-dimensional reasoning questions related to one or more topics; And Processing circuitry configured to perform operations including: Obtaining a free-form answer to a given multi-dimensional reasoning question of the plurality of multi-dimensional reasoning questions, the free-form answer including partial text corresponding to each respective part of a predetermined partial structure of the free-form answer; For each respective part of the predetermined partial structure, obtaining at least one partial content score by applying one or more artificial intelligence (AI) models to the corresponding partial text of the respective part to perform respective part content evaluation; Obtaining at least one logical connection score by applying another AI model to the corresponding partial text of adjacent respective parts in the predetermined partial structure to perform logical connection evaluation; and Using the at least one partial content score and the at least one logical connection score to calculate at least one total score corresponding to the free-form answer.

2. The system according to claim 1, wherein, The one or more AI models include the another AI model.

3. The system according to claim 1, further comprising, for each respective part of the predetermined partial structure, formatting the corresponding partial text of the respective part into formatted partial text.

4. The system according to claim 1, wherein, The predetermined partial structure includes three parts.

5. The system according to claim 1, wherein: Performing the respective part content evaluation includes obtaining respective evaluation inferences of one or more partial content scores corresponding to the at least one partial content score from each AI model of at least one AI model of the one or more AI models; And The operations further include: Using the respective evaluation inferences of each AI model of the at least one AI model to prepare a feedback message; And Providing the feedback message to a remote computing device of an end user.

6. The system according to claim 5, wherein, The feedback message includes an explanation or actionable advice.

7. The system according to claim 5, wherein: The end user is a student; and The feedback message is provided in real time or near real time.

8. The system according to claim 1, wherein, The operations further include: selecting the one or more AI models at least partially based on evaluation criteria corresponding to the free-form answer.

9. The system according to claim 1, wherein The operations further include: Applying one or more human scoring rules to the at least one total score; and Based on the application indicating that the free-form answer is eligible for human scoring, queuing the free-form answer for human review.

10. The system according to claim 9, wherein, The operations further include: Obtaining at least one human-calculated score from the human review; and Providing the corresponding partial text of each respective part of the predetermined partial structure and the at least one human-calculated score to adjust i) the one or more AI models and / or ii) at least one AI model of the another AI model.

11. The system according to claim 9, wherein The one or more human scoring rules include identifying any perfect scores.

12. A method for automatically evaluating the content of free-form answers, the method comprising: Obtain the corresponding partial text of each of the multiple text parts of the free-form answer through the text input user interface of the e-learning platform, where the free-form answer is for a given multi-dimensional reasoning question among multiple multi-dimensional reasoning questions of the e-learning platform; Format the corresponding partial text of each respective text part of the multiple text parts through a processing circuit for submission to at least one corresponding artificial intelligence (AI) model among one or more AI models; For each respective text part of the multiple text parts, submit the formatted corresponding partial text of the respective text part by the processing circuit to at least one first AI model among the one or more AI models to form a content evaluation of the respective text part; Submit the formatted corresponding partial text of each respective text part of the multiple text parts by the processing circuit to at least one second AI model among the one or more AI models to form a logical connection evaluation of the formatted corresponding partial text of adjacent text parts among the multiple text parts; And Calculate at least one score corresponding to the free-form answer, at least in part based on the respective content evaluations and the logical connection evaluations of each of the respective text parts of the multiple text parts.

13. The method according to claim 12, further comprising presenting the text input user interface at a display of a remote computing device, wherein, The text input user interface includes a separate text input pane for each text part of the multiple text parts.

14. The method according to claim 12, wherein, Formatting the corresponding partial text of each respective part includes discarding punctuation and / or discarding determiners from multiple tags of the corresponding partial text of each respective text part.

15. The method according to claim 12, further comprising: Submit the corresponding partial text of each respective part of the multiple text parts by the processing circuit to at least one third AI model among the one or more AI models to form a style and / or grammar quality evaluation of the free-form answer; Wherein calculating the at least one score includes further calculating the at least one score using the style and / or grammar quality evaluation.

16. The method according to claim 12, further comprising: Adjust at least one model among the one or more AI models by the processing circuit using multiple sample answers and corresponding scores, thereby generating at least one adjusted AI model among the one or more AI models; Identify, by the processing circuit, multiple free-form answers, each free-form answer receiving one or more perfect scores for the at least one score through the calculation; And Update the adjustment of the at least one adjusted AI model by the processing circuit using the multiple free-form answers.

17. The method according to claim 16, wherein At least part of the multiple sample answers includes a set of answers submitted by multiple learners of the e-learning platform and manually scored by a group of professionals.

18. The method according to claim 12, further comprising providing the at least one score to a recommendation process by the processing circuit to select a next learning activity for a learner.

19. The method according to claim 12, wherein, Submitting the formatted corresponding partial text to the first at least one AI model includes: submitting the formatted corresponding partial text and engineering prompts designed to request the at least one first AI model to evaluate the formatted corresponding partial text.