A multi-agent collaborative evaluation system and method for student self-generated courses

CN122550322APending Publication Date: 2026-08-11JINAN VOCATIONAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-11

AI Technical Summary

Benefits of technology

[0009]与现有技术相比,本发明具有如下显著优点:1、通过多智能体协同机制,实现了学生从课程接受者向课程共建者的转变,显著提升了学生在自主学习场景下的主体性;2、通过素养映射智能体的语义匹配与补全功能,解决了个性化评价量规与国家核心素养标准之间的对齐问题,确保了评价的规范性与科学性;3、通过三方共识校准与规则自适配机制,能够自动化化解评分争议并持续优化评价精度,结合链式哈希审计技术,为教育治理提供了全流程可追溯、不可篡改的可信证据链。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550322A_ABST
    Figure CN122550322A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-agent collaborative evaluation systems and methods for student self-generating course, the system includes the perception interaction layer, for receiving the natural language prompt input by student end;Intelligent agent cooperation layer includes main control arrangement intelligent agent, course generation intelligent agent, literacy mapping intelligent agent, evaluation scoring intelligent agent, calibration dialogue intelligent agent and audit tracking intelligent agent;Algorithm engine layer includes core literacy bidirectional mapping algorithm, course structured generation algorithm, multi-source scoring difference detection algorithm, scoring rule self-adapting algorithm and chain audit track generation algorithm;It further includes data resource layer, business service layer and application layer;The application realizes student self-course generation, evaluation standard alignment, scoring dispute automation resolution and full-link credible audit, and further improves the individualization and standardization level of educational evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of evaluation system technology, specifically to a multi-agent collaborative evaluation system and method for student-generated courses. Background Technology

[0002] As the digital transformation of education deepens, traditional educational evaluation and management are gradually shifting towards learner-participatory construction, process-oriented assessment, collaborative decision-making, and trustworthy governance. Especially against the backdrop of the rapid development of generative artificial intelligence and multi-agent systems, modern education systems already possess the ability to support students in autonomously generating course content, such as learning resource recommendations, intelligent feedback, and automatic scoring. However, existing technologies still have significant shortcomings and limitations in supporting students' autonomous course generation.

[0003] First, existing course generation systems are often teacher- or platform-driven, primarily generating instructional designs, lesson plans, or learning resources based on pre-set teaching objectives, subject themes, and standardized templates. This approach limits student autonomy and creativity during the learning process, with most students passively learning within a predetermined course structure, rather than generating personalized syllabi and phased goals based on individual interests, time constraints, learning objectives, and preferences. Even when some course design systems begin to support personalized learning path adjustments, these adjustments mainly focus on resource recommendation and task distribution, failing to form a complete technological chain and lacking a closed-loop process from student input prompts to automatic course generation, and from objective decomposition to automatic generation of evaluation rubrics.

[0004] Secondly, existing student self-assessment rubrics often lack an effective automatic alignment mechanism with formal educational evaluation standards. Student-generated courses may produce evaluation dimensions with innovative and personalized characteristics, such as project practice, creative application, and interdisciplinary collaboration. However, during course evaluation, these personalized dimensions must be consistent with formal evaluation standards under the national curriculum standards and core competency framework to ensure the fairness and consistency of course evaluation. However, current technology cannot automatically map students' personalized evaluation dimensions to formal educational standards. This leads to problems such as fragmented evaluation standards and incomparability of evaluation results between courses, affecting the standardization and fairness of the overall educational evaluation system.

[0005] Third, traditional educational information systems focus more on results such as scores, transcripts, and simple comments, lacking structured records of the causal relationships between course generation prompts, version generation, rubric changes, scoring discrepancies, calibration dialogues, rule revisions, and final reports. This makes it difficult to support subsequent supervision and verification, teaching debriefing, dispute arbitration, and quality audits. Especially when generative artificial intelligence is involved in curriculum design and evaluation, without a fully traceable, verifiable, and tamper-proof technical mechanism, it will be difficult to meet the needs of digital governance, process evaluation, and accountability in education. Summary of the Invention

[0006] To address the aforementioned shortcomings and limitations in existing technologies, this invention provides a multi-agent collaborative evaluation system and method for student-generated courses. This system supports student-driven course generation, enables bidirectional mapping between personalized rubrics and core competency standards, and possesses multi-agent collaborative evaluation and rule self-adaptation capabilities.

[0007] According to one aspect of the present invention, a multi-agent collaborative evaluation system for student-generated courses is provided, comprising: The perception and interaction layer is used to receive natural language prompts from the student's input. The intelligent agent collaboration layer includes a master orchestration agent, a curriculum generation agent, a competency mapping agent, an assessment and scoring agent, a calibration dialogue agent, and an audit trail agent; The algorithm engine layer includes a core competency bidirectional mapping algorithm, a curriculum structure generation algorithm, a multi-source scoring difference detection algorithm, a scoring rule self-adaptation algorithm, and a chain-based audit trajectory generation algorithm; The process involves the following steps: First, the master orchestration agent schedules the course generation agent to generate a course outline, chapter structure, stage tasks, learning arrangements, and initial self-assessment rubrics based on natural language prompts. Second, the competency mapping agent maps the initial self-assessment rubrics to at least one dimension within the core competency standard framework using the core competency bidirectional mapping algorithm, generating a standardized mapping rubric. Third, the assessment and scoring agent integrates student self-assessments, teacher assessments, and agent machine assessments to obtain multi-source scores, and determines the difference detection results using the multi-source score difference detection algorithm. Fourth, if the difference detection results exceed a preset threshold, the calibration dialogue agent generates a disagreement report, a list of evidence, and calibration suggestions, initiates a three-party calibration dialogue, and outputs a rule correction signal. Fifth, the scoring rule self-adaptation algorithm outputs the corrected scoring results and rule correction signals based on the consensus reached during the dialogue. Sixth, the audit trail agent performs chained hash encapsulation and timestamp signing to generate an auditable assessment report.

[0008] According to another aspect of the present invention, a multi-agent collaborative evaluation method for student-generated courses is provided, the method comprising: S1: Receive natural language prompts from the student's input, perform intent recognition and task decomposition, and assign the corresponding tasks to the corresponding intelligent agents. Among them, the course generation intelligent agent generates the course outline, chapter structure, stage tasks, learning arrangements and initial self-assessment rubric. S2: The initial self-assessment rubric is semantically vectorized and matched with the standard items in the preset core competency standard framework library to establish a positive mapping relationship. The key items of the core competencies that are not covered are completed in reverse to form a standardized mapping rubric. S3: Integrate student self-scores obtained from standardized mapping metric, teacher external scores obtained from formal evaluation rules, and AI agent machine scores to obtain multi-source scores, and determine the difference detection results through a multi-source score difference detection algorithm; S4: When the difference detection result is lower than the preset threshold, generate the result; when the difference detection result is higher than the preset threshold, generate the divergence report and initiate a three-party calibration dialogue, and output the corrected scoring result and rule correction signal based on the consensus formed in the dialogue. S5: Dynamically and adaptively update the scoring rules according to the rule correction signal, encapsulate the entire process event using chain hashing, and generate an auditable evaluation report.

[0009] Compared with existing technologies, this invention has the following significant advantages: 1. Through a multi-agent collaborative mechanism, it realizes the transformation of students from course recipients to course co-builders, significantly enhancing students' subjectivity in self-directed learning scenarios; 2. Through the semantic matching and completion function of the competency mapping agent, it solves the alignment problem between personalized evaluation rubrics and national core competency standards, ensuring the standardization and scientific nature of the evaluation; 3. Through a tripartite consensus calibration and rule self-adaptation mechanism, it can automatically resolve scoring disputes and continuously optimize evaluation accuracy. Combined with chain hash auditing technology, it provides a fully traceable and tamper-proof credible evidence chain for education governance. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of a multi-agent collaborative evaluation system architecture for student-generated courses provided by an embodiment of the present invention;

[0011] Figure 2 This is a schematic diagram of the interaction process of the intelligent agent collaboration layer provided in an embodiment of the present invention;

[0012] Figure 3 This is a schematic diagram of the core competency bidirectional mapping algorithm provided in an embodiment of the present invention;

[0013] Figure 4 This is a schematic diagram of the difference detection and third-party calibration process provided in an embodiment of the present invention;

[0014] Figure 5This is a schematic diagram of the scoring rule self-adaptation engine structure provided in an embodiment of the present invention;

[0015] Figure 6 This is a schematic diagram of an auditable and evaluable trajectory chain-type evidence storage structure provided in an embodiment of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0017] Existing technologies typically generate instructional designs, lesson plans, or learning resources by having teachers input teaching objectives, subject topics, and standardized templates. Students mostly engage in passive learning within a predetermined curriculum structure, lacking the ability to autonomously drive the system to generate course outlines, stage goals, and assessment rubrics based on personal interests, learning objectives, time constraints, preferred learning formats, and outcome orientation. Even when some systems support personalized learning path adjustments, they mostly focus on resource recommendation and task distribution.

[0018] When students create their own rubrics, they often generate personalized assessment dimensions that match the course content, such as project practice, creative application, reflective ability, and collaborative expression. However, schools, teachers, and education supervisory departments typically need to conduct assessments based on national curriculum standards, core competency frameworks, and academic quality standards. If student-defined rubrics cannot automatically establish a mapping relationship with formal education standards, it will lead to fragmented assessment standards, incomparability between different courses, and difficulty in integrating personalized assessments into the formal education governance system, thereby affecting the fairness, consistency, and compliance of the assessment.

[0019] The evaluation rules of existing technologies are pre-set during the deployment phase, mainly including fixed indicators, fixed weights, fixed thresholds, and fixed feedback logic. It is difficult to dynamically adjust the evaluation rules based on course type, task complexity, students' existing abilities, historical calibration results, and the quality of evidence. For highly contextualized courses such as project-based learning, inquiry-based learning, interdisciplinary learning, and vocational skills training, static evaluation rules often fail to accurately reflect learners' true levels, easily leading to evaluation bias and accumulated disputes.

[0020] Most existing technologies only support evaluation comparison or evaluation summary. Especially when the evaluation difference exceeds the threshold, there is a lack of automatically triggered calibration protocols and interpretable correction basis, making it difficult for the system to learn from a dispute and optimize subsequent scoring strategies.

[0021] Current technologies focus more on outcome scores, transcripts, and simple comments, lacking structured records of the causal relationships between course generation prompts, generated versions, rubric changes, scoring discrepancies, calibration dialogues, rule revisions, and final reports. This makes it difficult to support subsequent supervision and verification, teaching debriefing, dispute arbitration, and quality audits. Especially when generative artificial intelligence is involved in curriculum design and evaluation, without a fully traceable, verifiable, and tamper-proof technical mechanism, it will be difficult to meet the needs of digital governance in education, process evaluation, and accountability.

[0022] Example 1 To address the aforementioned problems in the existing technology, please refer to... Figure 1 This invention provides a multi-agent collaborative evaluation system for student-generated courses. This system can be configured to be deployed on at least one computing device, which may include, but is not limited to, servers, cloud computing clusters, edge computing nodes, or distributed computing platforms. Specifically, it includes: The perception and interaction layer is used to receive natural language prompts from the student's input.

[0023] Specifically, the perception and interaction layer receives natural language prompts from students, standard configurations from teachers, audit query requests from supervisors, policy control instructions from school management, and results visualization requests. Natural language prompts may include at least one of the following: learning topic, expected goals, duration constraints, learning preferences, outcome format, and evaluation preferences. For example, natural language prompts may include constraints such as a learning cycle of approximately 4 weeks and approximately 3 class hours per week, but this disclosure is not limited to this. The perception and interaction layer can preprocess the received natural language prompts, including noise filtering, word segmentation, part-of-speech tagging, and preliminary intent recognition. In this way, the perception and interaction layer can significantly reduce the technical threshold for students to participate in curriculum design and improve the interactive experience.

[0024] The data resource layer is used to store and manage student profile databases, assessment standard databases, course knowledge bases (including course templates and course examples), educational resource databases, core competency standard framework databases, historical scoring databases, calibration session log databases, audit trail databases, and evidence material databases.

[0025] The agent collaboration layer is used to parse natural language prompts and schedule multiple agents to generate course outlines, chapter structures, stage tasks, learning arrangements, and initial self-assessment rubrics. The agents include master orchestration agent, course generation agent, goal decomposition agent, rubric generation agent, competency mapping agent, assessment and scoring agent, calibration dialogue agent, and audit trail agent.

[0026] For details, please refer to Figure 2The agent collaboration layer is responsible for decomposing the complex course generation task into multiple executable subtasks and assigning them to specialized agents. When a student submits a course generation prompt, the master orchestration agent performs semantic parsing, task decomposition, and process orchestration on the natural language prompt. In some embodiments, the master orchestration agent can utilize at least one pre-trained language model to extract key entities and constraints from the prompt. Based on the extracted information, the master orchestration agent can dynamically generate a directed acyclic graph of task execution, determining the dependencies and execution order of each subtask. For example, the master orchestration agent can first schedule the course generation agent to build the basic framework, and then schedule the target decomposition agent and metric generation agent in parallel for detail filling. This orchestration mechanism ensures efficient collaboration among multiple agents and avoids task conflicts.

[0027] The course-generating agent generates a course syllabus, chapter structure, and stage tasks according to the scheduling of the master orchestration agent. In some embodiments, the course-generating agent can retrieve at least one course template library and, in conjunction with students' personalized prompts, generate course content that conforms to pedagogical logic. For example, for a course on artificial intelligence, the course-generating agent can generate a chapter structure that includes an introduction, basic algorithms, network structures, and application projects, and allocate appropriate class hours and stage tasks to each chapter.

[0028] The goal decomposition agent breaks down curriculum objectives into cognitive, skill, and competency objectives. In some embodiments, the agent can further subdivide macro-level curriculum expectations into finer-grained categories based on at least one educational goal taxonomy theory. For example, cognitive objectives may include understanding specific concepts, skill objectives may include performing specific operations, and competency objectives may include possessing specific reflective or collaborative abilities. Through this multi-dimensional decomposition, the agent provides a structured basis for subsequent accurate assessment.

[0029] The metric generation agent generates an initial self-assessment metric that includes at least evaluation metrics, metric descriptions, evidence requirements, rating levels, and initial weights. In some embodiments, the metric generation agent can automatically generate corresponding evaluation dimensions based on the decomposed objectives. For example, the initial weight can be approximately 0.14. The metric generation agent ensures that each evaluation metric has clear evidence requirements, such as requiring the submission of a project poster or explanatory video as a basis for scoring.

[0030] See Figure 3The competency mapping agent performs bidirectional mapping and missing item completion on the initial self-assessment rubric and the core competency standard framework to generate a standardized mapping rubric. In existing technologies, student-defined rubrics are often difficult to align with formal national or school education standards. The competency mapping agent overcomes this limitation and achieves a balance between personalization and standardization by introducing a bidirectional mapping algorithm for core competencies. Subsequently, the competency mapping agent performs similarity matching between the vectorized evaluation indicators and the standard items in the core competency standard framework library to establish a forward mapping. In some embodiments, similarity matching can use at least one distance metric method such as cosine similarity or Euclidean distance. For example, a forward mapping can be established when the similarity score is greater than or equal to a first threshold; the first threshold can be approximately 0.78. Through forward mapping, the system can identify the core competency dimensions already covered in the student rubric. The competency mapping agent performs reverse completion on the standard items not covered in the core competency standard framework library, supplementing the missing indicators. In some embodiments, if the maximum matching score of a key core competency item in the forward mapping is lower than a second threshold, a reverse completion mechanism is triggered; the second threshold can be approximately 0.62. The competency mapping agent can automatically generate evaluation indicators corresponding to the missing competency item and recommend corresponding initial weights. For example, if the initial rubric lacks an assessment of the information awareness dimension, the competency mapping agent can automatically supplement the information retrieval and discrimination ability indicators. Finally, the competency mapping agent performs indicator conflict resolution and weight normalization, outputting a standardized mapping rubric. In some embodiments, conflict resolution may include merging semantically highly overlapping indicators or eliminating redundant indicators irrelevant to the curriculum objectives. Weight normalization ensures that the sum of the weights of all indicators is a specific value, such as 1.0. Through the above steps, the competency mapping agent can output a standardized mapping rubric that retains students' individual characteristics while conforming to formal education standards.

[0031] The assessment and scoring module intelligent agent acquires student self-assessment results, teacher peer assessment results, and intelligent agent machine assessment results based on standardized mapping rubrics, and performs multi-source score difference detection. Existing technologies typically rely on only a single evaluator, which is prone to subjective bias. The assessment and scoring intelligent agent overcomes this limitation and improves the objectivity and credibility of the assessment by fusing multi-source evaluation data and performing in-depth difference analysis. The assessment and scoring intelligent agent calculates the overall deviation rate, dimensional deviation rate, structural difference degree, and evidence consistency index of the multi-source score vector. In some embodiments, the overall deviation rate can reflect the degree of difference in total scores among different evaluators; the dimensional deviation rate can locate the specific evaluation indicators where significant disagreements exist; the structural difference degree can measure the differences in the distribution pattern of scores among different subjects; and the evidence consistency index is used to assess the logical fit between the scoring results and the submitted evidence. By calculating these multi-dimensional difference indicators, the assessment and scoring intelligent agent can comprehensively characterize the state of disagreement among multi-source scores. The assessment and scoring intelligent agent dynamically adjusts the preset threshold based on course complexity, evidence completeness, and historical calibration variance. In some embodiments, the preset threshold is not fixed but is adaptively calculated according to the current assessment context. For example, for courses with high complexity, insufficient evidence, or significant historical controversy, the system can appropriately lower the threshold for triggering calibration to more sensitively capture potential scoring biases. The preset threshold can be approximately 16.50. This dynamic adjustment mechanism effectively balances the system's calibration sensitivity with operational efficiency.

[0032] See Figure 4When the result of multi-source score difference detection exceeds a preset threshold, the calibration dialogue agent triggers a three-way consensus calibration involving the student, teacher, and agent, and generates a corrected result. Existing technologies lack automated mechanisms for resolving scoring disagreements. The calibration dialogue agent according to embodiments of this disclosure overcomes this limitation and achieves interpretable resolution of disputes by introducing a multi-round dialogue and evidence questioning mechanism. Specifically, when the result of multi-source score difference detection exceeds a preset threshold, the calibration dialogue agent automatically generates a disagreement report, an evidence list, and calibration recommendations. In some embodiments, the disagreement report can visually display the evaluation dimension with the greatest difference, the evidence list can list all supporting materials related to that dimension, and the calibration recommendations can be provided by the agent based on similar historical cases, offering preliminary mediation solutions. Subsequently, the calibration dialogue agent organizes the three-way consensus calibration dialogue involving the student, teacher, and agent. During the dialogue, the student can supplement their self-evaluation basis, the teacher can explain the specific meaning of the scoring criteria, and the agent can retrieve relevant evidence in real time and provide objective reference information. This three-way interactive mechanism promotes understanding and alignment among the evaluation subjects. After the three parties reach a consensus, the calibration dialogue agent outputs the revised scoring result, the reason for the revision, and the scoring rule update signal as the revision result. In some embodiments, the revised scoring result may be a consensus score reached through negotiation, and the reason for the revision may be a structured summary of the disagreement resolution process. The scoring rule update signal is used to indicate any irrationality in the current scoring rules, providing triggering conditions for subsequent rule optimization.

[0033] See Figure 5 The rule-adaptive engine updates the scoring rules adaptively based on the correction results. Existing scoring rules are typically fixed after deployment and cannot learn and evolve from historical assessments. The rule-adaptive engine overcomes this limitation by constructing a closed-loop feedback optimization mechanism, enabling the system to continuously evolve. Specifically, the rule-adaptive engine further dynamically adjusts the weights of evaluation indicators, pass thresholds, grade boundaries, and evidence priorities based on at least two of the following: calibration consensus strength, course characteristics, student profiles, historical calibration stability, and evidence quality. If an evaluation indicator proves difficult to measure objectively in multiple calibrations, the rule-adaptive engine can reduce its weight; if a certain type of evidence plays a crucial role in resolving disagreements, its priority can be increased. For example, the rule-adaptive engine can dynamically adjust the weight of a practical skill from 0.17 to 0.18.

[0034] After the adjustments are completed, the rule self-adaptation engine feeds back the adjusted scoring rules to the competency mapping agent and the assessment scoring agent to update the mapping and evaluation parameters in subsequent course generation and scoring processes. In some embodiments, this feedback mechanism can ensure that the system can apply the optimized rules when processing the next batch of similar courses, thereby reducing repeated scoring disputes and improving overall assessment efficiency.

[0035] See Figure 6 The audit trail agent performs chained hash encapsulation and timestamp signing on natural language prompts, course version generation, rubric mapping logs, scoring records, discrepancy detection results, calibration session records, and rule update records to form an immutable audit trail. Existing systems often lack a full-process traceability mechanism, making it difficult to meet the needs of educational supervision and accountability. The audit trail agent overcomes this limitation and provides highly reliable evaluation records by introducing cryptographic evidence storage technology. In some embodiments, the audit trail agent can employ at least one distributed ledger technology or hash chain list structure to bind the summary information of each key event with the hash value of the preceding event, and attach a timestamp issued by an authoritative timestamp server and the digital signature of the operating entity. This mechanism ensures that any tampering with historical evaluation data can be easily detected, thus providing a solid data trust foundation for educational governance.

[0036] The algorithm engine layer includes a core competency bidirectional mapping algorithm, a curriculum structure generation algorithm, a multi-source scoring difference detection algorithm, a scoring rule self-adaptation algorithm, and a chain-style audit trajectory generation algorithm.

[0037] Specifically, the core competency bidirectional mapping algorithm is as follows: Let r be the semantic vector of the i-th indicator of the student rubric. i The vector of the j-th standard item in the core competency standard library is c. j Then semantic similarity is defined as: ; Further, let the structural coverage factor be Cov. ij It is used to represent the degree of structural consistency of the indicator across four fields: behavioral object, ability type, evidence form, and task context. This is achieved by calculating a comprehensive mapping score M. ij To establish a mapping relationship, the overall mapping score M is calculated. ij The calculation formula is: M ij =α·Sim ij +(1-α)·Cov ij ; Among them, Sim ij Cov represents semantic similarity. ijThe structural coverage factor is represented by α, and the weighting factor is represented by α, where 0 < α < 1; in M ij A positive mapping is established when the maximum mapping score of a core competency item is greater than or equal to the preset first threshold θ1, and missing item completion is triggered when the maximum mapping score of a core competency item is lower than the preset second threshold θ2.

[0038] Let the student self-evaluation vector be S={s} i ,...,s m The teacher rating vector is T={t}. i ,...,t m The intelligent agent's scoring vector is A={a}. i ,...,a m}, then the three-source difference degree D can be defined as: ; Where β1+β2+β3=1.

[0039] The single-dimensional bias rate is: ;

[0040] Among them, F i This represents the full score of the i-th indicator.

[0041] The dynamic preset threshold δ is calculated as follows: δ = δ0 + η1C - η2E + η3V; Where δ0 represents the basic threshold, C represents the course complexity coefficient, E represents the evidence completeness coefficient, V represents the historical calibration variance coefficient, and η1, η2, and η3 represent adjustment parameters.

[0042] The self-adaptive algorithm for the scoring rules is as follows: ;

[0043] Among them, w i (t) represents the weight of the i-th evaluation indicator in round t, q i Denotes the consensus coefficient, d i e represents the difficulty level of the course. i λ represents the evidence quality coefficient, and λ, μ, and ν represent the learning rate parameters.

[0044] The algorithm for generating chain audit trails is as follows: Let the record of the t-th event be E. t Its digest hash is H t ,but: ; Let E be the record of the t-th event. t H t Represents digest hash, Role t Indicates the main role in the event, Timet Indicates a timestamp, Sign t This represents a digital signature, which can form an immutable, verifiable, and searchable chain of events.

[0045] The course structure generation algorithm is used to parse key information such as learning topics, time constraints, objectives, preferences, and output formats from the natural language prompts on the student's end, and automatically generate a structured framework including course outline, stage modules, time allocation, tiered objectives, and task outputs, as detailed below: S1. Classify the natural language prompts input by students (knowledge learning / skills training / project practice / interdisciplinary inquiry), extract key structured elements, including: topic T, total duration H, intensity P, outcome type O, preference dimension F, and output structured element vector E=[T,H,P,O,F].

[0046] S2. Match a standard module library based on the subject area (or use a general split if none exists): Basic Cognition Module, Core Skills Module, Comprehensive Application Module, and Outcome Output Module. Allocate learning hours according to total time: Basic 20%–30%; Core 40%–50%; Application 20%–30%. Generate a module sequence M = [M1, M2, …, M k [, and indicate the learning hours for each module.]

[0047] S3. Sort by progressive logic: Introduction, Principles, Training, Comprehensive, Defense; Slice by week / stage, generating stage number, stage name, core content points, class hours, and deliverables. Output course timeline.

[0048] S4. Learning objectives are broken down into three levels: cognitive objectives, skill objectives, and competency objectives. Cognitive objectives include knowing, understanding, describing, distinguishing, and summarizing. Skill objectives include operating, implementing, designing, debugging, and completing. Competency objectives include reflecting, critiquing, collaborating, taking responsibility, innovating, and standardizing. Furthermore, each objective is observable, verifiable, and assessable.

[0049] S5. Calculate the course complexity coefficient C. The calculation formula is: C = α・number of modules + β・skill ratio + γ・complexity of outcome + δ・collaboration requirements.

[0050] S6. Verification: Total class hours = required student duration, no module overlap or omission, goals are achievable, outputs are deliverable, output the final structured course package: course outline, stage task table, three-level goal list, deliverable list, complexity coefficient C.

[0051] The business service layer includes at least course generation services, competency alignment services, scoring services, calibration services, rule update services, auditing services, and visualization services.

[0052] The application layer includes at least the student end, teacher end, supervisor end, school management end, and the interface layer for connecting to third-party learning platforms.

[0053] Example 2 See Figure 2 This invention discloses a multi-agent collaborative evaluation method for student-generated courses, the method comprising: S1: Receives natural language prompts from the student's input, performs intent recognition and task decomposition, and assigns the corresponding tasks to the corresponding agents. Among them, the course generation agent generates the course outline, chapter structure, stage tasks, learning arrangements and initial self-assessment rubric.

[0054] Specifically, the perception and interaction layer receives natural language prompts from students, the agent collaboration layer parses these prompts, and multiple agents are scheduled to generate a course syllabus and initial self-assessment rubrics. In some embodiments, this step may include capturing student input through at least one user interface and performing necessary format validation and data cleaning. The received natural language prompts are converted into standard data structures that the system can process internally.

[0055] This can further include: a master orchestration agent performing semantic parsing, task decomposition, and process orchestration on natural language prompts; a course generation agent generating the course outline, chapter structure, and stage tasks according to the scheduling of the master orchestration agent; a goal decomposition agent decomposing the course goals into cognitive goals, skill goals, and competency goals; and a rubric generation agent generating the initial self-assessment rubric, which includes at least evaluation indicators, indicator descriptions, evidence requirements, rating levels, and initial weights. In some embodiments, these agents can work collaboratively in parallel or serial manner, exchanging intermediate states through shared memory or message queues.

[0056] S2: The initial self-assessment rubric is semantically vectorized and matched with the standard items in the preset core competency standard framework library to establish a positive mapping relationship. The key items of the core competencies that are not covered are completed in reverse to form a standardized mapping rubric.

[0057] Specifically, the competency mapping agent performs bidirectional mapping and missing item completion between the initial self-assessment rubric and the core competency standard framework library to generate a standardized mapping rubric. In some embodiments, this may further include: semantically vectorizing the evaluation indicators in the initial self-assessment rubric; performing similarity matching between the vectorized evaluation indicators and the standard items in the core competency standard framework library to establish a forward mapping; performing reverse completion on the standard items not covered in the core competency standard framework library to supplement missing indicators; and performing indicator conflict resolution and weight normalization to output the standardized mapping rubric.

[0058] S3: Integrate student self-scores obtained from standardized mapping metric, teacher external scores obtained from formal evaluation rules, and AI agent machine scores to obtain multi-source scores, and determine the difference detection results through a multi-source score difference detection algorithm.

[0059] Specifically, the assessment and scoring module obtains student self-assessment results, teacher peer assessment results, and agent machine assessment results based on standardized mapping rubrics, and performs multi-source scoring difference detection. In some embodiments, the student self-assessment results, teacher peer assessment results, and agent machine assessment results are integrated to construct a multi-source scoring vector; the overall bias rate, dimensional bias rate, structural variability, and evidence consistency index of the multi-source scoring vector are calculated; and the dynamically preset threshold is dynamically adjusted based on course complexity, evidence completeness, and historical calibration variance.

[0060] S4: When the difference detection result is lower than the preset threshold, generate the result; when the difference detection result is higher than the preset threshold, generate a divergence report and initiate a three-party calibration dialogue, and output the corrected scoring result and rule correction signal based on the consensus formed in the dialogue.

[0061] Specifically, when the result of multi-source score difference detection exceeds a dynamically preset threshold, the calibration dialogue agent triggers a three-way consensus calibration involving students, teachers, and the agent, and generates a corrected result. In some embodiments, when the result of multi-source score difference detection exceeds the dynamically preset threshold, a disagreement report, a list of evidence, and calibration recommendations are automatically generated; the three-way consensus calibration dialogue involving students, teachers, and the agent is organized; and after the three-way consensus is reached, the corrected score result, the reason for the correction, and the score rule update signal are output as the corrected result. In some embodiments, if the three parties still cannot reach a consensus after a predetermined number of rounds of dialogue, the system can introduce a higher-level arbitration mechanism.

[0062] S5: Dynamically and adaptively update the scoring rules according to the rule correction signal, encapsulate the entire process event using chain hashing, and generate an auditable evaluation report.

[0063] Specifically, the rule-adaptive engine adaptively updates the scoring rules based on the correction results. In some embodiments, the weights of evaluation indicators, pass thresholds, grade boundaries, and evidence priorities are dynamically adjusted based on at least two of the following: calibration consensus strength, course characteristics, student profiles, historical calibration stability, and evidence quality. The adjusted scoring rules are then fed back to the competency mapping agent and the evaluation scoring agent to update the mapping and evaluation parameters in subsequent course generation and scoring processes. In some embodiments, rule updates can employ at least one reinforcement learning algorithm, using calibration results as reward signals to optimize the scoring strategy.

[0064] Finally, the method may further include: the audit trail agent performing chained hash encapsulation and timestamp signing on natural language prompts, course version generation, rubric mapping logs, scoring records, discrepancy detection results, calibration session records, and rule update records to form an immutable audit trail. In some embodiments, the generated audit trail may be periodically archived to at least one secure cloud storage service and provided with a structured query interface for supervisors to verify.

[0065] Example 3 The system and method disclosed herein can be applied to various educational scenarios. For example, in vocational education, students can input natural language prompts regarding industrial robot programming. The agent collaboration layer can generate a syllabus including modules such as instructional programming, trajectory planning, and safety regulations. The competency mapping agent can automatically supplement missing indicators such as equipment log records. During the assessment phase, if there is a significant discrepancy between student self-assessment and teacher peer assessment regarding safe operating procedures, the calibration dialogue agent can trigger a dialogue to clarify the difference between the execution of safe actions and the expression of safety awareness. Subsequently, the rule self-adaptation engine can dynamically increase the weight of safety-related indicators and add a veto rule. This process not only resolves current scoring disputes but also optimizes the evaluation criteria for similar courses in the future.

[0066] This embodiment illustrates a specific scenario. The server configuration includes: at least 16 CPU cores, at least 32GB of memory, at least 1TB of storage, and support for GPU-accelerated inference. The network environment supports HTTPS communication protocol and WebSocket real-time messaging mechanism. Terminal devices include PCs, tablets, and mobile phones. The software environment includes: an operating system of Linux CentOS 7.6 or later; a backend development framework of Spring Boot or a similar service framework; a frontend framework of Vue.js or a similar framework; and databases of MySQL 8.0 and Redis 6.0. The model service uses a combination of a large language model, a text embedding model, and a scoring model. The audit trail agent can use a consortium blockchain evidence storage component or a chained hash evidence storage component.

[0067] (1) High School Artificial Intelligence and Life Course:

[0068] S1: A high school student enters the following prompt: I want to learn the basics of artificial intelligence, including machine learning, neural networks and real-life application cases. The learning period is 4 weeks, with 3 class hours per week. I hope there will be project practice, results presentation and self-evaluation criteria.

[0069] S2: The course generation agent generates a 4-week course outline, totaling 12 class hours. The course outline structure is shown in Table 1 below.

[0070] Table 1. Four-week course structure

[0071] Week 1 Introduction to AI and Life Analysis of Artificial Intelligence Concepts, Application Scenarios, and Case Studies 3 Concept Map Week 2 Machine Learning Fundamentals Classification, Regression, Samples and Labels 3 Data labeling practice Week 3 Introduction to Neural Networks Neuron, network layer, image recognition example 3 Network Structure Description Week 4 Life Application Projects Campus AI Application Solution Design and Demonstration 3 Project posters and explanatory videos

[0072] Six learning objectives are generated simultaneously: (1) Understand the basic concepts of artificial intelligence and machine learning; (2) Be able to distinguish the input, processing and output of common AI applications; (3) Be able to describe the basic structure of neural networks; (4) Be able to complete a small project design for a life scenario; (5) Be able to reflect on the deviations and ethical issues in AI applications; (6) Be able to conduct self-evaluation based on rubrics and revise the work accordingly.

[0073] Eight initial self-assessment rubrics were generated, as shown in Table 2 below.

[0074] Table 2 Initial Self-Assessment Ratio Table

[0075] R1 Conceptual understanding 0.14 R2 Algorithm Logic Understanding 0.12 R3 Practical skills 0.18 R4 Case analysis skills 0.10 R5 Innovative application capabilities 0.14 R6 Collaborative expression skills 0.10 R7 Ethical reflection ability 0.08 R8 Self-correction ability 0.14

[0076] S3: The eight initial self-assessment rubrics were semantically vectorized and matched with the standard items in the pre-set core competency standard framework library using similarity matching with α=0.72, θ1=0.78, and θ2=0.62 to establish a positive mapping relationship, resulting in the results shown in Table 3 below:

[0077] Table 3. Mapping Relationship Table.

[0078] R1 Scientific spirit - rational thinking 0.89 R2 Scientific spirit - courage to explore 0.84 R3 Practical Innovation - Technology Application 0.91 R4 Scientific spirit - critical questioning 0.81 R5 Practical Innovation - Problem Solving 0.87 R6 Responsibility and Social Responsibility 0.76 R7 Responsibility and Commitment - National Identity / International Understanding 0.73 R8 Learn to learn - Be diligent in reflection 0.90

[0079] Due to insufficient information awareness coverage, the system automatically supplements indicator R9, "Information Retrieval and Discrimination Ability," with a suggested weight of 0.06. After normalization, the standardized mapping metric is shown in Table 4.

[0080] Table 4 Standardized Mapping Gauge

[0081] R1 0.13 R2 0.11 R3 0.17 R4 0.1 R5 0.13 R6 0.10 R7 0.08 R8 0.12 R9 0.06

[0082] S3: The student's self-score obtained according to the standardized mapping metric, the teacher's external score obtained according to the formal evaluation rules, and the agent's machine score are integrated to obtain a three-source score. The difference detection results are determined by the three-source score difference detection algorithm, as shown in Table 5:

[0083] Table 5. Results of Difference Detection R1 88 78 80 R2 82 70 73 R3 92 84 86 R4 85 72 75 R5 90 76 79 R6 87 80 82 R7 79 68 70 R8 91 77 81 R9 75 66 69

[0084] The weighted total scores are as follows: self-assessment 86.87 points; teacher assessment 74.66 points; and machine assessment of the intelligent agent 77.64 points.

[0085] S4: The calculated overall deviation rate between self-assessment and teacher assessment was 12.21%, but significant local differences were found in dimensions R5 and R8, with a comprehensive variance D=18.42, exceeding the system's preset threshold of 16.50, thus triggering third-party calibration. Students submitted project prototypes, explanation videos, and reflection documents as evidence; teachers pointed out deficiencies in the students' theoretical accuracy. The agent suggested splitting R5 into innovative ideas and application demonstrations, and increasing the weight of version iteration evidence in R8.

[0086] S5: Perform dynamic self-adaptive updates, and the total score is corrected to 79.20 after calibration. The rule update results are as follows: the weight of practical operation ability is increased from 0.17 to 0.18; the weight of innovative application ability is decreased from 0.13 to 0.11; and the total weight of sub-items related to theoretical accuracy is increased from 0.24 to 0.27.

[0087] The entire process of events is encapsulated using chained hashing to generate an auditable evaluation report. This embodiment generates a total of 126 audit events, including: 1 input prompt event; 3 course version events; 18 mapping log events; 27 scoring events; 42 calibration events; 9 rule revision events; 2 report export events; and the remainder are signature and timestamp verification events. The chained verification pass rate is 100%.

[0088] Introductory course to industrial robot programming in vocational education S1: A vocational school student inputs the following prompt: "I need to learn the basics of industrial robot programming within 6 weeks, focusing on teachable programming, trajectory planning, and safety regulations. I hope to ultimately complete a material handling simulation, with an emphasis on practical skills and safety awareness during the evaluation." The prompt is then processed for intent recognition and task breakdown. The corresponding tasks are assigned to the appropriate agents. The course generation agent generates a 6-week course outline, totaling 24 class hours, including: Module M1: Robot Structure and Coordinate System.

[0089] Module M2: Teach pendant usage.

[0090] Module M3: Trajectory Planning.

[0091] Module M4: I / O control.

[0092] Module M5: Safety Procedures.

[0093] Module M6: Comprehensive Simulation of Transportation Tasks.

[0094] S2: Initially, 10 evaluation indicators were generated, with a total weight of 0.42 for practical indicators. After mapping, the system found that equipment log records and team handover instructions were missing in the original metric and automatically added them as new indicators. After addition, the total number of indicators reached 12. Some key indicators after standardization are shown in Table 6 below:

[0095] Table 6. Some key indicators after standardization Practical Implementation 0.18 Safety regulations 0.14 Trajectory optimization 0.12 Troubleshooting 0.11 Log recording 0.07 Handover instructions 0.06

[0096] S3: The student's self-score obtained from the standardized mapping metric, the teacher's external score obtained from the formal evaluation rules, and the agent's machine score are integrated to obtain a three-source score. The difference detection result is determined by the three-source score difference detection algorithm. The comprehensive simulation score data of the three-source scores is shown in Table 7 below:

[0097] Table 7. Comprehensive Simulation Scoring Data for Three Sources Teaching Programming Accuracy 93 88 90 trajectory smoothness 90 79 83 Safe operating procedures 95 76 80 Task completion efficiency 88 82 84 Troubleshooting capabilities 84 71 73 Log record completeness 72 65 68 Team handover clarity 78 70 72

[0098] S4: The total score is: self-evaluation 88.54 points; teacher evaluation 77.36 points; intelligent agent machine evaluation 80.11 points. The system identified that the maximum difference in the safety operation specification dimension was 19 points, which was 12 points higher than the course dynamic preset threshold, triggering a three-party calibration. After calibration, it was found that students understood "no accident occurred" as fully meeting the standard, while teachers scored according to the four detailed rules of "predicting risks, executing commands, stopping the machine and confirming the area isolation". Based on this, the system performed the following updates: (1) The safety operation specification was split into the execution of safety actions and the expression of safety awareness. (2) The total weight of safety indicators was increased from 0.14 to 0.20. (3) An excellent judgment constraint was added: if the safety item score is lower than 75 points, it cannot be rated as excellent. (4) The completeness of log records was included in the passability indicator. The final score after correction was 81.43 points.

[0099] S5: A test was conducted on 36 students and 72 comprehensive assignments. The results showed that traditional individual teacher grading took an average of 18.6 minutes per assignment. After using the system of this invention, the average time was 6.9 minutes per assignment. The proportion of assignments with discrepancies exceeding the threshold requiring calibration was 22.2%. After calibration, the teacher review time decreased by 61.8% compared to purely manual review.

[0100] (3) Data Journalism Visualization Course in Universities S1: An undergraduate team inputs the following prompt: We want to complete a data journalism visualization course within 8 weeks. The content should include data collection, cleaning, narrative structure, and interactive chart design. The final product should be a data story addressing campus issues. The evaluation should take into account technical implementation, journalistic ethics, and teamwork.

[0101] S2: The course generation agent generates an 8-week course outline, including: W1: Topic selection and topic justification; W2: Data source retrieval and legality review; W3: Data cleaning and structuring; W4: Narrative framework design; W5: Visual coding; W6: Interaction design; W7: Ethical review and revision; W8: Publication and defense.

[0102] Eleven evaluation indicators were generated simultaneously, including data authenticity, legality of data collection, quality of data cleaning, narrative integrity, clarity of visual expression, interactive experience, issue insight, teamwork, news ethics, iterative revision ability, and public expression ability.

[0103] S3: The system identified that news ethics are mapped to the dimension of responsibility and accountability. Issue insights are mapped to the combined dimensions of humanistic qualities and scientific spirit. The interactive experience does not adequately match the core competency standards; therefore, the system automatically supplemented the audience comprehensibility sub-indicator. The team's completed work was scored as shown in Table 8 below.

[0104] Table 8 Three-Source Scoring Table Data authenticity 94 91 90 Legality of data collection 89 85 87 Cleaning quality 86 80 83 Narrative integrity 92 78 81 Visual expression 90 82 85 Interactive experience 91 79 84 Issue Insights 93 76 78 Teamwork 95 88 86 Journalism Ethics 88 73 79 Iterative revision 91 77 80 Public expression 87 81 83

[0105] S4: Total Score: Self-assessment 90.67; Teacher's peer assessment 80.02; Machine assessment 82.63. The system identified high-discrepancy clusters in the three dimensions of "Issue Insight," "Journalistic Ethics," and "Iterative Revision." Calibration indicates that the team placed more emphasis on the final product presentation, while the teacher focused more on the overall process standards and intermediate version evidence.

[0106] S5: Based on this, the system implemented the following rule adjustments: the total weight of outcome-based indicators decreased from 0.58 to 0.49. The total weight of process-based indicators increased from 0.42 to 0.51. A bottom-line rule was set for news ethics: when the score is below 75, the overall rating cannot be higher than "Good". The weight of iterative revision capability was increased from 0.07 to 0.10, and a version difference rate factor was introduced. After the corrections, the team's final score was 83.95, with a rating of "Good". The system generated a 27-page "Course Consensus Assessment Audit Report", which included: course version tree; indicator mapping matrix; difference heatmap; calibration dialogue excerpt; rule update log; hash checksum; and supervisor query index page.

Claims

1. A multi-agent collaborative evaluation system for student self-generated courses, characterized in that, include: The perception and interaction layer is used to receive natural language prompts from the student's input. The intelligent agent collaboration layer includes a master orchestration agent, a curriculum generation agent, a competency mapping agent, an assessment and scoring agent, a calibration dialogue agent, and an audit trail agent; The algorithm engine layer includes a core competency bidirectional mapping algorithm, a curriculum structure generation algorithm, a multi-source scoring difference detection algorithm, a scoring rule self-adaptation algorithm, and a chain-based audit trajectory generation algorithm; The process involves the following steps: First, the master orchestration agent schedules the course generation agent to generate a course outline, chapter structure, stage tasks, learning arrangements, and initial self-assessment rubrics based on natural language prompts. Second, the competency mapping agent maps the initial self-assessment rubrics to at least one dimension within the core competency standard framework using the core competency bidirectional mapping algorithm, generating a standardized mapping rubric. Third, the assessment and scoring agent integrates student self-assessments, teacher assessments, and agent machine assessments to obtain multi-source scores, and determines the difference detection results using the multi-source score difference detection algorithm. Fourth, the calibration dialogue agent generates a disagreement report, a list of evidence, and calibration suggestions when the difference detection results exceed a preset threshold, and initiates a three-party calibration dialogue, outputting a rule correction signal. Fifth, the scoring rule self-adaptation algorithm outputs the corrected scoring results and rule correction signals based on the consensus formed during the dialogue. Sixth, the audit trail agent performs chained hash encapsulation and timestamp signing to generate an auditable assessment report.

2. The multi-agent collaborative evaluation system of claim 1, wherein, The natural language prompts include at least one or more of the following: learning topic, learning objective, learning duration, learning preferences, outcome format, and evaluation preferences.

3. The multi-agent collaborative evaluation system of claim 1, wherein, The system also includes a data resource layer for storing and managing a student profile database, a course template and course example database, a self-assessment rubric database, a core competency standard framework database, a historical scoring database, a calibration session log database, an audit track database, and an evidence material database.

4. The multi-agent collaborative evaluation system of claim 1, wherein, The system also includes a business service layer and an application layer; the business service layer includes at least course generation service, competency alignment service, scoring service, calibration service, rule update service, audit service, and visualization service; the application layer includes at least student client, teacher client, supervisor client, school management client, and an interface layer for connecting to third-party learning platforms.

5. A multi-agent collaborative evaluation method for student self-generated courses, characterized in that, The method includes: S1: Receive natural language prompts from the student's input, perform intent recognition and task decomposition, and assign the corresponding tasks to the corresponding intelligent agents. Among them, the course generation intelligent agent generates the course outline, chapter structure, stage tasks, learning arrangements and initial self-assessment rubric. S2: The initial self-assessment rubric is semantically vectorized and matched with the standard items in the preset core competency standard framework library to establish a positive mapping relationship. The key items of the core competencies that are not covered are completed in reverse to form a standardized mapping rubric. S3: Integrate student self-scores obtained from standardized mapping metric, teacher external scores obtained from formal evaluation rules, and AI agent machine scores to obtain multi-source scores, and determine the difference detection results through a multi-source score difference detection algorithm; S4: When the difference detection result is lower than the preset threshold, generate the result; when the difference detection result is higher than the preset threshold, generate the divergence report and initiate a three-party calibration dialogue, and output the corrected scoring result and rule correction signal based on the consensus formed in the dialogue. S5: Dynamically and adaptively update the scoring rules according to the rule correction signal, encapsulate the entire process event using chain hashing, and generate an auditable evaluation report.

6. The multi-agent cooperative evaluation method according to claim 5, characterized in that, In step S2, a mapping relationship is established by calculating a comprehensive mapping score M ij The calculation formula of the comprehensive mapping score M ij is: M ij = a - Sim ij + (1 - a) - Cov ij ; wherein Sim ij represents semantic similarity, Cov ij represents structural coverage coefficient, and a represents weight coefficient; in M ij greater than or equal to a preset first threshold value, and when the maximum mapping score of the core literacy item is lower than a preset second threshold value, triggering missing item completion.

7. The multi-agent collaborative evaluation method according to claim 5, wherein, In step S3, the calculation formula of the multi-source scoring difference detection algorithm is as follows: ; Where D={} represents the multi-source rating dissimilarity, S={s i ,...,s m } represents the student's self-assessment vector, T={t i ,...,t m } represents the teacher's rating vector, A={a i ,...,a m } represents the machine score vector; The single-dimensional bias rate is: ; where F i represents the full score of the i-th indicator.

8. The multi-agent collaborative evaluation method of claim 5, wherein, In step S4, the preset threshold δ is calculated as follows: δ = δ0 + η1C - η2E + η3V; Where δ0 represents the basic threshold, C represents the course complexity coefficient, E represents the evidence completeness coefficient, V represents the historical calibration variance coefficient, and η1, η2, and η3 represent adjustment parameters.

9. The multi-agent collaborative evaluation method of claim 5, wherein, In step S5, the scoring rules are dynamically updated using a scoring rule self-adaptation algorithm, as shown in the following expression: ; Among them, w i (t) represents the weight of the i-th evaluation indicator in round t, q i Denotes the consensus coefficient, d i e represents the difficulty level of the course. i λ represents the evidence quality coefficient, and λ, μ, and ν represent the learning rate parameters.

10. The multi-agent collaborative evaluation method of claim 5, wherein, The method further includes: using an audit trail agent to perform chained hash encapsulation and timestamp signature on prompt words, course versions, mapping logs, scoring records, discrepancy reports, calibration sessions, rule revisions, and final reports to generate an auditable evaluation report, as shown in the following expression: ; Let E be the record of the t-th event. t H t Represents digest hash, Role t Indicates the main role in the event, Time t Indicates a timestamp, Sign t This refers to a digital signature.