Online education path optimization method and system based on socrates learning situation four tuple perception

By constructing the Socratic Learning Awareness Quadruple Perception System, the shortcomings of online education platforms in learning awareness and dynamic teaching decision-making were addressed. This enabled personalized learning path planning and adaptive teaching loops, improving the accuracy of teaching feedback and the ability to dynamically adjust.

CN122114319AInactive Publication Date: 2026-05-29SICHUAN QIMINGDAREN TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN QIMINGDAREN TECH CO LTD
Filing Date
2026-04-24
Publication Date
2026-05-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing online education platforms are inadequate in terms of deep understanding of students' learning progress and dynamic teaching decisions. They are unable to identify students' learning bottlenecks, lack granular learning information, provide insufficient teaching feedback, and lack unified quantitative indicators for learning progress and cognitive status. The teaching process also lacks a closed-loop mechanism, making it difficult to achieve personalized path planning.

Method used

By constructing a Socratic learning situation quadruple perception system, including logarithmic round Socratic dialogue encoding, dialogue graph construction, learning situation quadruple extraction model, and multimodal learning situation state vector fusion, the learning situation quadruple extraction model is used to extract emotional, cognitive, and behavioral features, construct a multi-objective comprehensive function, optimize the learning path, and form a closed-loop feedback mechanism.

Benefits of technology

It improved the accuracy of student learning checkpoint recognition, enabled personalized learning path planning, enhanced the accuracy and dynamic adjustment capabilities of teaching feedback, formed an adaptive teaching closed loop, and balanced teaching benefits and student experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114319A_ABST
    Figure CN122114319A_ABST
Patent Text Reader

Abstract

The application discloses an online education path optimization method and system based on Socrates learning condition four-tuple perception, which comprises the following steps: constructing a dialogue graph containing time, question and follow-up question relationship, and establishing a learning condition four-tuple extraction model to convert unstructured dialogue into structured information such as learning object, dimension, student original evidence and learning condition polarity; fusing emotion, cognitive and behavioral characteristics to generate a multi-modal unified learning condition state vector; calculating action utility based on a strategy action set, selecting the optimal teaching strategy, and performing path optimization when the trigger condition is met; constructing a knowledge point level vector through global learning condition aggregation, combining a multi-objective comprehensive function and a group iterative search algorithm to generate a personalized learning path, and adjusting the strategy parameter set in a closed loop. The application realizes closed-loop adaptive teaching of dialogue collection, state estimation, strategy selection, path optimization and re-dialogue, improves the accuracy of card point identification and the effectiveness of teaching intervention, and is suitable for online education scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of online education technology, and in particular to an online education path optimization method and system based on Socrates learning information quadruple perception. Background Technology

[0002] As online education platforms continue to expand, teaching models have evolved from initially providing uniform content to dynamic, adaptive teaching that relies on real-time data. Early systems primarily assessed students' learning status based on structured data such as scores, accuracy rates, and time taken. While these metrics are easy to collect, they only reflect students' final performance and fail to reveal their true thoughts or cognitive difficulties during the learning process.

[0003] While current online education platforms have made significant progress in data collection, intelligent explanation, and question recommendation, they still fall short in two core capabilities: deep understanding of student emotions and dynamic instructional decision-making. Traditional understanding of student emotions remains at a coarse-grained level, unable to pinpoint specific learning difficulties. Most existing systems rely on sentiment classification models to label student texts as positive, negative, or neutral. This coarse-grained approach fails to differentiate whether a student is stuck on conceptual understanding, calculation steps, strategy selection, or psychological stress, nor can it identify specific misunderstandings made by students in particular steps of a problem. This lack of fine-grained information makes it difficult for existing systems to provide targeted instructional feedback.

[0004] In recent years, large-scale model-based dialogue interaction has been applied in intelligent teaching scenarios, with the most typical form being "Socratic dialogue interactive teaching." In traditional educational contexts, Socratic dialogue emphasizes guiding students' thinking through continuous questioning. Although existing teaching systems using "large-scale model + Socratic dialogue" can collect high-density student natural language expressions, they lack methods to transform unstructured dialogues into computable, traceable, and personalized path planning-enabled structured emotional signals. This results in dialogue cues being unable to be understood by the system and used for decision-making.

[0005] In addition, most platforms store three types of data—behavioral logs, question structure, and dialogue text—independently, lacking a unified modeling framework to map them to the same feature space. As a result, it is difficult to reveal the coupling relationship between "emotion-cognition-knowledge" and to dynamically adjust teaching strategies accordingly.

[0006] Existing path planning algorithms rely solely on hard metrics such as score, difficulty, and time consumption. They have not yet established unified quantitative metrics for soft factors and cognitive states, such as emotional stability, changes in confidence, and confusion about knowledge points. Consequently, they cannot achieve real-time dynamic adjustment in response to changes in students' psychological and cognitive states.

[0007] Finally, the existing teaching process is a one-way open loop of "dialogue → assessment → planning", lacking a closed loop mechanism of "dialogue acquisition → state estimation → strategy selection → path optimization → dialogue regeneration". The system cannot continuously use students' natural expressions to correct the teaching direction, nor can it actively change the learning state through strategy intervention, thus limiting adaptive teaching.

[0008] The "Socratic dialogue" in this technology refers to a multi-round question-and-answer process automatically initiated by the dialogue model within the system. The system does not provide a complete answer all at once, but rather generates follow-up questions or hints based on the student's current response, guiding the student to gradually explain their thought process and intermediate steps. This type of dialogue has the following technical characteristics: each round of question-and-answer is recorded in a structured format, including timestamps, role labels, and contextual relationships, which can serve as input data for subsequent algorithms, rather than merely representing a teaching philosophy.

[0009] Therefore, there is an urgent need to propose a logically simple, accurate and reliable method and system for optimizing online education paths based on the Socratic learning information quadruple perception. Summary of the Invention

[0010] To address the aforementioned problems, the present invention aims to provide an online education path optimization method and system based on Socratic learning information quadruple perception. The technical solution adopted by the present invention is as follows: The first part of this technology provides an online education path optimization method based on the Socratic learning information quadruple perception, which includes the following steps: Step S1: Encode several rounds of Socratic dialogue and construct a dialogue diagram that includes time, questioning, and follow-up questioning relationships; Step S2: Label relationships on the dialogue graph and construct a learning information quadruple extraction model; the learning information quadruple includes learning object, learning dimension, student verbal evidence, and learning information polarity; Step S3: The emotional features, cognitive features, and behavioral features extracted by the learning situation quadruple extraction model are fused to construct a multimodal unified learning situation state vector. Step S4: Preset a set of strategy actions and a set of strategy parameters corresponding to the set of strategy actions, and calculate the action utility value of each strategy action in the set of strategy actions based on the multimodal unified learning state vector, and select the strategy action with the highest action utility value to execute. When at least one of the preset cycle triggering condition, state offset triggering condition, or risk event triggering condition is met, the learning path optimization is triggered and the process proceeds to step S5; otherwise, the process returns to step S1 and proceeds to the next round of Socratic dialogue. Step S5: For the multimodal unified learning status vectors corresponding to multiple rounds of Socratic dialogues on the same knowledge point, perform global learning information aggregation to form a global learning information vector for that knowledge point; Based on a preset set of knowledge points, filter and sort the knowledge points in the set of knowledge points according to the global learning information vector to construct a set of candidate learning paths. Step S6: Construct a multi-objective comprehensive function; for each knowledge point of the candidate learning path in the candidate learning path set, calculate the mastery gap cost, emotional risk cost, fatigue load cost, and time cost based on the global learning situation vector, and accumulate them to form the comprehensive objective value of the path; under the condition of satisfying the preset learning time budget constraint and knowledge point coverage constraint, use a path update algorithm based on group iterative search to obtain the optimal solution of the multi-objective comprehensive function, and output the personalized learning path; Step S7: Based on the optimal solution of the multi-objective synthesis function of the current round, perform closed-loop feedback on the set of strategy parameters in step S4.

[0011] The second part of this technology provides a system for optimizing online education pathways using a Socratic learning information quadruple perception method, which includes: The dialogue graph construction module encodes several rounds of Socratic dialogues and constructs a dialogue graph that includes time, questioning, and follow-up questioning relationships. The learning information quadruple extraction model is connected to the dialogue graph construction module. Relationships are labeled on the dialogue graph to construct the learning information quadruple extraction model. The learning information quadruple includes learning object, learning dimension, student verbal evidence, and learning information polarity. The multimodal unified learning state vector construction module, together with the learning state quadruple extraction model, integrates the emotional features, cognitive features, and behavioral features extracted by the learning state quadruple extraction model to construct a multimodal unified learning state vector. The action utility calculation module is connected to the multimodal unified learning state vector construction module, the global learning vector solving module, and the dialogue graph construction module. It presets a set of strategy actions and a set of strategy parameters corresponding to the set of strategy actions, and calculates the action utility value of each strategy action in the set of strategy actions based on the multimodal unified learning state vector. The strategy action with the highest action utility value is selected for execution. When at least one of the preset periodic triggering condition, state offset triggering condition, and risk event triggering condition is met, the learning path optimization is triggered and the module enters the global learning vector solving module; otherwise, the module returns to the dialogue graph construction module to enter the next round of Socratic dialogue. The global learning vector calculation module is connected to the action utility calculation module. It performs global learning aggregation on the multimodal unified learning state vector corresponding to multiple rounds of Socratic dialogue for the same knowledge point to form a global learning vector for that knowledge point. Based on a preset set of knowledge points, it filters and sorts the knowledge points in the set according to the global learning vector to construct a set of candidate learning paths. A multi-objective comprehensive function module is connected to a global learning vector solving module to construct a multi-objective comprehensive function. For each knowledge point of a candidate learning path in the candidate learning path set, the module calculates the mastery gap cost, emotional risk cost, fatigue load cost, and time cost based on the global learning vector, and accumulates them to form the comprehensive objective value of the path. Under the condition of satisfying the preset learning time budget constraint and knowledge point coverage constraint, the optimal solution of the multi-objective comprehensive function is obtained by using a path update algorithm based on group iterative search, and a personalized learning path is output. The closed-loop feedback module is connected to the multi-objective synthesis function module and the action utility calculation module. It provides closed-loop feedback to the strategy parameter set in the action utility calculation module based on the optimal solution of the multi-objective synthesis function in the current round.

[0012] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention encodes several rounds of Socratic dialogue and constructs a dialogue graph containing time, question and follow-up relationship. By constructing the dialogue graph structure and introducing graph convolution processing, the system can more accurately capture the student's real position in the context, thereby improving the accuracy of checkpoint recognition.

[0013] (2) The present invention performs relation labeling on the dialogue graph and constructs a learning information quadruple extraction model. Through the modified quadruple extraction model, the natural language dialogue is transformed into structured learning information, including learning objects (knowledge points, question types, step positions, learning dimensions (concept understanding, calculation ability, strategy selection, time pressure, emotional state, etc.), student opinions or evidence phrases, learning polarity (mastery / not mastery, etc.), and the quadruple as the basic unit for subsequent state vector modeling.

[0014] (3) This invention utilizes the emotional features, cognitive features, and behavioral features extracted by the learning situation quadruple extraction model to construct a multimodal unified learning situation state vector. This vector can fully describe the student's real learning state at a certain moment and is the direct basis for strategy control.

[0015] (4) This invention adopts global learning information aggregation and path planning optimization. The state vector of a single dialogue will be aggregated to the knowledge point level to update the student's long-term learning information data. Then, a personalized learning path is generated through a multi-objective optimization algorithm. The goal is to take into account the improvement of mastery, emotional stability, reasonable learning load, learning time constraints, and the fact that the path optimization result will affect the guidance method of the next round of Socratic dialogue, thus forming a complete closed loop.

[0016] (5) The present invention obtains the action utility corresponding to any strategy action in the set of strategy actions by pre-setting a set of strategy actions and based on the multimodal unified learning state vector; selects the strategy action corresponding to the highest action utility at present according to the action utility, and introduces the framework of action value + action cost, which can balance teaching benefits and student experience.

[0017] In summary, this invention has the advantages of simple logic and high accuracy and reliability, and has high practical and promotional value in the field of online education technology. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope of protection. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a logic flowchart of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the present invention will be further described below with reference to the accompanying drawings and embodiments. The embodiments of the present invention include, but are not limited to, the following embodiments. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0021] like Figure 1 As shown, this embodiment provides an online education path optimization method based on Socratic learning information quadruple perception. It continuously collects key information from students' expressions through Socratic dialogue and transforms this information into structured signals that the system can calculate and execute, forming a fully adaptive teaching closed loop from dialogue to path planning.

[0022] The first step involves encoding several rounds of Socratic dialogue and constructing a dialogue graph that includes time, questioning, and follow-up questions. Here, the original dialogue sequence is constructed as a directed graph, and contextual relationships are encoded using graph convolution. This includes the following steps: (11) Text encoding is performed using a text encoding layer: Suppose there are a total of 1 dialogue, number ; For the Dialogue text Encode and construct text vectors. ;in, This represents a text encoding function, which can correspond to the encoding layers of RoBERTa and Qwen, and outputs a vector of fixed dimensions. Indicates the first The initial vector of each dialogue node; Indicates the number of dialogue nodes.

[0023] (12) Construct a dialogue graph, which includes time sequence edges, question pointing edges, and answer pointing edges.

[0024] (121) The establishment of time-sequence edges includes: if the first edge is... The dialogue node immediately follows the first one in time. If there are 10 dialogue nodes, then: ; in, An adjacency matrix representing temporal order relationships.

[0025] (122) The establishment of the question-pointing edge includes: if the first... The first dialogue node is a question issued by the system, and the second... If each dialogue node corresponds to a student's answer, then: ; in, An adjacency matrix representing the relationship between system questions and student answers.

[0026] (123) The establishment of the answer-directing edge includes: when the system generates subsequent questions based on preset judgment conditions in response to the student's answer in order to obtain the student's supplementary explanations, specific evidence, or intermediate reasoning process for the existing answer, an answer-directing follow-up question relationship is established, then: ; in, This represents the adjacency matrix that indicates the relationships between student responses and subsequent system follow-up questions.

[0027] Here, the follow-up question relationship is triggered when at least one of the following conditions is met: the student's answer does not contain complete problem-solving steps or intermediate reasoning information; the student's answer contains multiple possible interpretations, and the system cannot uniquely determine the learning object or problem-solving approach it refers to; the student's answer is inconsistent with the expected problem-solving path or known intermediate conclusions; the student's answer contains expressions of uncertainty (possibly, probably, not quite certain, etc.); the system determines based on the historical dialogue context that it needs to further obtain the student's explanation or reasoning. When any of the above conditions are met, the system automatically generates a follow-up question in response to the student's answer and uses this question as a follow-up question node, establishing a response-to-follow-up question relationship with the original student answer node.

[0028] (124) Merge the time sequence edges, question-pointing edges, and answer-pointing edges to obtain the adjacency matrix of the dialogue graph. Its expression is: ;in, .

[0029] (13) Perform convolutional encoding on the dialogue graph. In order to further characterize the semantic dependencies between nodes, the constructed dialogue graph is subjected to graph convolution processing.

[0030] The first layer of graph convolution uses the following formula: ; The second layer of graph convolution uses the following formula: ; The final node representation is as follows: ; in, This represents the final embedding vector of dialogue node i in the dialogue graph structure; Represents a dialogue node The vector after the second layer of graph convolution; Indicates the activation function; The adjacency matrix of the dialogue graph represents the first... and Each element, if a dialog node With dialogue nodes If there is an edge in any of the relationships of time sequence, question, or follow-up question, the value is 1; otherwise, the value is 0. Represents a dialogue node The vector after the first layer of graph convolution; This represents the weight matrix of the second-layer graph convolution, used to perform a linear mapping on the node vectors; This represents the weight matrix of the first layer graph convolution, used to perform a linear mapping on the node vectors.

[0031] The second step involves labeling relationships on the dialogue graph to construct a learning information quadruple extraction model. The learning information quadruple includes the learning object, learning dimension, student verbal evidence, and learning information polarity. Specifically, this includes the following steps: (21) For any pair of dialogue nodes Constructing relational feature vectors Its expression is: ; in, Represents a dialogue node The vector after convolution of the first layer graph; This indicates a vector concatenation operation.

[0032] (22) Obtain the relation label scoring vector using nonlinear mapping: ; in, Indicates dialogue node pair Confidence scores belonging to different categories of learning information field relationships; Represents a non-linear activation function; This represents the weight matrix of the relational mapping layer; This represents the bias vector of the relational mapping layer.

[0033] (23) Based on the dialogue node pair Confidence scores belonging to different learning information field relationship categories The grid annotation matrix elements are obtained by annotation: ;in, Indicates dialogue node pair The relationship score value on any type of learning information field. Each element is a relationship label score vector, used to represent the score of the corresponding node pair on the relationship of various types of learning information fields; the learning information fields include at least the learning target, learning aspect, student verbal evidence, and learning polarity.

[0034] This embodiment adopts a fixed process of stable and controllable candidate interval → selecting the optimal span → mapping text → assembling quadruples, to facilitate subsequent parameter tuning. The expression for filtering the candidate interval is as follows: ; in, Field representing learning situation The set of candidate intervals; Indicates dialogue node pair In the learning situation field The relationship score; This indicates the target field for learning information. Preset confidence threshold; The values ​​can be Target (learning target), Aspect (learning dimension), Opinion (student's original evidence), or Polarity (learning polarity).

[0035] Furthermore, the optimal span interval is selected, and its expression is: ; in, Indicates in the learning situation field The optimal span range within.

[0036] Mapped text: will be in the learning information field The optimal span range within Mapping back to the original dialogue text: ;in, Field representing learning situation The actual content in the original dialogue text; This represents a sequence of dialogue text.

[0037] Finally, assemble the quadruple: from the learning information field The actual content in the original dialogue text The model for extracting learning information using a four-tuple is constructed, and its expression is as follows: ; in, This represents the k-th learning information quadruple; This represents the field content of the learning object in the k-th learning information quadruple; This represents the field content of the learning dimension in the k-th learning information quadruple; The field representing the original words of the students in the k-th learning information quadruple; The field representing the polarity of the learning information in the k-th learning information quadruple.

[0038] The third step involves fusing the emotional, cognitive, and behavioral features extracted by the learning situation quadruple extraction model to construct a multimodal unified learning situation state vector, which includes the following steps: (31) Extract emotional features, the expression of which is: ;in, This represents an emotion feature vector, used to express the emotional state of students in recent conversations, such as positive, neutral, anxious, tired, uncertain and other multi-dimensional components; This represents the emotion feature extraction function, from which emotion-related Polarity information is extracted; The learning situation is represented by a quadruple.

[0039] (32) Extract cognitive features, the expression of which is: ;in, Represents the feature vector of cognitive state; This represents a cognitive understanding analysis function that extracts the Target and Aspect fields from the four-tuple, corresponding to cognitive dimensions such as understanding, transfer, computation, and strategy selection. (33) Construct behavioral characteristics, including the time taken for the current question step, the number of prompts, the number of incorrect steps, whether skipping steps, whether repeated queries, etc., the expression of which is: ;in, Represents a behavioral feature vector; Indicates the time elapsed for the current step; This indicates the number of prompts, used to measure whether students rely on prompts. This indicates the number of incorrect steps, used to illustrate the frequency of local errors made by students in the current problem.

[0040] (34) Construct a unified learning state vector for multimodal learning Its expression is: ; in, This represents the feature fusion matrix, used to map vectors from three different sources to a unified space (the dimension can be set by the user, such as 32, 64, or 128).

[0041] The fourth step involves presetting a set of policy actions and a set of policy parameters for policy configuration, and calculating the action utility value of each policy action in the set of policy actions based on the multimodal unified learning state vector. The policy action with the highest action utility value is then selected for execution, including the following steps: (41) Preset strategy action set For Socratic teaching scenarios, a finite set of strategy actions is predefined, expressed as: ; in, This indicates a continued exploration of the questioning action, guiding students to further explain their thought process or supplement their reasoning; This indicates a prompting action, such as emphasizing key conditions or pointing out common errors; This indicates a summary action of the currently completed content; This indicates a teaching action that directly explains the current step or the entire question, and is used to prevent students from getting stuck for too long; This indicates a teaching action that involves switching questions or adjusting the difficulty of learning tasks, such as changing to a similar but simpler question.

[0042] (42) Based on the multimodal unified learning state vector Estimate the set of selected strategy actions The action value of any action in the sequence is expressed as: ; in, Represents the set of policy actions The first in Each strategic action, ; Representation and policy action set The first in Each strategy action The corresponding value weight vector; T represents the matrix transpose operation; Indicates the set of policy actions The first in Each strategy action The theoretical payoff score under the current state.

[0043] (43) For any set of policy actions The first in Each strategy action The dynamic cost is obtained by the following expression: ; in, This represents the set of strategy actions under the current learning situation. The first in Each strategy action The dynamic cost; Represents the set of policy actions The first in Each strategy action The basic cost; Indicates the emotional sensitivity coefficient; This represents an index of negative emotion intensity calculated based on dialogue content and behavioral data. This represents the cognitive sensitivity coefficient, which controls the degree to which positive cognitive signals reduce costs. This indicates positive cognitive indicators, such as the student's mastery of the current knowledge point.

[0044] Here. The more negative the students' emotions ( The higher the cognitive level, the greater the overall cost of the action, and the more cautious the system will be in using "pressure-based" actions (such as delving deeper or tackling more complex problems). The more positive the student's cognition (…), the greater the overall cost of the action. The higher the difficulty, the lower the cost of some actions, and the more the system tends to increase the challenge or continue to delve deeper.

[0045] (44) The action utility corresponding to any strategy action in the set of strategy actions is obtained, and its expression is: ; in, This represents the set of strategy actions under the current learning situation. The first in Each strategy action The overall utility value.

[0046] (45) Use strategy selection to obtain the optimal strategy action. Its expression is: .

[0047] The fifth step involves globally aggregating the multimodal unified learning status vectors corresponding to multiple rounds of Socratic dialogues on the same knowledge point to form a global learning vector for that knowledge point. Based on a preset set of knowledge points, the knowledge points in the set are filtered and sorted according to the global learning vector to construct a set of candidate learning paths.

[0048] Here, the expression for the global learning vector of any knowledge point is: ; in, Indicates the first Global learning progress vector for each knowledge point; This indicates the number of Socratic dialogue rounds related to the knowledge point within the current analysis window, and it is an integer greater than or equal to 1. This represents the multimodal unified learning state vector corresponding to the Socratic dialogue in the kth round that is related to the knowledge point.

[0049] The sixth step is to construct a multi-objective comprehensive function. For each knowledge point in the candidate learning path set, the mastery gap cost, emotional risk cost, fatigue load cost, and time cost are calculated based on the global learning situation vector and accumulated to form the comprehensive objective value of the path. Under the condition of satisfying the preset learning time budget constraint and knowledge point coverage constraint, the optimal solution of the multi-objective comprehensive function is obtained by using a path update algorithm based on group iterative search, and the personalized learning path is output.

[0050] (61) Construct a multi-objective synthesis function Its expression is: ; in, This represents the loss term related to knowledge mastery; This indicates a loss of emotional stability. This represents the learning load loss term, used to characterize the intensity of cognitive load reflected by behavioral data during the learning process; This represents the total learning time loss item; This represents the loss item related to knowledge mastery. Weighting coefficients; Items representing loss of emotional stability Weighting coefficients; Represents the learning load loss term Weighting coefficients; This represents the total learning time loss item. The weighting coefficients.

[0051] (62) Suppose that in the t-th iteration, there are a total of There are 10 candidate learning paths, denoted as: ;in, In the t-th iteration, the first... There are 10 candidate learning paths.

[0052] (63) Global guidance and random perturbation path update are adopted: ; in, In the (t+1)th iteration, the th... Candidate learning paths; Denotes the multi-objective synthesis function in the t-th iteration. The shortest path; This represents the global guiding weight parameter; Indicates the random disturbance intensity coefficient; This represents a random perturbation vector with the same dimension as the path sequence.

[0053] (64) Based on the multi-objective synthesis function Find the optimal path: .

[0054] The seventh step is to apply closed-loop feedback to the strategy parameter set based on the optimal solution of the multi-objective synthesis function in the current round.

[0055] The strategy model should dynamically adjust the strategy parameters for the next round based on the path optimization results to better align with the long-term path objectives. Its expression is: ; in, This represents the set of policy parameters in the t-th iteration, including the set of policy actions. The first in Each strategy action Corresponding value weight vector Emotional sensitivity coefficient (Used to characterize the strength of the strategy's response to changes in the learner's emotional state), cognitive sensitivity coefficient (Used to characterize the strength of the strategy's response to changes in learners' cognitive load), negative emotion intensity index (Used to represent the degree of influence of negative emotions on strategy decisions during the learning process), cost parameter vector (Used to characterize the weighting of various cost factors during strategy execution). This represents the set of policy parameters in the t=1th iteration. This represents the feedback update rate, which is a positive number. It is usually set to a value less than 1 and is used to control the step size of each round of policy updates to prevent the policy from becoming unstable due to excessively rapid adjustments. This indicates the sensitivity of the multi-objective synthesis function to policy parameters, and how the policy affects the long-term objective.

[0056] The above embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any changes made based on the design principles of the present invention, or any non-creative modifications made thereon, shall fall within the scope of protection of the present invention.

Claims

1. An online education path optimization method based on Socratic learning information quadruple perception, characterized in that, Includes the following steps: Step S1: Encode several rounds of Socratic dialogue and construct a dialogue diagram that includes time, questioning, and follow-up questioning relationships; Step S2: Label relationships on the dialogue graph and construct a learning information quadruple extraction model; the learning information quadruple includes learning object, learning dimension, student verbal evidence, and learning information polarity; Step S3: The emotional features, cognitive features, and behavioral features extracted by the learning situation quadruple extraction model are fused to construct a multimodal unified learning situation state vector. Step S4: Preset a set of strategy actions and a set of strategy parameters corresponding to the set of strategy actions, and calculate the action utility value of each strategy action in the set of strategy actions based on the multimodal unified learning state vector, and select the strategy action with the highest action utility value to execute. When at least one of the preset cycle triggering condition, state offset triggering condition, and risk event triggering condition is met, the learning path optimization is triggered and the process proceeds to step S5. Otherwise, return to step S1 to proceed to the next round of the Socratic dialogue; Step S5: For the multimodal unified learning status vectors corresponding to multiple rounds of Socratic dialogues on the same knowledge point, perform global learning information aggregation to form a global learning information vector for that knowledge point; Based on a preset set of knowledge points, filter and sort the knowledge points in the set of knowledge points according to the global learning information vector to construct a set of candidate learning paths. Step S6: Construct a multi-objective comprehensive function; for each knowledge point of the candidate learning path in the candidate learning path set, calculate the mastery gap cost, emotional risk cost, fatigue load cost, and time cost based on the global learning situation vector, and accumulate them to form the comprehensive objective value of the path; under the condition of satisfying the preset learning time budget constraint and knowledge point coverage constraint, use a path update algorithm based on group iterative search to obtain the optimal solution of the multi-objective comprehensive function, and output the personalized learning path; Step S7: Based on the optimal solution of the multi-objective synthesis function of the current round, perform closed-loop feedback on the set of strategy parameters in step S4.

2. The online education path optimization method based on Socratic learning information quadruple perception as described in claim 1, characterized in that, The multi-objective synthesis function The expression is: ; in, This represents the loss term related to knowledge mastery; This indicates a loss of emotional stability. This represents the learning load loss term; This represents the total learning time loss item; This represents the loss item related to knowledge mastery. Weighting coefficients; Items representing loss of emotional stability Weighting coefficients; Represents the learning load loss term Weighting coefficients; This represents the total learning time loss item. The weighting coefficients.

3. The online education path optimization method based on Socratic learning information quadruple perception as described in claim 2, characterized in that, Encode several rounds of Socratic dialogues and construct a dialogue graph that includes time, questioning, and follow-up questions, including the following steps: Text encoding is performed using a text encoding layer: Suppose there are a total of 1 dialogue, number ; For the first Dialogue text Encode and construct text vectors. ;in, Represents a text encoding function; Indicates the first The initial vector of each dialogue node; Indicates the number of dialogue nodes; Construct a dialogue graph; the dialogue graph includes time sequence edges, question pointing edges, and answer pointing edges; The establishment of the time sequence edge includes: if the first The dialogue node immediately follows the first one in time. If there are 10 dialogue nodes, then: ; in, An adjacency matrix representing temporal order relationships; The establishment of the question-pointing edge includes: if the first... The first dialogue node is a question issued by the system, and the second... If each dialogue node corresponds to a student's answer, then: ; in, An adjacency matrix representing the relationship between system questions and student answers; The establishment of the answer-pointing edge includes: if the student's answer is followed up or clarified, then: ; in, The adjacency matrix represents the system's follow-up questions after a student answers. By fusing the time-sequence edges, question-pointing edges, and answer-pointing edges, we obtain the adjacency matrix of the dialogue graph. Its expression is: ;in, ; The expression for convolutional encoding of the dialogue graph is as follows: ; ; ; in, This represents the final embedding vector of dialogue node i in the dialogue graph structure; Represents a dialogue node The vector after the second layer of graph convolution; Indicates the activation function; The adjacency matrix of the dialogue graph represents the first... and Each element, if a dialog node With dialogue nodes If there is an edge in any of the relationships of time sequence, question, or follow-up question, the value is 1; otherwise, the value is 0. Represents a dialogue node The vector after the first layer of graph convolution; This represents the weight matrix of the second-layer graph convolution; This represents the weight matrix of the first layer graph convolution.

4. The online education path optimization method based on Socratic learning information quadruple perception as described in claim 3, characterized in that, Relationships are labeled on the dialogue graph, and a learning information four-tuple extraction model is constructed, including: For any pair of dialogue nodes Constructing relational feature vectors Its expression is: ; in, Represents a dialogue node The vector after convolution of the first layer graph; This represents a vector concatenation operation; The relationship label scoring vector is obtained by using a non-linear mapping: ; in, Indicates dialogue node pair Confidence scores belonging to different categories of learning information field relationships; Represents a nonlinear activation function; This represents the weight matrix of the relational mapping layer; Represents the bias vector of the relational mapping layer; According to the dialogue node pair Confidence scores belonging to different learning information field relationship categories The grid annotation matrix elements are obtained by annotation: ;in, Indicates dialogue node pair The relationship score on any type of learning performance field; The expression for filtering candidate ranges is: ; in, Field representing learning situation The set of candidate intervals; Indicates dialogue node pair In the learning situation field The relationship score; This indicates the target field for learning information. Preset confidence threshold; The optimal span interval is selected, and its expression is: ; in, Indicates in the learning situation field The optimal span range within; In the learning situation field The optimal span range within Mapping back to the original dialogue text: ;in, Field representing learning situation The actual content in the original dialogue text; Represents a sequence of dialogue text; From the learning situation field The actual content in the original dialogue text The model for extracting learning information using a four-tuple is constructed, and its expression is as follows: ; in, This represents the k-th learning information quadruple; This represents the field content of the learning object in the k-th learning information quadruple; This represents the field content of the learning dimension in the k-th learning information quadruple; The field representing the original words of the students in the k-th learning information quadruple; The field representing the polarity of the learning information in the k-th learning information quadruple.

5. The online education path optimization method based on Socratic learning information quadruple perception as described in claim 4, characterized in that, The emotional, cognitive, and behavioral features extracted by the learning situation four-tuple extraction model are fused to construct a multimodal unified learning situation state vector, including: The expression for extracting emotional features is as follows: ;in, Represents the emotional feature vector; This represents the emotion feature extraction function; Represents the learning situation in quadruplets; Extracting cognitive features, its expression is: ;in, Represents the feature vector of cognitive state; This represents a cognitive understanding and analysis function; Construct behavioral features, the expression of which is: ;in, Represents a behavioral feature vector; Indicates the time elapsed for the current step; Indicates the number of prompts; Indicates the number of incorrect steps; Constructing a unified learning state vector for multiple modalities Its expression is: ; in, This represents the feature fusion matrix.

6. The online education path optimization method based on Socratic learning information quadruple perception as described in claim 5, characterized in that, A set of preset strategy actions and a set of strategy parameters corresponding to the set of strategy actions are defined. The action utility value of each strategy action in the set of strategy actions is calculated based on the multimodal unified learning state vector. The strategy action with the highest action utility value is selected for execution. When at least one of the preset cycle triggering condition, state offset triggering condition, and risk event triggering condition is met, the learning path optimization is triggered and the process proceeds to step S5. Otherwise, return to step S1 to proceed to the next round of the Socratic dialogue, which includes: Preset strategy action set Its expression is: ; in, This indicates a continued inquiry and further questioning. Indicates a teaching action that provides a prompt; This action indicates a summary of the content that has been completed. This indicates a teaching action that directly explains the current step or the entire problem; This indicates a teaching action that switches topics or adjusts the difficulty of learning tasks; Based on a multimodal unified learning state vector Estimate the set of selected strategy actions The action value of any action in the sequence is expressed as: ; in, Represents the set of policy actions The first in Each strategic action, ; Representation and policy action set The first in Each strategy action The corresponding value weight vector; T represents the matrix transpose operation; Indicates the set of policy actions The first in Each strategy action The theoretical payoff score under the current state; For any set of strategy actions The first in Each strategy action The dynamic cost is obtained by the following expression: ; in, This represents the set of strategy actions under the current learning situation. The first in Each strategy action The dynamic cost; Represents the set of policy actions The first in Each strategy action The basic cost; Indicates the emotional sensitivity coefficient; This represents an index of negative emotion intensity calculated based on dialogue content and behavioral data. Indicates the cognitive sensitivity coefficient; Indicators representing positive cognitive levels; The action utility corresponding to any strategy action in the set of strategy actions is expressed as follows: ; in, This represents the set of strategy actions under the current learning situation. The first in Each strategy action The overall utility value; Optimal policy action is obtained by using policy selection. Its expression is: 。 7. The online education path optimization method based on Socratic learning information quadruple perception as described in claim 6, characterized in that, Global learning information aggregation is performed on the multimodal unified learning information state vectors corresponding to multiple rounds of Socratic dialogues for the same knowledge point, forming a global learning information vector for that knowledge point, the expression of which is: ; in, Indicates the first Global learning progress vector for each knowledge point; This indicates the number of Socratic dialogue rounds related to the knowledge point within the current analysis window, and it is an integer greater than or equal to 1. This represents the multimodal unified learning state vector corresponding to the Socratic dialogue in the kth round that is related to the knowledge point.

8. The online education path optimization method based on Socratic learning information quadruple perception as described in claim 2 or 7, characterized in that, The optimal solution of the multi-objective synthesis function in the current round is obtained using a population-based iterative search path update algorithm, including: Suppose that in the t-th iteration, there are a total of There are 10 candidate learning paths, denoted as: ;in, In the t-th iteration, the first... Candidate learning paths; Global guidance and random perturbation path update are employed: ; in, In the (t+1)th iteration, the th... Candidate learning paths; Denotes the multi-objective synthesis function in the t-th iteration. The shortest path; This represents the global guiding weight parameter; Indicates the random disturbance intensity coefficient; Represents a random perturbation vector with the same dimension as the path sequence; Based on the multi-objective synthesis function Find the optimal path: 。 9. The online education path optimization method based on Socratic learning information quadruple perception as described in claim 8, characterized in that, Based on the optimal solution of the multi-objective synthesis function in the current round, and by performing closed-loop feedback on the strategy parameter set in step S4, the closed-loop feedback includes: ; in, This represents the set of policy parameters in the t-th iteration, including the set of policy actions. The first in Each strategy action Corresponding value weight vector Emotional sensitivity coefficient Cognitive sensitivity coefficient Negative emotion intensity index Cost parameter vector ; This represents the set of strategy parameters in the (t+1)th iteration; Indicates the feedback update rate; This indicates the sensitivity of the multi-objective synthesis function to policy parameters.

10. A system employing the online education path optimization method based on the Socratic learning information quadruple perception as described in any one of claims 1 to 9, characterized in that, include: The dialogue graph construction module encodes several rounds of Socratic dialogues and constructs a dialogue graph that includes time, questioning, and follow-up questioning relationships. The learning information quadruple extraction model is connected to the dialogue graph construction module. Relationships are labeled on the dialogue graph to construct the learning information quadruple extraction model. The learning information quadruple includes learning object, learning dimension, student verbal evidence, and learning information polarity. The multimodal unified learning state vector construction module, together with the learning state quadruple extraction model, integrates the emotional features, cognitive features, and behavioral features extracted by the learning state quadruple extraction model to construct a multimodal unified learning state vector. The action utility calculation module is connected to the multimodal unified learning state vector construction module, the global learning vector solving module, and the dialogue graph construction module. It presets a set of strategy actions and a set of strategy parameters for strategy configuration, and calculates the action utility value of each strategy action in the set of strategy actions based on the multimodal unified learning state vector. The strategy action with the highest action utility value is selected for execution. When at least one of the preset periodic triggering condition, state offset triggering condition, and risk event triggering condition is met, the learning path optimization is triggered and the system enters the global learning vector solving module; otherwise, it returns to the dialogue graph construction module to enter the next round of Socratic dialogue. The global learning vector calculation module is connected to the action utility calculation module. It performs global learning aggregation on the multimodal unified learning state vector corresponding to multiple rounds of Socratic dialogue for the same knowledge point to form a global learning vector for that knowledge point. Based on a preset set of knowledge points, it filters and sorts the knowledge points in the set according to the global learning vector to construct a set of candidate learning paths. A multi-objective comprehensive function module is connected to a global learning vector solving module to construct a multi-objective comprehensive function. For each knowledge point of a candidate learning path in the candidate learning path set, the module calculates the mastery gap cost, emotional risk cost, fatigue load cost, and time cost based on the global learning vector, and accumulates them to form the comprehensive objective value of the path. Under the condition of satisfying the preset learning time budget constraint and knowledge point coverage constraint, the optimal solution of the multi-objective comprehensive function is obtained by using a path update algorithm based on group iterative search, and a personalized learning path is output. The closed-loop feedback module is connected to the multi-objective synthesis function module and the action utility calculation module. It provides closed-loop feedback to the strategy parameter set in the action utility calculation module based on the optimal solution of the multi-objective synthesis function in the current round.