Teaching information processing method and electronic equipment
By modeling teachers, students, and teaching content as game players, and using multimodal data to quantify the probability distribution and payoff of strategies, this method solves the problem of the inability to accurately judge the teaching status in existing technologies, and enables real-time perception and precise intervention of the teaching status.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU HAILIANG DIGITAL TECH CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies are insufficient to reflect the autonomous decision-making and mutual influence among teachers, students, and teaching content during the teaching process, resulting in an inability to accurately assess the teaching status and to promptly perceive and implement precise interventions.
Teachers, students, and teaching content are modeled as game players. By collecting multimodal data in real time, the strategy probability distribution and strategy payoff of each game player are determined. Evolutionary equilibrium and stability are analyzed using numerical methods to select the intervention strategy with the lowest cost and best effect.
It enables real-time perception and dynamic assessment of the teaching status, and can accurately screen and push targeted intervention strategies when identifying adverse situations, ensuring that the teaching process tends to be in a virtuous balance and avoiding inefficient or chaotic states.
Smart Images

Figure CN122045845A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a teaching information processing method and electronic device. Background Technology
[0002] With the development of artificial intelligence, multimodal perception, and big data technologies, timely classroom analysis, intelligent assessment of teaching status, and scientific intervention have become essential requirements for improving teaching quality.
[0003] Currently, there are various methods and means for assessing and intervening in teaching quality, such as classroom analysis based on post-class video playback and questionnaire feedback, or classroom analysis based on machine learning methods, such as facial expression recognition and speech analysis. However, methods based on post-class video playback and questionnaire feedback suffer from time lag, making real-time decision-making and dynamic adjustments impossible. While methods based on facial expression recognition and speech analysis possess a certain degree of automation, they often focus on isolated monitoring of individual behaviors or simple correlation analysis, neglecting the fact that classroom teaching is essentially a complex and dynamic interactive process involving teachers, students, and the teaching content. Therefore, they fail to reveal the interaction and dynamic balance among these three parties, leading to an inability to promptly perceive the teaching status and implement timely and precise interventions for teaching adjustments.
[0004] Therefore, existing methods struggle to capture the interrelationships of autonomous decision-making and mutual influence among teachers, students, and teaching content during the teaching process, making it impossible to accurately assess the teaching status. Thus, there is an urgent need for a method that can effectively characterize the dynamic equilibrium and game-theoretic characteristics of these three parties, enabling more effective real-time evaluation and precise intervention of teaching quality. Summary of the Invention
[0005] The purpose of this application is to address the shortcomings of the prior art by providing a teaching information processing method and electronic device, so as to solve the problem that the prior art is unable to reflect the relationship between the autonomous decision-making and mutual influence of teachers, students and teaching content in the teaching process, and therefore cannot accurately judge the teaching status.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, embodiments of this application provide a method for processing teaching information, the method comprising: Real-time collection of multimodal data at the current time step of the target teaching activity, including the teacher's teaching data, the student group's response data, and the teaching content; The teacher, the student group, and the teaching content are each modeled as game players, and the strategy probability distribution of each game player in the corresponding strategy space is determined based on the multimodal data, and the strategy payoff of each game player is determined based on the multimodal data. Each of the teacher, the student group, and the teaching content corresponds to a strategy space, and each strategy space includes at least one strategy dimension, and each strategy dimension includes multiple dimension types. Based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player, the state evaluation result of the target teaching activity is determined. Based on the status assessment results, target intervention strategies are selected from the preset intervention strategy library and sent to the user terminal held by the teacher.
[0007] As one possible implementation, determining the strategy probability distribution of each player in the corresponding strategy space based on the multimodal data includes: extracting behavioral features of each player from the multimodal data, and performing quantitative analysis on the behavioral features of each player to obtain the quantitative results of the behavioral features of each player; for each player, traversing each strategy dimension in the strategy space corresponding to the player, and for the current strategy dimension, matching the quantitative results of the player's behavioral features with the benchmark data corresponding to each dimension type under the current strategy dimension to obtain the matching degree between the player and each dimension type under the current strategy dimension; and determining the strategy probability distribution of each player in the corresponding strategy space based on the matching degree between each player and each dimension type under each strategy dimension in the corresponding strategy space.
[0008] As one possible implementation, determining the strategy probability distribution of each player in the corresponding strategy space based on the matching degree between each player and the dimension types under each strategy dimension in the corresponding strategy space includes: combining the dimension types under each strategy dimension in the strategy space corresponding to the player to obtain multiple dimension combinations, each dimension combination including one dimension type under multiple strategy dimensions; determining the probability result of each dimension combination based on the matching degree between each player and the dimension types under each strategy dimension in the corresponding strategy space; and using the vector composed of the probability results of all dimension combinations as the strategy probability distribution of the player in the corresponding strategy space.
[0009] As one possible implementation, determining the strategy payoffs of each player based on the multimodal data includes: determining the teaching progress deviation and the average test scores of students based on the teaching data; and using a formula... The strategy payoff for the teacher in the game is calculated; where, This represents the strategic payoff for the teacher in the game. This represents the average test score of the students. This indicates the deviation in the teaching progress. Indicates the first weight. This indicates the second weight.
[0010] As one possible implementation, determining the strategy payoffs of each player based on the multimodal data includes: determining the overall understanding and overall cognitive load based on the response data; and based on the formula... The strategic payoffs of the student group in the game are calculated; among them, This represents the strategic payoff of the student group in the game. This indicates the overall level of understanding. This indicates the overall cognitive load. Indicates the third weight. This indicates the fourth weight.
[0011] As one possible implementation, the response data includes: cognitive response data, text interaction data, behavioral trajectory data, and physiological visual data; determining the overall comprehension and overall cognitive load based on the response data includes: determining the question-and-answer accuracy and standard response time based on the cognitive response data, determining the cognitive state based on the text interaction data, determining the standard behavioral entropy based on the behavioral trajectory data, and determining the standard physiological information based on the physiological visual data; based on the formula The overall comprehension level is calculated; where, This indicates the overall level of understanding. This indicates the accuracy rate of the question and answer response. This indicates the cognitive state. Indicates the fifth weight. Represents the sixth weight; based on the formula The comprehensive cognitive load was calculated; wherein, This indicates the overall cognitive load. This indicates the standard response time. This represents the standard behavioral entropy. This represents the standard physiological information. Indicates the seventh weight. Indicates the eighth weight. This indicates the ninth weight.
[0012] As one possible implementation, determining the strategy payoffs of each player based on the multimodal data includes: determining the information decay rate of the players based on the teaching content; and based on the formula... The strategic payoffs of the players in the teaching content game are calculated; among them, This represents the strategic payoff of the player in the game based on the teaching content. This represents the information attenuation rate. This indicates the tenth weight.
[0013] As one possible implementation, determining the state evaluation result of the target teaching activity based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player includes: determining a first result based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player, and determining multiple second results based on the first result; predicting the strategy payoff of the next time step of the current time step based on the first result and each of the second results, iteratively executing until the difference between the predicted strategy payoff and the player's strategy payoff is less than a preset threshold, and taking the predicted strategy payoff at the end of the iteration as the target payoff; calculating the eigenvalue of the strategy evolution equation of each player at the target payoff; determining whether the strategy evolution equation is stable based on the eigenvalue; if so, determining the state evaluation result based on the average of the strategy payoffs corresponding to multiple time steps.
[0014] As one possible implementation, the step of selecting target intervention strategies from a preset intervention strategy library based on the state assessment result includes: if the state assessment result indicates that the target teaching activity has an unfavorable situation, then for each candidate intervention strategy in the preset intervention strategy library, the current teaching state characteristics are concatenated with the identification information of the candidate intervention strategy to obtain concatenated information, and the concatenated information is input into a pre-trained benefit prediction model, which then performs benefit prediction based on the concatenated information to obtain the target strategy benefit after executing the candidate intervention strategy; based on the target strategy benefit corresponding to each candidate intervention strategy, candidate intervention strategies that meet preset constraints are selected from multiple candidate intervention strategies as feasible intervention strategies, wherein the preset constraints are that the target strategy benefit after executing the candidate intervention strategy is greater than the strategy benefit before executing the candidate intervention strategy; and the feasible intervention strategy with the lowest intervention intensity is selected as the target intervention strategy.
[0015] Secondly, embodiments of this application provide a teaching information processing device, the device comprising: The data acquisition module is used to collect multimodal data of the target teaching activity in real time at the current time step. The multimodal data includes the teacher's teaching data, the student group's response data, and the teaching content. The processing module is used to model the teacher, the student group, and the teaching content as game players, determine the strategy probability distribution of each game player in the corresponding strategy space based on the multimodal data, and determine the strategy payoff of each game player based on the multimodal data. The teacher, the student group, and the teaching content each correspond to a strategy space, and each strategy space includes at least one strategy dimension, and each strategy dimension includes multiple dimension types. The determination module is used to determine the state evaluation result of the target teaching activity based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player. The intervention module is used to select target intervention strategies from a preset intervention strategy library based on the status assessment results, and send the target intervention strategies to the user terminal held by the teacher.
[0016] As one possible implementation, the processing module is specifically used for: extracting behavioral features of each player from the multimodal data, and performing quantitative analysis on the behavioral features of each player to obtain quantitative results of behavioral features of each player; for each player, traversing each strategy dimension in the strategy space corresponding to the player, and for the current strategy dimension, matching the quantitative results of the player's behavioral features with the benchmark data corresponding to each dimension type under the current strategy dimension to obtain the matching degree between the player and each dimension type under the current strategy dimension; and determining the strategy probability distribution of each player in the corresponding strategy space based on the matching degree between each player and each dimension type under each strategy dimension in the corresponding strategy space.
[0017] As one possible implementation, the processing module is specifically used to: combine the dimension types under each strategy dimension in the strategy space corresponding to the player to obtain multiple dimension combinations, each dimension combination including one dimension type under multiple strategy dimensions; determine the probability result of each dimension combination based on the matching degree between each player and each dimension type under each strategy dimension in the corresponding strategy space; and use the vector composed of the probability results of all dimension combinations as the strategy probability distribution of the player in the corresponding strategy space.
[0018] As one possible implementation, the processing module is specifically used to: determine the teaching progress deviation and the average test score of students based on the teaching data; and based on the formula... The strategy payoff for the teacher in the game is calculated; where, This represents the strategic payoff for the teacher in the game. This represents the average test score of the students. This indicates the deviation in the teaching progress. Indicates the first weight. This indicates the second weight.
[0019] As one possible implementation, the processing module is specifically used to: determine the overall comprehension level and overall cognitive load based on the response data; and based on the formula... The strategic payoffs of the student group in the game are calculated; among them, This represents the strategic payoff of the student group in the game. This indicates the overall level of understanding. This indicates the overall cognitive load. Indicates the third weight. This indicates the fourth weight.
[0020] As one possible implementation, the response data includes: cognitive response data, text interaction data, behavioral trajectory data, and physiological visual data; the processing module is specifically used to: determine the question-and-answer accuracy rate and standard response time based on the cognitive response data, determine the cognitive state based on the text interaction data, determine the standard behavioral entropy based on the behavioral trajectory data, and determine the standard physiological information based on the physiological visual data; based on the formula The overall comprehension level is calculated; where, This indicates the overall level of understanding. This indicates the accuracy rate of the question and answer response. This indicates the cognitive state. Indicates the fifth weight. Represents the sixth weight; based on the formula The comprehensive cognitive load was calculated; wherein, This indicates the overall cognitive load. This indicates the standard response time. This represents the standard behavioral entropy. This represents the standard physiological information. Indicates the seventh weight. Indicates the eighth weight. This indicates the ninth weight.
[0021] As one possible implementation, the processing module is specifically used to: determine the information decay rate of the game players based on the teaching content; and based on the formula... The strategic payoffs of the players in the teaching content game are calculated; among them, This represents the strategic payoff of the player in the game based on the teaching content. This represents the information attenuation rate. This indicates the tenth weight.
[0022] As one possible implementation, the determining module is specifically configured to: determine a first result based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player; and determine multiple second results based on the first result; predict the strategy payoff of the next time step of the current time step based on the first result and each of the second results, iteratively execute until the difference between the predicted strategy payoff and the player's strategy payoff is less than a preset threshold, and take the predicted strategy payoff at the end of the iteration as the target payoff; calculate the eigenvalue of the strategy evolution equation of each player at the target payoff; determine whether the strategy evolution equation is stable based on the eigenvalue; if so, determine the state evaluation result based on the average of the strategy payoffs corresponding to multiple time steps.
[0023] As one possible implementation, the intervention module is specifically used for: if the state assessment result indicates that the target teaching activity has a poor state, then for each candidate intervention strategy in the preset intervention strategy library, concatenating the current teaching state characteristics with the identification information of the candidate intervention strategy to obtain concatenated information, and inputting the concatenated information into a pre-trained benefit prediction model, which then performs benefit prediction based on the concatenated information to obtain the target strategy benefit after executing the candidate intervention strategy; based on the target strategy benefit corresponding to each candidate intervention strategy, selecting candidate intervention strategies that meet preset constraints from multiple candidate intervention strategies as feasible intervention strategies, wherein the preset constraints are that the target strategy benefit after executing the candidate intervention strategy is greater than the strategy benefit before executing the candidate intervention strategy; and selecting the feasible intervention strategy with the lowest intervention intensity as the target intervention strategy.
[0024] Thirdly, embodiments of this application provide an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the teaching information processing method as described in any of the first aspects above.
[0025] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the teaching information processing method as described in any of the first aspects above.
[0026] According to the teaching information processing method and electronic device of this application, multimodal data of the current time step of the target teaching activity is collected in real time. Teachers, students, and teaching content are modeled as game players, and the strategy probability distribution and strategy payoff of each player in the corresponding strategy space are determined based on the multimodal data. Based on the strategy probability distribution and strategy payoff of each player in the corresponding strategy space, the state evaluation result of the target teaching activity is determined. Based on the state evaluation result, target intervention strategies are selected from a preset intervention strategy library and sent to the user terminal held by the teacher. According to the embodiments of this application, by modeling teachers, students, and teaching content as game players and quantifying their strategy probability distribution and strategy payoff in real time using multimodal data, the complex teaching process is transformed into a dynamic evolutionary game system. This not only accurately depicts the real-time game characteristics of the three parties based on strategy interaction and payoff tradeoffs, but also analyzes evolutionary equilibrium and stability through numerical methods, thereby dynamically revealing whether the teaching state is trending towards a benign equilibrium, falling into inefficiency and instability, or in a state of chaos and instability. Based on this, this application can not only assess the teaching situation in a timely manner, but also accurately select the target intervention strategy with the lowest cost and the best expected improvement effect when identifying adverse situations, and push the target intervention strategy to the teacher's terminal, so as to realize the real-time perception, dynamic assessment and precise control of the teaching status. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart illustrating a teaching information processing method provided in an embodiment of this application is shown; Figure 2 A flowchart illustrating a method for determining a strategy probability distribution according to an embodiment of this application is shown. Figure 3 A flowchart illustrating another strategy probability distribution method provided in an embodiment of this application is shown; Figure 4 A flowchart illustrating a method for determining the state assessment results of a target teaching activity according to an embodiment of this application is shown. Figure 5 A flowchart illustrating a target intervention strategy screening method provided in an embodiment of this application is shown; Figure 6 This paper shows a schematic diagram of the structure of a teaching information processing device provided in an embodiment of this application; Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0030] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0031] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0032] Figure 1 A flowchart illustrating a teaching information processing method provided in an embodiment of this application is shown. (Refer to...) Figure 1 As shown, the method specifically includes the following steps: S101. Real-time acquisition of multimodal data of the current time step of the target teaching activities.
[0033] Multimodal data includes teachers' teaching data, students' response data, and teaching content.
[0034] Optionally, teaching data includes teaching progress, voice explanation data, and operational behavior data, such as PPT page-turning speed and blackboard writing frequency, which reflect the teacher's teaching rhythm and style. Response data includes cognitive response data (such as answer accuracy and response time), text interaction data (such as text content in discussion forums, bullet comments, and notes), behavioral trajectory data (such as mouse movement and interface dwell area), and physiological visual data (such as head posture change frequency, facial expression confusion level, and blinking frequency). Teaching content includes data on the presentation of teaching content, such as the knowledge points, formulas, and charts in the current courseware, student feedback data, and teacher output data.
[0035] S102. Model teachers, students, and teaching content as game players, determine the strategy probability distribution of each game player in the corresponding strategy space based on multimodal data, and determine the strategy payoff of each game player based on multimodal data.
[0036] Optionally, teachers, student groups, and teaching content each correspond to a strategy space, each strategy space includes at least one strategy dimension, and each strategy dimension includes multiple dimension types.
[0037] For example, a teacher's strategy space includes two strategy dimensions: teaching pace and teaching style. The teaching pace strategy dimension includes three types: fast, medium, and slow. The teaching style strategy dimension includes two types: abstract and concrete. A student's strategy space includes one strategy dimension: cognitive engagement. This cognitive engagement strategy dimension includes three types: high, medium, and low. A teaching content's strategy space includes two strategy dimensions: knowledge density and structural complexity. The knowledge density strategy dimension includes three types: high, medium, and low. The structural complexity strategy dimension also includes three types: high, medium, and low.
[0038] Optionally, all possible pure strategy combinations within the strategy space are generated for each player. For the single-dimensional strategy space corresponding to the student group players, the dimension type under each strategy dimension can directly constitute a dimension combination. For the multi-dimensional strategy space corresponding to the teacher players and the teaching content players, the complete combination set containing the dimension type under each strategy dimension is generated through the Cartesian product operation between dimensions, thereby enumerating all alternative strategy states.
[0039] Furthermore, for each generated dimension combination, based on the types of each dimension included in that combination, the matching degree between the player's behavioral characteristics and the corresponding dimension types, and according to preset fusion rules (such as calculating the matching degree product if dimensions are assumed to be independent, or performing weighted summation), the joint probability of that dimension combination is calculated. This joint probability represents the likelihood that the current actual teaching behavior conforms to that specific strategy pattern. Finally, the probability values calculated from all dimension combinations are arranged into a probability vector in a predefined order. The dimension of this probability vector is equal to the total number of pure strategies, thus forming a standardized mixed strategy probability distribution. This precisely quantifies the tendency of each player to adopt various strategies under the current teaching situation, providing a direct and structured input for subsequent evolutionary game analysis.
[0040] Optionally, for the teacher's strategic payoff, the process of determining their strategic payoff includes: dynamically collecting and integrating two types of core teaching data. One type reflects teaching progress, such as the difference between actual completed class hours and planned class hours, used to quantify teaching progress deviation. The other type reflects short-term teaching effectiveness, such as students' correct answer rate within a preset time window, used to calculate students' average test scores. By assigning different weights to these two core indicators and then performing a comprehensive calculation, the teacher's strategic payoff can be obtained. It is worth noting that the teacher's strategic payoff reflects the teacher's real-time trade-off between advancing teaching progress and ensuring student learning outcomes. A higher payoff value indicates that the teacher has achieved a better balance between progress and quality.
[0041] Optionally, for the student group as the game player, the process of determining their strategic payoff includes: calculating the comprehensive comprehension level, which reflects the degree of knowledge mastery, by analyzing cognitive response data and text interaction data; and calculating the comprehensive cognitive load, which reflects students' cognitive effort and psychological burden, by analyzing behavioral trajectory data and physiological visual data, combined with standardized response time. Based on this, the comprehensive comprehension level and comprehensive cognitive load are weighted and calculated to obtain the strategic payoff for the student group as the game player. It is worth noting that the strategic payoff for the student group as the game player comprehensively considers learning effectiveness and learning experience; higher payoffs indicate that students are in a highly efficient and low-cost active learning state.
[0042] Optionally, for the player in the teaching content game, the process of determining their strategic payoff includes: analyzing the teacher's output content, such as speech-to-text explanations, and the students' feedback content, such as notes and questions, respectively; calculating the information entropy of the teacher's output and the information entropy of the students' feedback to quantify their information richness and diversity; then, by comparing the ratio of the difference between the teacher's output information entropy and the students' feedback information entropy, calculating the information decay rate; and determining the player's strategic payoff based on this information decay rate. A lower information decay rate indicates a higher degree of effective reception and processing of the teaching content by students, and the player's strategic payoff is defined as a negative correlation function of this information decay rate, indicating that a higher payoff represents better information transmission efficiency and a more effective use of the teaching content.
[0043] S103. Based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player, determine the state evaluation results of the target teaching activity.
[0044] Optionally, in the embodiments of this application, based on the current strategy distribution and payoffs of each player, a numerical calculation method such as the fourth-order Runge-Kutta method is used to perform multi-round iterative prediction. The fourth-order Runge-Kutta method calculates an initial change and changes in multiple intermediate states to gradually deduce the dynamic changes of each player's strategy payoffs over time until the changes in the prediction results tend to stabilize, i.e., reach a convergence state, and the stable strategy payoff value at this time is determined as the target payoff.
[0045] Furthermore, based on the pre-established strategy evolution equations for teachers, students, and teaching content, the system eigenvalue at the target gain is calculated to determine whether the overall strategy evolution dynamics are in a stable equilibrium state. Specifically, if the eigenvalue indicates stability, the final state assessment result is determined based on the average level of strategy gains at multiple time steps during the iteration process. A higher average gain indicates a good teaching situation requiring no intervention, while a lower average gain indicates an inefficient or poor teaching state requiring intervention. Conversely, if the eigenvalue indicates instability, the state assessment result indicates that the teaching process is chaotic or unable to converge, necessitating immediate teaching intervention.
[0046] S104. Based on the status assessment results, select target intervention strategies from the preset intervention strategy library and send the target intervention strategies to the user terminals held by teachers.
[0047] Optionally, if the state assessment results indicate an unfavorable situation in the target teaching activity, i.e., when it is determined that the current teaching activity is in an unfavorable state, such as a sharp drop in student comprehension, excessive cognitive load, or an imbalance in teacher-student strategies, the intervention optimization process is immediately triggered. Specifically, based on the current real-time three-way game state, including the strategy distribution and payoffs of the teacher, students, and teaching content, and combined with each candidate intervention strategy in the pre-set intervention strategy library, a trained payoff prediction model is used to accurately predict the change in student payoffs after implementing each strategy. Then, all prediction results are screened, and only the intervention strategy that can ensure a clearly defined minimum increase in student payoffs is retained as a feasible intervention strategy, forming a feasible strategy set. Further, based on the principle of minimum intervention intensity, the feasible intervention strategy with the lowest cost and least interference intensity is selected from the feasible strategy set as the target intervention strategy.
[0048] Furthermore, after determining the target intervention strategy, specific and actionable prompts are generated and pushed to teachers' user terminals, such as computers, tablets, or smartwatches, via an interface. These prompts not only include a description of the specific intervention action, such as "suggest inserting a historical story about Faraday's law," but also clearly state the expected teaching effect of the intervention, such as "expected improvement in comprehension by 20%" and the required resources, such as "approximately 2 minutes," thereby assisting teachers in quickly and scientifically implementing interventions to reverse unfavorable teaching situations.
[0049] Based on this, the teaching information processing method according to the embodiments of this application models teachers, students, and teaching content as game players, and quantifies their strategy probability distribution and strategy payoffs in real time using multimodal data, thereby transforming the complex teaching process into a dynamic evolutionary game system. This not only accurately depicts the real-time game characteristics based on strategy interaction and payoff tradeoffs among the three parties, but also analyzes evolutionary equilibrium and stability through numerical methods, dynamically revealing whether the teaching state is trending towards a benign equilibrium, falling into inefficiency and poor performance, or in a state of chaos and instability. Based on this, this application can not only promptly assess the teaching situation, but also accurately select the lowest-cost and most effective target intervention strategy when identifying unfavorable situations, and push this target intervention strategy to the teacher's terminal, achieving real-time perception, dynamic assessment, and precise control of the teaching state.
[0050] Figure 2 A flowchart illustrating a method for determining a strategy probability distribution according to an embodiment of this application is shown. (Refer to...) Figure 2 As shown, step S102 above determines the strategy probability distribution of each player in the corresponding strategy space based on multimodal data, specifically including the following steps: S201. Extract the behavioral characteristics of each player from the multimodal data, and perform quantitative analysis on the behavioral characteristics of each player to obtain the quantitative results of the behavioral characteristics of each player.
[0051] Optionally, behavioral features specific to teachers, students, and teaching content can be extracted from the collected real-time multimodal teaching data. Specifically, for teachers, behavioral features include speech rate (words / minute), PPT page-turning interval, and frequency of words in abstract and concrete expressions. For students, behavioral features include average response time, keyword frequency in interactive text, mouse trajectory entropy value, and facial expression confusion score. For teaching content, behavioral features include vocabulary diversity in the explanatory text, frequency of knowledge points, and density of logical connectives.
[0052] Furthermore, the behavioral characteristics of each player extracted above are standardized and normalized to convert them into a unified numerical feature vector, forming a quantified result of the behavioral characteristics of each player that can be calculated and compared, thus laying a data foundation for subsequent strategy matching.
[0053] S202. For each player, traverse each strategy dimension in the strategy space corresponding to the player. For the current strategy dimension, match the quantitative results of the player's behavioral characteristics with the benchmark data corresponding to each dimension type under the current strategy dimension to obtain the matching degree between the player and each dimension type under the current strategy dimension.
[0054] Optionally, for each player, each strategy dimension defined in the strategy space corresponding to that player is traversed. Taking the teacher player as an example, the two strategy dimensions of "teaching pace" and "teaching style" are traversed. For the currently traversed strategy dimension, such as "teaching pace," the parts of the teacher player's behavioral characteristics quantified that are related to that strategy dimension, such as the teacher's actual speaking speed and page-turning frequency, are compared with the preset benchmark data range or feature templates for each dimension type (fast, medium, slow) under that strategy dimension. For example, the actual speaking speed is compared with the benchmark speaking speed range of "fast pace," and then the matching degree is calculated. This matching degree is used to reflect the degree of probability that the current actual behavioral characteristics conform to a certain specific dimension type.
[0055] S203. Based on the matching degree between each player and the corresponding strategy dimension type in the strategy space, determine the strategy probability distribution of each player in the corresponding strategy space.
[0056] Optionally, after obtaining the matching degree of each player with all dimension types under each strategy dimension, the final overall strategy probability distribution is determined through comprehensive calculation. Specifically, for a player, its strategy space consists of multiple dimensions. By comprehensively processing the matching degrees under each dimension, for example, through the multiplication principle or weighted fusion, the joint matching probability that simultaneously satisfies "fast pace" and "abstract style" is calculated, ultimately generating the probability distribution vector of the player over all possible pure strategy combinations. For example, each component in the teacher player's strategy probability distribution represents the probability that the current teaching behavior is judged to adopt a certain pure strategy, such as "fast pace, abstract style," and the sum of the probabilities of all components in the strategy probability distribution is 1. In this way, a complete mixed strategy description is formed, characterizing the strategy tendency of each player in the current teaching state.
[0057] Optionally, Figure 3 A flowchart illustrating another strategy probability distribution method provided in an embodiment of this application is shown. (Refer to...) Figure 3 As shown, step S203 above determines the strategy probability distribution of each player in the corresponding strategy space based on the matching degree between each player and the corresponding strategy dimension type in the strategy space. Specifically, it includes the following steps: S301. Combine the dimension types of each strategy dimension in the strategy space corresponding to the game players to obtain multiple dimension combinations.
[0058] Optionally, in the embodiments of this application, a Cartesian product operation is performed on the strategy dimensions defined by each player, that is, the multidimensional strategy space is enumerated into a series of discrete, evaluable alternative pure strategies, thereby generating all possible strategy combinations, and each dimension combination includes a dimension type under multiple strategy dimensions.
[0059] For example, taking the teacher as the game player, its strategy space includes two strategy dimensions: "teaching pace" (including three dimensions: fast, medium, and slow) and "teaching style" (including two dimensions: abstract and concrete). By combining all the dimension types under these two strategy dimensions in pairs, six complete dimension combinations can be generated: "fast pace - abstract style", "fast pace - concrete style", "medium pace - abstract style", "medium pace - concrete style", "slow pace - abstract style", and "slow pace - concrete style". Each dimension combination contains one concrete dimension type under each strategy dimension in the strategy space.
[0060] For example, taking the student group as a game player, their strategy space contains only one strategy dimension: cognitive engagement. This strategy dimension has three sub-dimensions: high, medium, and low. During combination processing, since there is only one strategy dimension, each sub-dimension can be directly considered an independent combination. Therefore, three dimension combinations are generated: "high engagement," "medium engagement," and "low engagement." Each dimension combination fully represents a specific pure strategy state of the student group on the cognitive engagement dimension.
[0061] For example, taking the game of teaching content as an example, the strategy space of the game player of teaching content includes two strategy dimensions: knowledge density (dimension type: high, medium, low) and structural complexity (dimension type: high, medium, low). When processing these two dimensions, each dimension type under the first dimension, i.e., knowledge density, is paired with each dimension type under the second dimension, i.e., structural complexity. Through pairwise combinations, nine different dimension combinations are finally generated, specifically "high density-high complexity", "high density-medium complexity", "high density-low complexity", "medium density-high complexity", "medium density-medium complexity", "medium density-low complexity", "low density-high complexity", "low density-medium complexity", and "low density-low complexity".
[0062] S302. Determine the probability results of each dimension combination based on the matching degree between each player and each dimension type under each strategy dimension in the corresponding strategy space.
[0063] Optionally, for each player, based on each dimension combination corresponding to each player, according to the types of dimensions included in the dimension combination, the matching degree corresponding to these dimension types calculated by the player in the aforementioned steps, and the preset fusion rules, such as assuming that the dimensions are independent of each other, the product of the matching degrees is used, or a weighted synthesis is performed according to the importance of the dimensions, to calculate the overall probability result of the dimension combination. This probability result represents the joint probability that the actual teaching behavior reflected by the current multimodal data matches the strategy pattern defined by the specific dimension combination.
[0064] S303. The vector formed by combining the probability results of all dimensions is taken as the strategy probability distribution of the player in the corresponding strategy space.
[0065] Optionally, based on the probability results of the combinations of each dimension, they are arranged in a predefined order to form a probability vector. The dimension of this probability vector is equal to the number of all possible pure strategies of the player. For example, the probability vector of the teacher player is 6-dimensional, and each element in the probability vector corresponds to the probability of a pure strategy. In addition, the probability vector satisfies the constraint that the sum of all elements is 1, thus completely and formally representing the mixed strategy adopted by the player in the current state, that is, the strategy probability distribution of the player in its own strategy space.
[0066] Based on this, the embodiments of this application combine the dimensional types under each strategy dimension to obtain multiple dimensional combinations, ensuring that all possible pure strategy states are covered. Furthermore, by calculating the joint probability based on the matching degree, a standardized strategy probability distribution vector is generated, realizing the mapping from low-dimensional behavioral indicators to high-dimensional strategy space. In this way, multi-dimensional and complex teaching behavior characteristics are efficiently and structurally mapped into a standard game-theoretic hybrid strategy expression, providing reliable input for subsequent dynamic analysis and equilibrium solving of three-party evolutionary games, and ensuring the feasibility of the entire teaching state evaluation and optimization.
[0067] As one possible implementation, this application models teachers, students, and teaching content as teacher players, student players, and teaching content players, respectively. For the teacher player, the above steps determine the player's strategy payoff based on multimodal data, including: determining the teaching progress deviation and average student test scores based on teaching data, and calculating the teacher player's strategy payoff based on the following formula (1): (1) in, This represents the strategic payoff for the teacher in the game. This represents the average test score of students. This indicates a deviation in the teaching progress. Indicates the first weight. This indicates the second weight.
[0068] For example, regression analysis of historical teaching data is used to determine , This indicates that teachers value scores more, but also pay attention to progress.
[0069] Optionally, the teaching data includes teaching progress data, such as the actual number of class hours completed, the planned number of class hours completed, and the total planned number of class hours. Based on the teaching progress data extracted from the teaching data, the teaching progress deviation can be calculated according to the following formula (2). : (2) Optionally, the teaching data also includes real-time classroom question and answer data, such as the number of questions answered correctly within a preset time window and the total number of questions within the preset time window. Based on the real-time classroom question and answer data extracted from the teaching data, the average test score of students can be calculated according to the following formula (3). : (3) in, This represents the average test score of students. Indicates the preset time window The number of questions answered correctly. Indicates the preset time window Total number of questions within, preset time window For example, 2 minutes.
[0070] For example, based on the teaching progress data and real-time classroom Q&A data extracted from the teaching data, the teaching progress deviation is calculated accordingly using the above formulas (2) and (3). and average test scores of students and the deviation in teaching progress and average test scores of students Substituting into formula (1) above, the strategic payoff of the teacher in the game can be calculated. .
[0071] Based on this, the embodiments of this application integrate teaching progress data (actual and planned lesson completion) and real-time classroom question-and-answer data (answer accuracy within a preset time window), quantifying them into teaching progress deviation and average student test scores, respectively. These are then combined with weighted coefficients to calculate strategy returns. This not only dynamically reflects the teacher's trade-off between teaching efficiency and quality but also avoids subjective evaluation bias. For example, when a teacher excessively pursues progress while neglecting mastery, the teaching progress deviation may be small, but the average student test score may be low, leading to a decrease in strategy returns. Conversely, if teaching is solid but significantly behind schedule, the strategy returns will also decrease. Therefore, the quantification mechanism based on multimodal real-world teaching behavior data provided in this application offers a reliable basis for optimizing teaching strategies, effectively supporting the adjustment of teacher strategies and the improvement of overall teaching effectiveness in subsequent evolutionary game analysis.
[0072] As a possible implementation, for the student group as the game player, the above steps determine the game player's strategy payoff based on multimodal data, including: determining the comprehensive comprehension and comprehensive cognitive load based on the response data, and calculating the strategy payoff of the student group as the game player based on the following formula (4): (4) in, This represents the strategic payoff of the student group in the game. Indicates overall comprehension level. Indicates overall cognitive load. Indicates the third weight. This indicates the fourth weight.
[0073] For example, third weight For example, a weight of 1.0, the fourth weight. For example, it is 0.9.
[0074] Optionally, response data includes: cognitive response data, text interaction data, behavioral trajectory data, and physiological visual data. Cognitive response data includes, for example, students' test scores, the current average response time of the student group, historical minimum response time, and historical maximum response time. Text interaction data includes, for example, interaction data entered by students in text interaction areas such as chat areas, bullet comments, and notes. Positive keywords such as "understand," "comprehend," and "moved" can be further extracted from the text interaction data, as well as negative keywords such as "why," "don't understand," and "confused." Behavioral trajectory data includes, for example, the distribution of mouse movement trajectories or click areas (for example, if a student is listening attentively, the mouse will remain in a relatively fixed area; if a student is confused, the mouse will wander randomly). Physiological visual data includes, for example, the frequency of students' head posture changes and facial expression confusion scores.
[0075] Based on this, the above steps determine the overall comprehension and overall cognitive load according to the response data, including: determining the question-and-answer accuracy and standard response time based on cognitive response data, determining the cognitive state based on text interaction data, determining the standard behavioral entropy based on behavioral trajectory data, and determining the standard physiological information based on physiological visual data. Then, based on the question-and-answer accuracy, standard response time, cognitive state, standard behavioral entropy, and standard physiological information, the overall comprehension and overall cognitive load are determined.
[0076] Specifically, the question-and-answer accuracy rate is calculated using the formula (3) above, which shows the average test score of students. The calculation method is similar. For example, for a single test, the average test score can be calculated by dividing the test score of each student in that test by the total number of students. For multiple tests, the average test score can be determined by dividing the average test score of students across multiple tests by the total number of tests. The standard response time is calculated as shown in the following formula (5): (5) in, Indicates the standard response time. This indicates the current average response time for the student population. Indicates the historical minimum response time. This indicates the historical maximum response time.
[0077] Specifically, the text interaction data is analyzed, and positive keywords such as "understand," "comprehend," and "moved" are extracted from the text interaction data, as well as negative keywords such as "why," "don't understand," and "confused." The frequency of occurrence of positive and negative keywords is counted, and then the cognitive state is calculated based on the following formula (6): (6) in, Indicates cognitive state. This indicates the frequency of positive keywords. This indicates the number of times negative keywords appear.
[0078] Specifically, standard behavioral entropy can be calculated based on the information entropy of behavioral trajectory data such as the distribution of mouse movement trajectories or click areas. Information entropy can be used to reflect attention concentration. The calculation method of information entropy is shown in the following formula (7): (7) in, Represents information entropy. This indicates that the mouse is over the interface area. The probability of.
[0079] Furthermore, after obtaining the aforementioned information entropy Based on this, the standard behavioral entropy can be calculated using the following formula (8): (8) in, Represents the standard behavioral entropy, Represents information entropy. Represents the minimum information entropy. This represents the maximum information entropy.
[0080] Specifically, raw physiological information is determined by collecting raw physiological visual data such as the frequency of students' head posture changes and facial expression confusion scores. This raw physiological visual data can be analyzed based on computer vision algorithms to capture classroom footage from cameras, identify students' heads, and calculate the number of times significant angle changes occur within a preset time period, such as 30 seconds or 1 minute, and convert it into a frequency per minute to obtain raw physiological information. For example, if a student's head has moved 15 times in the past minute, the standard physiological information is calculated using the following formula (9): (9) in, Represents standard physiological information, This indicates raw physiological information, such as a head twitching frequency of 15 times per minute. This represents the theoretical lower limit when cognitive load is very low and attention is highly focused on the learning content, such as a head twitching frequency of 2 times / minute. This represents the theoretical upper limit of cognitive load, feeling extremely confused or anxious, such as a head tetany frequency of 40 times per minute.
[0081] Furthermore, based on the question-and-answer accuracy, standard response time, cognitive state, standard behavioral entropy, and standard physiological information obtained from the response data, the overall comprehension can be calculated using the following formula (10), and the overall cognitive load can be calculated using the following formula (11): (10) in, Indicates overall comprehension level. Indicates the accuracy rate of the question and answer. Indicates cognitive state. Indicates the fifth weight. This indicates the sixth weight.
[0082] For example, the fifth weight For example, 0.7, the sixth weight For example, it is 0.3.
[0083] (11) in, Indicates overall cognitive load. Indicates the standard response time. Represents the standard behavioral entropy, Represents standard physiological information, Indicates the seventh weight. Indicates the eighth weight. This indicates the ninth weight.
[0084] For example, the seventh weight Eighth weight and the ninth weight It can be set according to the actual situation, but it must meet the seventh weight. Eighth weight and the ninth weight The sum is 1.
[0085] For example, the comprehensive cognitive load is calculated based on the above formulas (10) and (11). and overall comprehension Then, the overall cognitive load will be considered. and overall comprehension Substituting into formula (4) above, the strategic payoff of the student group in the game can be calculated. .
[0086] Based on this, the embodiments of this application comprehensively utilize cognitive response data, text interaction data, behavioral trajectory data, and physiological visual data. Through standardization and weighted fusion, the overall comprehension level and overall cognitive load are determined respectively, and then the strategy gains are calculated based on the overall comprehension level and overall cognitive load. This not only avoids the one-sidedness of relying on a single indicator, such as only looking at scores, but also dynamically captures subtle changes in students' knowledge mastery, cognitive input, and psychological state. For example, even with a high accuracy rate, if behavioral entropy is high and expressions of confusion are frequent, the gains will still be reduced due to high cognitive load, thus truly reflecting learning efficiency. Therefore, this ensures that the obtained strategy gains can more comprehensively and accurately depict the learning experience and effectiveness of the student group, providing reliable data support for real-time adjustment of teaching strategies and evolutionary game analysis.
[0087] As a possible implementation, for the game player regarding teaching content, the above steps determine the game player's strategy payoff based on multimodal data, including: determining the information decay rate of the game player based on the teaching content, and calculating the game player's strategy payoff based on the following formula (12): (12) in, This represents the strategic payoff of the players in the game of teaching content. Indicates the information decay rate. This indicates the tenth weight.
[0088] For example, the teaching content includes teaching progress, voice explanation data, and operational behavior data, such as PPT page-turning speed and blackboard writing frequency, which can reflect the teacher's teaching rhythm and style, as well as corresponding student feedback data. In this embodiment, the teacher's voice explanation data can be converted to obtain corresponding text, and preprocessing operations such as word segmentation and stop word removal can be performed on the text to calculate the word frequency distribution information entropy corresponding to the teacher's output. Similarly, for student feedback data, such as notes, questions, and bullet comments, preprocessing operations such as word segmentation and stop word removal can also be performed to calculate the word frequency distribution information entropy corresponding to the student feedback.
[0089] Specifically, taking the calculation of the word frequency distribution information entropy corresponding to the teacher's output as an example, the word frequency distribution information entropy can be calculated using the following formula (13): (13) in, This represents the entropy of the word frequency distribution information corresponding to the teacher's output. Words The probability of it occurring within a time window.
[0090] Furthermore, based on the information entropy of the word frequency distribution corresponding to the teacher's output and the information entropy of the word frequency distribution corresponding to the student's feedback, the above information decay rate can be calculated using the following formula (14). : (14) in, Indicates the information decay rate. This represents the entropy of the word frequency distribution information corresponding to the teacher's output. This represents the information entropy of the word frequency distribution corresponding to student feedback.
[0091] For example, the information decay rate is calculated based on the above formula (14). Based on this, information decay rate Substituting into the above formula (12), the strategic payoff of the player in the teaching content game can be calculated. .
[0092] Based on this, this application's embodiments analyze the converted text of the teacher's speech and the student's feedback text, calculate the information entropy of their word frequency distributions, and determine the information decay rate based on the difference in information entropy between the two word frequency distributions. Then, the strategy benefit is calculated using the information decay rate. Since the information entropy of the word frequency distribution measures the richness and uncertainty of semantic content, when the semantic diversity of student feedback is significantly lower than that of the teacher's output, it indicates that a large amount of teaching information is lost or not absorbed during transmission, i.e., the information decay rate is high and the strategy benefit is reduced. This quantitative mechanism based on multimodal text data can objectively and in real-time assess the effective transmission of teaching content, providing reliable data for optimizing teaching pace, adjusting knowledge density, or improving expression methods, thereby improving the overall collaborative efficiency and knowledge transformation effectiveness of teaching.
[0093] Figure 4 This illustration shows a flowchart of a method for determining the state assessment results of a target teaching activity according to an embodiment of this application. (Refer to...) Figure 4 As shown, step S103 above determines the state evaluation result of the target teaching activity based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player. Specifically, it includes the following steps: S401. Based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player, determine the first outcome, and based on the first outcome, determine multiple second outcomes.
[0094] Optionally, for each player, the fourth-order Runge-Kutta method (RK4) can be used, combined with the strategy probability distribution of the players in the corresponding strategy space and the strategy payoff of each player, to perform iterative solution.
[0095] For example, taking the student group as an example, the first outcome can be determined based on the strategy probability distribution of the student group in the corresponding strategy space and the strategy payoff of each player. The specific expression is shown in the following expression (15): (15) in, Indicates the first result. Indicates the time step. This represents the strategic payoff of the student group at the current time step, specifically manifested as... The first player in the game among the student groups The proportions of the dimension types under each strategy dimension This indicates the current distribution of teacher strategies. Distribution of teaching content strategies The following strategy is adopted. The expected payoff of the student player in the game. express The average expected payoff for each student group in the game.
[0096] Furthermore, after obtaining the aforementioned first result... Based on the first result Calculate multiple second results, such as those described below. , , The specific calculation process is shown in the following formulas (16), (17), and (18): (16) (17) (18) in, , , Indicates the second result. Indicates the time step. , , Indicating different intermediate states The first player in the game among the student groups The proportions of the dimension types under each strategy dimension This indicates the current distribution of teacher strategies. Distribution of teaching content strategies The following strategy is adopted. The expected payoff of the student player in the game. , , This represents the new average return corresponding to different intermediate states.
[0097] Specifically, the above , , .
[0098] S402. Based on the first result and each of the second results, predict the strategy payoff for the next time step of the current time step, iterate until the difference between the predicted strategy payoff and the player's strategy payoff is less than a preset threshold, and take the predicted strategy payoff at the end of the iteration as the target payoff.
[0099] Optionally, continuing to take the student group as the game player as an example, after obtaining the above first result... and each second result , , Based on this, the policy payoff for the next time step can be predicted using the following expression (19): (19) in, This indicates the strategy payoff at the next time step. This represents the strategy payoff of the student group at the current time step. Indicates the first result. , , This indicates the second result.
[0100] Furthermore, for each player, based on the above method, the player's strategy payoff at the current time step and the strategy payoff at the next time step are calculated. The change norm, i.e., the difference between the predicted strategy payoff and the player's strategy payoff, is then calculated. If this difference is less than a preset threshold, such as... If the iteration stops, the predicted policy return at the end of the iteration is taken as the target return.
[0101] S403. Calculate the eigenvalues of the strategy evolution equations of each player at the target payoff.
[0102] Optionally, the strategy evolution equation for the student group players is shown in the following expression (20): (20) in, The first player in the student group game represents the third player. The proportions of the dimension types under each strategy dimension This indicates the current distribution of teacher strategies. Distribution of teaching content strategies The following strategy is adopted. The expected payoff of the student player in the game. This represents the average expected payoff for the student group in the game.
[0103] Optionally, the strategy evolution equation for the teacher player is shown in the following expression (21): (twenty one) in, The teacher's game partner is represented by the first... The proportions of the dimension types under each strategy dimension This indicates the strategy distribution among the current student group. Distribution of teaching content strategies The following strategy is adopted. The expected payoff of the teacher in the game. This represents the average expected payoff for the teacher player in the game.
[0104] Optionally, the strategy evolution equation of the game players in the teaching content is specifically shown in the following expression (22): (twenty two) in, The first player in the game represents the teaching content. The proportions of the dimension types under each strategy dimension This indicates the current distribution of teacher strategies. Student Group Strategy Distribution The following strategy is adopted. The expected payoff of the game players in terms of the teaching content. This represents the average expected payoff for each player in the game of teaching content.
[0105] Furthermore, based on the strategy evolution equations of each player, the eigenvalues of the strategy evolution equations of each player at the target payoff are calculated, that is, the first derivative of each strategy evolution equation is obtained, and the result of the first derivative is used as the eigenvalue.
[0106] S404. Determine whether the strategy evolution equation is stable based on the eigenvalues.
[0107] Optionally, if the eigenvalue of the player's strategy evolution equation at the target payoff is less than zero, then the player's strategy evolution equation is determined to be stable; conversely, if the eigenvalue of the player's strategy evolution equation at the target payoff is greater than zero, then the player's strategy evolution equation is determined to be unstable.
[0108] S405. If so, the state evaluation result is determined based on the average of the strategy returns corresponding to multiple time steps.
[0109] Optionally, under stable conditions, if the average of the policy returns across multiple time steps is less than a preset value, such as 0.5, the state evaluation result is determined to be a poor state, indicating that teaching is inefficient, the teacher is overworked, and the students cannot understand, requiring intervention. If the average of the policy returns across multiple time steps is greater than the preset value, such as 0.5, the state evaluation result is determined to be a good state, indicating that teaching is effective, the students understand, and no intervention is needed. Furthermore, if unstable, the state evaluation result indicates non-convergence / chaos, suggesting that teaching is in a chaotic and unpredictable state, requiring intervention.
[0110] Based on this, the embodiments of this application employ the fourth-order Runge-Kutta method for iterative solution, effectively avoiding the numerical instability and slow convergence problems that may occur in traditional evolutionary game analysis. This allows for a more accurate and efficient capture of the evolutionary trends of the strategic interactions among teachers, students, and teaching content, providing a reliable basis for teaching evaluation. Furthermore, this application also clearly distinguishes between steady-state outcomes and unsteady chaotic states in the teaching process through eigenvalue stability judgment, thereby achieving scientific evaluation and precise intervention of the effectiveness of teaching activities and significantly improving the simulation accuracy and stability of the dynamic evolution process of teaching activities.
[0111] Figure 5 A flowchart illustrating a target intervention strategy screening method provided in an embodiment of this application is shown. (Refer to...) Figure 5 As shown, the above steps, based on the state assessment results, select target intervention strategies from the pre-set intervention strategy library, specifically including the following steps: S501. If the state assessment result indicates that the target teaching activity has a poor state, then for each candidate intervention strategy in the preset intervention strategy library, the current teaching state characteristics and the identification information of the candidate intervention strategy are spliced together to obtain spliced information. The spliced information is then input into the pre-trained benefit prediction model, and the benefit prediction model makes benefit prediction based on the spliced information to obtain the benefit of the target strategy after executing the candidate intervention strategy.
[0112] Optionally, when a negative situation is identified in the current teaching activity, the dynamic characteristics of the current teaching state are obtained, including the strategy distribution ratios of teachers, students, and teaching content, and their corresponding real-time benefit values, which together constitute a multi-dimensional state feature vector. Then, for each candidate intervention strategy in the preset intervention strategy library, such as "prompting the teacher to slow down their speech" or "inserting a specific case," its unique identification information is concatenated with the above state feature vector to obtain concatenated information. This concatenated information is a complete input information that integrates the current teaching context and the proposed intervention action.
[0113] Furthermore, the spliced information is input into a pre-trained benefit prediction model, which, like a lightweight neural network, learns a complex mapping relationship between intervention actions and benefit changes based on historical teaching data. By processing the input spliced information, the benefit prediction model predicts the new strategy benefit that the student group may achieve after implementing the candidate intervention strategy, i.e., the target strategy benefit.
[0114] S502. Based on the target strategy benefit corresponding to each candidate intervention strategy, select candidate intervention strategies that meet the preset constraints from multiple candidate intervention strategies as feasible intervention strategies.
[0115] The preset constraint is that the target strategy return after implementing the candidate intervention strategy must be greater than the strategy return before implementing the candidate intervention strategy. It is worth noting that this preset constraint requires that the target strategy return after implementing the candidate intervention strategy must be strictly greater than the baseline return before the intervention, and it usually also needs to meet a minimum improvement threshold, such as an increase in return of at least 0.2, to ensure that the intervention has a real improvement effect.
[0116] Optionally, after obtaining the target strategy return corresponding to each candidate intervention strategy, it is compared with the strategy return before the intervention, and filtered according to preset constraints. Specifically, all candidate intervention strategies are traversed, and their predicted returns are checked one by one to see if they meet the above preset constraints. All candidate intervention strategies that meet the above preset constraints are marked as feasible intervention strategies, forming a set of currently available effective interventions, laying the foundation for the selection of the final strategy.
[0117] S503. The feasible intervention strategy with the lowest intervention intensity shall be the target intervention strategy.
[0118] Optionally, from the selected set of feasible intervention strategies, the final target intervention strategy is further determined based on the principle of minimizing intervention cost or intensity. Specifically, each feasible intervention strategy is assigned an intensity weight, such as a simple cue with an intensity of 1, a complex multimedia interaction with an intensity of 3, and the intensity of a combined intervention is the sum of the intensities of its constituent actions. Based on this, the intensity values of all feasible intervention strategies are calculated, and the strategy with the lowest intensity is selected as the target intervention strategy.
[0119] Optionally, if there are multiple feasible intervention strategies with the same intensity, the one with the highest predicted benefit can be selected first, thereby optimizing the intervention cost while ensuring the constraint of improving teaching benefits, and ensuring the accuracy and economy of teaching intervention.
[0120] Based on this, the embodiments of this application jointly encode the current teaching state and candidate intervention actions and input them into a pre-trained lightweight neural network for benefit prediction, thereby realizing the prediction of intervention effects. Furthermore, by combining constraint screening and intensity ranking mechanisms, it ensures that the selected target intervention strategy meets educational goals while taking into account implementation efficiency and resource conservation.
[0121] The following will provide a detailed explanation of the instructional information processing method provided in this application using a complete example: Target teaching activity: High school physics "Electromagnetic Induction", current teaching segment: explaining Lenz's Law, data collection and monitoring cycle. arrive It lasts for 5 minutes; Regarding the initial state and data collection for a three-way game: assuming that in time... The monitored three-party strategy states are as follows: the strategy probability distribution of the teacher's side is (0.1, 0.1, 0.6, 0.1, 0.1, 0.0), that is, the third strategy (medium, abstract) has the highest probability; the strategy probability distribution of the student group's side is (0.15, 0.20, 0.65), that is, the proportion of low-investment students is 65%; the strategy probability distribution of the teaching content's side is (0.8, 0.05, ..., 0.05), that is, the first strategy has the highest probability.
[0122] The collected raw data are as follows: Regarding the teaching progress deviation, if 2 knowledge points were planned to be explained but 1.5 were actually completed, the teaching progress deviation is 0.05. Regarding the test accuracy rate, if 2 questions were pushed within 5 minutes, the accuracy rate is 0.5. Regarding the discussion forum text, 20 bullet comments were analyzed, including 12 confusing words and 8 comprehension words. Regarding student response time, the current average response time is 4.5 seconds, the historical minimum response time is 1 second, and the historical maximum response time is 10 seconds. Regarding behavioral entropy, the current behavioral entropy is calculated to be 1.8, the historical minimum behavioral entropy is 0.5, and the historical maximum behavioral entropy is 2.5. Regarding information entropy, H (teacher output) is calculated from the teacher's speech to text as 4.2 bits, and H (student feedback) is calculated from student notes and questions as 1.5 bits.
[0123] Furthermore, real-time payout calculations show that the strategic payout for the teacher is 0.46, for the student group is 0.025, and for the teaching content is -0.643. Based on this, the state payout for the current time step is (0.46, 0.025, -0.643). Among these, the strategic payout for the student group is extremely low, and the strategic payout for the teaching content is negative, indicating that the teaching effect is very poor and equilibrium calculation and intervention are required.
[0124] For example, taking the student group as a game player, we replicate dynamics and Nash equilibrium calculations. Specifically, based on the current state, we calculate the strategy evolution of the student players. Assuming that based on the current teacher strategy and teaching content strategy, the expected payoffs for the three student strategies are calculated as follows: high input 0.40, medium input 0.15, and low input -0.10, then the corresponding average strategy payoff is 0.025. Based on this, we perform RK4 iterations, assuming a time step of 0.5, to calculate the change in strategy 1 (high input), obtaining the first result. The value is 0.028125. For simplicity, we assume the average return estimated from the intermediate state is approximately 0.05, and thus obtain the second result. The value is 0.0287. For simplicity, we assume the average return estimated from the intermediate state is approximately 0.06, and then calculate... The value is 0.0287, which is further calculated to obtain... The value is 0.0304. Based on this, the predicted difference between the policy reward at the current time step and the policy reward at the next time step is 0.0289. Convergence is then determined. Assuming that after several steps, the policy ratio no longer changes significantly and converges to (0.25, 0.25, 0.50), the calculated average reward is 0.10. This indicates stability, and according to the evaluation rules, the state evaluation result is a negative situation. Therefore, intervention optimization is triggered to find the target intervention policy with the minimum intensity.
[0125] For example, suppose the following candidate intervention strategies exist in the preset intervention strategy library: prompting teachers to slow down their speech (intensity=1), inserting a specific case (intensity=2), highlighting formulas and pushing animations (intensity=3), and a combined intervention strategy of prompting teachers to slow down their speech and inserting a specific case (intensity=3). Based on this, a benefit prediction model was used to predict the benefits, resulting in predicted benefits of 0.22, 0.35, 0.40, and 0.45 for each of the candidate intervention strategies. Taking a threshold of 0.30 as an example for comparison, the feasible intervention strategy that satisfies the preset constraints among all candidate intervention strategies is determined to be a combination of the following intervention strategies: inserting a specific case, highlighting the formula and pushing an animation, prompting the teacher to slow down their speech, and inserting a specific case. By comparing the strength of these feasible intervention strategies, the feasible intervention strategy with the lowest strength, "inserting a specific case," is selected as the target intervention strategy, and suggestions are sent to the teacher's user terminal, such as "Students have difficulty understanding abstract laws and have high cognitive load. Suggestion: Insert a demonstration video of the experiment of 'magnets passing through and out of copper tubes,' which takes 1.5 minutes and is expected to improve comprehension to 0.35+."
[0126] Based on the same inventive concept, this application also provides a teaching information processing device corresponding to the teaching information processing method. Since the principle of the teaching information processing device in this application is similar to that of the teaching information processing method described above in this application, the implementation of the teaching information processing device can refer to the implementation of the method, and the repeated parts will not be described again.
[0127] Reference Figure 6 The diagram shown is a structural schematic of a teaching information processing device provided in an embodiment of this application. The teaching information processing device 600 includes: a data acquisition module 601, a processing module 602, a determination module 603, and an intervention module 604, wherein: The data acquisition module 601 is used to collect multimodal data of the target teaching activity in real time at the current time step. The multimodal data includes the teacher's teaching data, the student group's response data, and the teaching content. The processing module 602 is used to model teachers, students and teaching content as players, determine the strategy probability distribution of each player in the corresponding strategy space based on multimodal data, and determine the strategy payoff of each player based on multimodal data. Each player, student and teaching content corresponds to a strategy space, each strategy space includes at least one strategy dimension, and each strategy dimension includes multiple dimension types. The determination module 603 is used to determine the state evaluation result of the target teaching activity based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player. The intervention module 604 is used to select target intervention strategies from the preset intervention strategy library based on the status assessment results and send the target intervention strategies to the user terminal held by the teacher.
[0128] Based on this, the teaching information processing device according to the embodiments of this application models the teacher, student group, and teaching content as game players, and quantifies their strategy probability distribution and strategy payoff in real time using multimodal data, thereby transforming the complex teaching process into a dynamic evolutionary game system. In this way, it can not only accurately characterize the real-time game characteristics based on strategy interaction and payoff tradeoffs among the three parties, but also analyze evolutionary equilibrium and stability through numerical methods, thereby dynamically revealing whether the teaching state is trending towards a benign equilibrium, falling into inefficiency and poor performance, or in a state of chaos and instability. Based on this, this application can not only timely assess the teaching situation, but also, when identifying unfavorable situations, accurately select the target intervention strategy with the lowest cost and the best expected improvement effect, and push the target intervention strategy to the teacher's terminal, realizing real-time perception, dynamic assessment, and precise control of the teaching state.
[0129] In one possible implementation, the processing module 602 is specifically used for: extracting behavioral features of each player from multimodal data, and performing quantitative analysis on the behavioral features of each player to obtain the quantitative results of the behavioral features of each player; for each player, traversing each strategy dimension in the strategy space corresponding to the player, and for the current strategy dimension, matching the quantitative results of the player's behavioral features with the benchmark data corresponding to each dimension type under the current strategy dimension to obtain the matching degree between the player and each dimension type under the current strategy dimension; and determining the strategy probability distribution of each player in the corresponding strategy space based on the matching degree between each player and each dimension type under each strategy dimension in the corresponding strategy space.
[0130] As one possible implementation, the processing module 602 is specifically used to: combine the dimension types under each strategy dimension in the strategy space corresponding to the player to obtain multiple dimension combinations, each dimension combination including one dimension type under multiple strategy dimensions; determine the probability result of each dimension combination based on the matching degree between each player and each dimension type under each strategy dimension in the corresponding strategy space; and use the vector composed of the probability results of all dimension combinations as the strategy probability distribution of the player in the corresponding strategy space.
[0131] As one possible implementation, the aforementioned processing module 602 is specifically used to: determine the teaching progress deviation and the average test score of students based on teaching data; and based on the formula... The strategy payoff for the teacher in the game is calculated; where, This represents the strategic payoff for the teacher in the game. This represents the average test score of students. This indicates a deviation in the teaching progress. Indicates the first weight. This indicates the second weight.
[0132] As one possible implementation, the aforementioned processing module 602 is specifically used to: determine the overall comprehension level and overall cognitive load based on the response data; and based on the formula... The strategic payoffs of the student group in the game are calculated; among them, This represents the strategic payoff of the student group in the game. Indicates overall comprehension level. Indicates overall cognitive load. Indicates the third weight. This indicates the fourth weight.
[0133] As one possible implementation, the response data includes: cognitive response data, text interaction data, behavioral trajectory data, and physiological visual data; the aforementioned processing module 602 is specifically used to: determine the question-and-answer accuracy rate and standard response time based on the cognitive response data, determine the cognitive state based on the text interaction data, determine the standard behavioral entropy based on the behavioral trajectory data, and determine the standard physiological information based on the physiological visual data; based on the formula The overall comprehension level is calculated; among which, Indicates overall comprehension level. Indicates the accuracy rate of the question and answer. Indicates cognitive state. Indicates the fifth weight. Represents the sixth weight; based on the formula The overall cognitive load was calculated; among which, Indicates overall cognitive load. Indicates the standard response time. Represents the standard behavioral entropy, Represents standard physiological information, Indicates the seventh weight. Indicates the eighth weight. This indicates the ninth weight.
[0134] As one possible implementation, the aforementioned processing module 602 is specifically used for: determining the information decay rate of the game players based on the teaching content; and based on the formula... The strategic payoffs of the players in the teaching content game are calculated; among them, This represents the strategic payoff of the players in the game of teaching content. Indicates the information decay rate. This indicates the tenth weight.
[0135] As one possible implementation, the aforementioned determining module 603 is specifically used for: determining a first result based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player; determining multiple second results based on the first result; predicting the strategy payoff of the next time step based on the first result and each second result, iteratively executing until the difference between the predicted strategy payoff and the player's strategy payoff is less than a preset threshold, and taking the predicted strategy payoff at the end of the iteration as the target payoff; calculating the eigenvalues of the strategy evolution equations of each player at the target payoff; determining whether the strategy evolution equations are stable based on the eigenvalues; if so, determining the state evaluation result based on the average of the strategy payoffs corresponding to multiple time steps.
[0136] As one possible implementation, the aforementioned intervention module 604 is specifically used for: if the state assessment result indicates that the target teaching activity has an unfavorable situation, then for each candidate intervention strategy in the preset intervention strategy library, the current teaching state characteristics and the identification information of the candidate intervention strategy are concatenated to obtain concatenated information, and the concatenated information is input into a pre-trained benefit prediction model, which then performs benefit prediction based on the concatenated information to obtain the benefit of the target strategy after implementing the candidate intervention strategy; based on the benefit of the target strategy corresponding to each candidate intervention strategy, candidate intervention strategies that meet preset constraints are selected from multiple candidate intervention strategies as feasible intervention strategies, wherein the preset constraint is that the benefit of the target strategy after implementing the candidate intervention strategy is greater than the benefit of the strategy before implementing the candidate intervention strategy; and the feasible intervention strategy with the lowest intervention intensity is selected as the target intervention strategy.
[0137] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0138] This application also provides an electronic device 700, such as... Figure 7 The diagram shown is a structural schematic of an electronic device 700 provided in an embodiment of this application, including: a processor 701 and a memory 702, and optionally, a bus 703. The memory 702 stores machine-readable instructions executable by the processor 701. When the electronic device 700 is running, the processor 701 and the memory 702 communicate via the bus 703. When the machine-readable instructions are executed by the processor 701, the steps of the teaching information processing method described above are performed.
[0139] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the teaching information processing method described in any of the preceding claims.
[0140] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0141] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0142] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for processing teaching information, characterized in that, include: Real-time collection of multimodal data at the current time step of the target teaching activity, including the teacher's teaching data, the student group's response data, and the teaching content; The teacher, the student group, and the teaching content are each modeled as game players, and the strategy probability distribution of each game player in the corresponding strategy space is determined based on the multimodal data, and the strategy payoff of each game player is determined based on the multimodal data. Each of the teacher, the student group, and the teaching content corresponds to a strategy space, and each strategy space includes at least one strategy dimension, and each strategy dimension includes multiple dimension types. Based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player, the state evaluation result of the target teaching activity is determined. Based on the status assessment results, target intervention strategies are selected from the preset intervention strategy library and sent to the user terminal held by the teacher.
2. The method according to claim 1, characterized in that, Determining the strategy probability distribution of each player in the corresponding strategy space based on the multimodal data includes: The behavioral characteristics of each player are extracted from the multimodal data, and the behavioral characteristics of each player are quantitatively analyzed to obtain the quantitative results of the behavioral characteristics of each player. For each player, each strategy dimension in the strategy space corresponding to the player is traversed. For the current strategy dimension, the behavioral feature quantification result of the player is matched with the benchmark data corresponding to each dimension type under the current strategy dimension to obtain the matching degree between the player and each dimension type under the current strategy dimension. Based on the matching degree between each player and the corresponding strategy dimension type in the strategy space, the strategy probability distribution of each player in the corresponding strategy space is determined.
3. The method according to claim 2, characterized in that, The step of determining the strategy probability distribution of each player in the corresponding strategy space based on the matching degree between each player and the corresponding strategy dimension type in the strategy space includes: The dimension types under each strategy dimension in the strategy space corresponding to the game player are combined to obtain multiple dimension combinations, and each dimension combination includes one dimension type under multiple strategy dimensions. Based on the matching degree between each player and the corresponding strategy space, the probability result of each dimension combination is determined. The vector formed by combining the probability results of all dimensions is taken as the strategy probability distribution of the player in the corresponding strategy space.
4. The method according to claim 1, characterized in that, The step of determining the strategy payoffs of each player based on the multimodal data includes: Based on the teaching data, the deviation in teaching progress and the average test scores of students were determined; Based on formula The strategic payoff for the teacher in the game is calculated. in, This represents the strategic payoff for the teacher in the game. This represents the average test score of the students. This indicates the deviation in the teaching progress. Indicates the first weight. This indicates the second weight.
5. The method according to claim 1, characterized in that, The step of determining the strategy payoffs of each player based on the multimodal data includes: Based on the response data, determine the overall comprehension level and overall cognitive load; Based on formula The strategic payoffs of the student group in the game are calculated. in, This represents the strategic payoff of the student group in the game. This indicates the overall level of understanding. This indicates the overall cognitive load. Indicates the third weight. This indicates the fourth weight.
6. The method according to claim 5, characterized in that, The response data includes: cognitive response data, text interaction data, behavioral trajectory data, and physiological visual data; The determination of overall comprehension and overall cognitive load based on the response data includes: The question-and-answer accuracy rate and standard response time are determined based on the cognitive response data, the cognitive state is determined based on the text interaction data, the standard behavioral entropy is determined based on the behavioral trajectory data, and the standard physiological information is determined based on the physiological visual data. Based on formula The overall comprehension level is calculated. Among them, This indicates the overall level of understanding. This indicates the accuracy rate of the question and answer response. This indicates the cognitive state. Indicates the fifth weight. Indicates the sixth weight; Based on formula The comprehensive cognitive load was calculated. in, This indicates the overall cognitive load. This indicates the standard response time. This represents the standard behavioral entropy. This represents the standard physiological information. Indicates the seventh weight. Indicates the eighth weight. This indicates the ninth weight.
7. The method according to claim 1, characterized in that, The step of determining the strategy payoffs of each player based on the multimodal data includes: Based on the teaching content, determine the information decay rate of the game players in the teaching content; Based on formula The strategic payoffs of the players in the game of teaching content are calculated. in, This represents the strategic payoff of the player in the game based on the teaching content. This represents the information attenuation rate. This indicates the tenth weight.
8. The method according to claim 1, characterized in that, The determination of the state evaluation result of the target teaching activity based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player includes: Based on the strategy probability distribution of each player in the corresponding strategy space and the strategy payoff of each player, a first result is determined, and based on the first result, multiple second results are determined. Based on the first result and each of the second results, the strategy payoff for the next time step of the current time step is predicted, and the process is iterated until the difference between the predicted strategy payoff and the player's strategy payoff is less than a preset threshold. The predicted strategy payoff at the end of the iteration is then taken as the target payoff. Calculate the eigenvalues of the strategy evolution equations for each player at the target payoff; Based on the eigenvalues, determine whether the strategy evolution equation is stable; If so, the state evaluation result is determined based on the average of the strategy returns corresponding to multiple time steps.
9. The method according to claim 1, characterized in that, The step of selecting target intervention strategies from a pre-set intervention strategy library based on the state assessment results includes: If the state assessment result indicates that the target teaching activity has a poor state, then for each candidate intervention strategy in the preset intervention strategy library, the current teaching state characteristics are concatenated with the identification information of the candidate intervention strategy to obtain concatenated information, and the concatenated information is input into the pre-trained benefit prediction model. The benefit prediction model makes benefit prediction based on the concatenated information to obtain the benefit of the target strategy after executing the candidate intervention strategy. Based on the target strategy benefit corresponding to each of the candidate intervention strategies, candidate intervention strategies that meet preset constraints are selected from multiple candidate intervention strategies as feasible intervention strategies, wherein the preset constraints are that the target strategy benefit after executing the candidate intervention strategy is greater than the strategy benefit before executing the candidate intervention strategy. The feasible intervention strategy with the lowest intervention intensity is taken as the target intervention strategy.
10. An electronic device, characterized in that, include: A processor and a memory, the memory storing machine-readable instructions executable by the processor, wherein when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the instructional information processing method as described in any one of claims 1 to 9.