Teaching strategy processing method and apparatus, electronic device, and computer storage medium

By receiving learning information and generating candidate teaching strategies using a profiling model, and using a deviation risk identification model to detect and adjust strategy deviations, the problems of low transparency and entrenched biases in AI education are solved, enabling personalized and effective adjustment of teaching strategies.

CN121581489BActive Publication Date: 2026-05-29BEIJING SANSAN SMART EDUCATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SANSAN SMART EDUCATION TECHNOLOGY CO LTD
Filing Date
2025-11-17
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing AI recommendation engines in education suffer from problems such as low decision-making transparency, entrenched biases, and lagging risk control. This makes it difficult for teachers to trust them, biases may become entrenched, and risks cannot be intervened in real time, affecting educational equity and teaching effectiveness.

Method used

By receiving learning information and profiling models, candidate teaching strategies are generated, and deviation risk identification models are used to detect strategy deviations, generating deviation risk root cause reports. Intervention decision information is received from educational terminals, and adjustments are made to the target teaching strategy to achieve real-time intervention and personalized teaching.

Benefits of technology

It improves the transparency and fairness of teaching strategies, ensures the personalization and effectiveness of teaching strategies, reduces the impact of bias, and enables collaborative decision-making between teachers and AI.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121581489B_ABST
    Figure CN121581489B_ABST
Patent Text Reader

Abstract

The present disclosure provides a teaching strategy processing method and device, and relates to the technical fields of explainable artificial intelligence, educational technology and the like. The specific implementation scheme is as follows: receiving learning information and a portrait model of a teaching object; generating a candidate teaching strategy based on the portrait model; detecting whether the candidate teaching strategy has a teaching strategy bias risk based on the learning information and / or strategy information of the candidate teaching strategy; in response to detecting that the candidate teaching strategy has a teaching strategy bias risk, generating and sending a bias risk root cause report of the candidate teaching strategy to an education terminal; receiving teaching intervention decision information sent by the education terminal, and generating and sending a target teaching strategy to the teaching object based on the teaching intervention decision information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure pertains to the field of computer science, specifically relating to interpretable artificial intelligence, educational technology, and human-computer collaborative decision-making, and particularly to a teaching strategy processing method and apparatus, electronic equipment, and computer-readable storage medium. Background Technology

[0002] Current systems centered around AI (Artificial Intelligence) recommendation engines analyze students' historical learning data, such as homework accuracy, test scores, knowledge mastery, and video viewing time. They then use machine learning models—such as knowledge tracking models, reinforcement learning, or collaborative filtering—to predict students' current knowledge status and recommend the most suitable learning content for them, such as practice questions, instructional videos, or reading materials.

[0003] While existing AI recommendation engines improve the personalization efficiency of recommending learning content to students, they sacrifice transparency, fairness, and security in the decision-making process, creating a technological dilemma where teachers are difficult to trust, biases may become entrenched, and risks cannot be intervened in real time. Therefore, there is an urgent need in this field for a new technological paradigm that can solve these problems. Summary of the Invention

[0004] This disclosure provides a teaching strategy processing method and apparatus, an electronic device, and a computer-readable storage medium.

[0005] According to the first aspect, a teaching strategy processing method is provided, the method comprising: receiving learning information and a profile model of the learning subjects; generating candidate teaching strategies based on the profile model; detecting whether the candidate teaching strategies have a teaching strategy deviation risk based on the learning information and / or the strategy information of the candidate teaching strategies; generating and sending a deviation risk root cause report of the candidate teaching strategies to an educational terminal in response to detecting that the candidate teaching strategies have a teaching strategy deviation risk; receiving teaching intervention decision information sent by the educational terminal, and generating and sending a target teaching strategy to the learning subjects based on the teaching intervention decision information.

[0006] According to a second aspect, a teaching strategy processing apparatus is provided, comprising: a receiving unit configured to receive learning information and a profile model of the learning subjects; a strategy generation unit configured to generate candidate teaching strategies based on the profile model; a detection unit configured to detect whether the candidate teaching strategies have a teaching strategy deviation risk based on the learning information and / or the strategy information of the candidate teaching strategies; a report generation unit configured to generate and send a deviation risk root cause report of the candidate teaching strategies to an educational terminal in response to detecting that the candidate teaching strategies have a teaching strategy deviation risk; and a sending unit configured to receive teaching intervention decision information sent by the educational terminal, and generate and send a target teaching strategy to the learning subjects based on the teaching intervention decision information.

[0007] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first aspect.

[0008] According to a fourth aspect, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the method described in any implementation of the first aspect.

[0009] The teaching strategy processing method and apparatus provided in the embodiments of this disclosure first receive the learning situation information and profile model of the learning object; second, generate candidate teaching strategies based on the profile model; third, detect whether the candidate teaching strategies have a teaching strategy deviation risk based on the learning situation information and / or the strategy information of the candidate teaching strategies; then, in response to the detection of a teaching strategy deviation risk, generate and send a deviation risk root cause report of the candidate teaching strategy to the education terminal; finally, receive teaching intervention decision information sent by the education terminal, and generate and send a target teaching strategy to the learning object based on the teaching intervention decision information. Thus, by identifying and intervening in teaching strategy deviation risks in real time, the education terminal can accurately adjust teaching strategies, tailoring a unique learning path for each learning object, thereby improving the pertinence and effectiveness of personalized teaching.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0012] Figure 1This is a flowchart of an embodiment of the teaching strategy processing method according to this disclosure;

[0013] Figure 2 This is a structural diagram of the process for detecting deviations in teaching strategies.

[0014] Figure 3 This is the overall architecture diagram corresponding to the teaching strategy processing method disclosed in this publication;

[0015] Figure 4 This is a schematic diagram of the structure of one embodiment of the instructional strategy processing device disclosed herein;

[0016] Figure 5 This is a block diagram of an electronic device used to implement the teaching strategy processing method of the embodiments of this disclosure. Detailed Implementation

[0017] Unless otherwise expressly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "comprises" shall be understood to include the stated elements or components without excluding other elements or other components.

[0018] The technical solutions of this disclosure are illustrated below through specific embodiments. It should be understood that one or more steps mentioned in this disclosure do not preclude the existence of other methods and steps before or after the combined steps, or that other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are for illustrative purposes only and are not intended to limit the scope of this disclosure. Unless otherwise stated, the numbering of each method step is only for the purpose of identifying each method step, and not to limit the order of each method or to limit the scope of implementation of this disclosure. Changes or adjustments to their relative relationships, without substantial changes to the technical content, can also be considered as within the scope of implementation of this disclosure.

[0019] The raw materials and instruments used in the examples are not subject to any specific restrictions on their source; they can be purchased from the market or prepared according to conventional methods known to those skilled in the art.

[0020] However, as AI plays an increasingly important role in educational decision-making, its role is quietly shifting from an efficiency tool to assist teachers to a teaching agent capable of autonomous decision-making. This fundamental role shift exposes three profound and urgent shortcomings in existing technologies. These shortcomings not only limit the depth of technology application but also bring serious ethical challenges, constituting the core pain points that this publication aims to address:

[0021] 1. Lack of trust and oversight

[0022] Existing AI teaching systems, especially those employing complex models like deep learning, essentially operate as "black boxes" in their decision-making processes. A system might tell a teacher to "recommend these 5 questions for student A," but it cannot explain "why these 5 questions, and not another 5." For teachers, who bear the heavy responsibility of educating students, this incomprehensible and unverifiable recommendation logic is unacceptable. This directly leads to a "crisis of trust": teachers find it difficult to entrust the core decision-making power of teaching to a machine that cannot "communicate" with them. Therefore, teachers either blindly follow AI instructions or completely abandon AI suggestions, neither of which achieves effective human-machine collaboration and severely limits the value of AI technology.

[0023] 2. The risk of algorithmic bias solidification and amplification – a potential threat to educational equity

[0024] AI models learn from historical data, but historical data itself often harbors existing biases from human society. Current technologies not only fail to identify and eliminate these biases, but may even unconsciously solidify and amplify them in the pursuit of a single optimization goal (such as "maximizing scores"), posing a serious threat to educational equity. Specifically, this manifests in the following ways:

[0025] If historical data shows that male students spend more time on science subjects, the model might learn this association and thus tend to recommend more science content to male students, and vice versa. This could inadvertently reinforce gender stereotypes and limit students' potential for holistic development.

[0026] Students from families or schools with abundant educational resources typically have more complete and higher-quality online learning data. Models may therefore consider these students to have greater potential and recommend richer, more challenging learning paths for them. Conversely, students with sparse data may be conservatively recommended basic and repetitive content, leading to a digital Matthew effect and further exacerbating inequality in educational resources.

[0027] Under optimization focused on short-term score improvement, the model is highly likely to discover that high-intensity, repetitive mechanical training (i.e., endless practice problems) is the fastest path to achieving the goal. This leads the system to push a large number of homogeneous problems to students, stifling their learning interest, spirit of inquiry, and critical thinking, and may cause academic anxiety and cognitive overload, which runs counter to the concept of quality education and holistic development advocated by modern education.

[0028] 3. The lag in risk prevention and control

[0029] Even with sporadic academic research attempting to analyze the fairness of educational AI after the fact, this is far from sufficient in real-world teaching scenarios. By the time an analysis report points out potential bias in an AI's recommendation strategy weeks or months later, hundreds or thousands of students have already learned according to that strategy for an extended period, and the potential negative impact has already been inflicted and is difficult to reverse. Current technologies generally lack a mechanism for pre-screening and real-time intervention. Teaching is a dynamic and continuous process; the system must be able to identify erroneous teaching decisions before they are implemented and immediately return decision-making power to teachers with professional competence and ethical judgment.

[0030] To address the shortcomings of traditional technologies, this disclosure proposes a teaching strategy approach. Figure 1 A flow 100 is shown as an embodiment of a teaching strategy processing method according to the present disclosure, which includes the following steps:

[0031] Step 101: Receive the learning information and profile model of the students.

[0032] In this embodiment, the teaching object is the object that learns from the teaching resources provided by the educational terminal. The teaching object can be a learner or a learner's terminal, such as a learning terminal. The educational terminal can be a teacher's terminal, through which relevant information can be sent to the teacher, such as the teacher's teaching courseware.

[0033] In this embodiment, learning information is a collection of "real-time or near-real-time learning process and result data" of the learner on a specific course / knowledge point / task. It is an objective record of the learner's learning behavior in the current round or several recent rounds. Learning information includes: the learner's behavioral process, the evaluation of the results of the specific course / knowledge point / task, context and equipment, emotion and workload, content mapping, etc.

[0034] In this embodiment, the profile model is a "multi-dimensional, structured, and computable" state representation of the learning object, serving as the system's internal "learner state" for decision-making and generation. It includes both relatively stable attributes and dynamic estimates that change with learning. The profile model includes: a basic profile, abilities and tendencies, a knowledge state graph, evaluation parameters, etc. In each round of decision-making and generation of the teaching strategy, the profile model serves as the current state, constraining the path selection and content generation of the teaching strategy, and is continuously updated by new learning information.

[0035] In this embodiment, step 101 includes: acquiring the original data of the teaching objects (such as front-end tracking and interaction logs, questions and content services, tools and peripherals, external system interface information, etc.), performing data cleaning and mapping on the original data (such as event normalization and noise reduction, question and knowledge point mapping, problem-solving path standardization processing, etc.) and / or feature processing and indicator generation to obtain learning data; and updating the previous round of profile model online based on the learning data to obtain the profile model.

[0036] In this embodiment, both the learning information and the profiling model are dynamic data structures. They both include the current state of the learning object. The current state is data obtained after summarizing the state of the learning object. For example, both include the learning object's latest knowledge mastery status, recent practice performance, and other current status.

[0037] Step 102: Generate candidate teaching strategies based on the profile model.

[0038] In this embodiment, there are multiple ways to generate candidate teaching strategies through the profiling model. One specific way is to treat the process of generating candidate teaching strategies as a text generation task based on rich context. Specifically, step 102 includes: translating the profiling model into teaching context prompts that include all the contexts required for the teaching plan; inputting the teaching context prompts into a large language model to obtain the candidate teaching strategies output by the large language model.

[0039] In this embodiment, the candidate teaching strategy is a temporary, pending teaching decision strategy, which is a machine-readable data record composed of multiple predefined fields. The candidate teaching strategy includes: the identifier of the learner, the teaching objective, the strategy type, and the teaching content. The teaching objective specifies the teaching purpose the strategy aims to achieve, such as reinforcing a specific knowledge point. The strategy type defines the form of the teaching activity, such as practice, watching videos, or reading materials. The teaching content includes a list or set of specific teaching resource identifiers; for a practice strategy, it is a list of question identifiers; for a video strategy, it is the identifier of video resources.

[0040] Step 103: Based on the learning information and / or the strategy information of the candidate teaching strategies, detect whether the candidate teaching strategies have the risk of teaching strategy deviation.

[0041] In this embodiment, step 103 includes: parsing the learning information (including multi-dimensional data such as knowledge mastery, cognitive level, learning preferences and emotional state) and the strategy information of the candidate teaching strategies (such as teaching objectives, content difficulty, interaction methods, feedback mechanisms and evaluation standards, etc.). Then, a pre-set deviation risk identification model is used to perform a matching degree analysis on the two, focusing on detecting whether the strategy exceeds the target's ability range, violates their preference characteristics, or ignores their emotional changes, thereby determining whether the candidate teaching strategy has a teaching strategy deviation risk, and outputting early warning prompts and optimization suggestions when a risk is detected.

[0042] Optionally, step 103 above also includes: constructing a "digital twin" virtual student cohort based on learning information. Specifically, using learning information, constructing profile data of real teaching subjects, using all real student profile data as training data, and training a generative model (e.g., variational autoencoder VAE or generative adversarial network GAN). This generative model can generate a large number of virtual student profiles that are highly consistent with the statistical distribution of the real student cohort. All simulations are conducted on virtual data, without infringing on the privacy of real students. Multiple virtual students can be easily generated, amplifying the cumulative effect of small deviations and making them easier to observe. Groups with specific characteristics (e.g., a 1:1 male-to-female ratio, a 3:7 urban-rural ratio) can also be generated as needed for more targeted stress testing. A strategy prototype template is defined, which includes strategy type, knowledge points, target groups for the knowledge points, and target requirements. After inputting the various feature parameters of the strategy prototype template, specific strategy prototype data can be obtained. The core features of the candidate teaching strategy S_cand are extracted, and based on the strategy prototype template, the core features are abstracted into a strategy prototype data P_arch. For example, a strategy prototype data set can be defined as: P_arch={type:“Physics Practice”, properties:[“High Math Requirements”, “Low Conceptuality”], target_group:“Math Prodigies”}. For this strategy prototype data set, a Fast-Forward Social Simulation is performed, specifically simulating a cyclical process: the system initiates a multi-round fast simulation. In each round, for each virtual student in the virtual student group, if their profile matches the target group characteristics of the strategy prototype data P_arch (e.g., “Math Prodigies”), the system simulates the impact of applying the prototype strategy. This simulation process completes hundreds or even thousands of iterations within seconds, equivalent to simulating weeks or months of teaching in the real world.

[0043] The abnormal evolution of group macro indicators is monitored through key macro indicators (each with its own threshold). These key macro indicators include: Knowledge Gini Coefficient: This measures the gap in knowledge levels within a group. A rapid increase in this coefficient during simulation indicates that the strategy prototype is exacerbating the knowledge gap among students (a manifestation of socioeconomic bias). Cohort Interest Entropy: This measures the diversity of learning interests within the entire group. A rapid decrease in entropy indicates that the strategy prototype is guiding all students to focus on a few areas, stifling personalized development (a manifestation of pedagogical bias). Cross-Group Mean Ability Gap: For example, calculating the difference in average scores between virtual male and virtual female groups on the dimension of "physical concept understanding." A continuously widening gap during simulation indicates gender bias in the strategy prototype. Time series anomaly detection algorithms (such as Isolation Forest and One-Class SVM) are used to determine whether the evolution curves of these macro indicators are "normal." If the evolution of any macroeconomic indicator is deemed abnormal (e.g., the Gini coefficient rises sharply after the simulation begins), the prototype data P_arch of the strategy is deemed to have a potential long-term bias risk, thus the original candidate strategy S_cand is deemed to be high-risk.

[0044] Step 104: In response to the detection that the candidate teaching strategy has a teaching strategy deviation risk, a deviation risk root cause report of the candidate teaching strategy is generated and sent to the education terminal.

[0045] In this embodiment, the educational terminal is the terminal used by the educational subject (such as a teacher), and the deviation risk root cause report is a report that records the content and causes of deviation risks in teaching strategies. Through the deviation risk root cause report, the educational subject can be informed whether the current candidate teaching strategy is reasonable or unreasonable, and the reasons for being reasonable or unreasonable can be given.

[0046] In this embodiment, when the execution entity running on the teaching strategy processing method detects a teaching strategy deviation risk in a candidate teaching strategy, it immediately activates the deviation tracing mechanism: First, based on the strategy content, historical teaching logs, and student feedback data, a pre-set deviation identification model is used to locate the root cause of the risk (such as incomplete knowledge graph coverage, unbalanced difficulty gradient design, or bias in personalized recommendation algorithms); then, a visual deviation risk root cause report is automatically generated, explaining the cause of the deviation in natural language (e.g., "This strategy relies too much on rote memorization templates for the 'geometric proof' knowledge point, neglecting logical reasoning training, which is inconsistent with regional teaching and research standards"), and is pushed to the education terminal through the student terminal or teacher workbench, along with correction suggestions and links to alternative strategies, supporting one-click replacement or manual review.

[0047] Step 105: Receive teaching intervention decision information sent by the education terminal, and based on the teaching intervention decision information, generate and send the target teaching strategy to the teaching object.

[0048] In this embodiment, the teaching intervention decision information is information for correcting candidate teaching strategies. This information may include: the identifier of the learner, the teaching objective, the intervention type, and the revised content. The intervention type is the target type of the strategy, and the revised content is the target content of the teaching content. Through the teaching intervention decision information, the corresponding content in the candidate teaching strategy can be revised, thereby modifying the candidate teaching strategy into the target teaching strategy. Optionally, the teaching intervention decision information may also include the diagnostic results after diagnosing the current state of the learner.

[0049] In this embodiment, the execution entity running on the teaching strategy processing method receives teaching intervention decision information sent by the educational terminal (such as a teacher or teaching system). This information includes the diagnostic results of the current learning status of the teaching object (such as a student) and the type of intervention to be taken (such as reinforcement practice, prompting guidance, or content re-examination). Subsequently, based on the built-in teaching strategy matching model, the system maps the intervention decision to a specific executable target teaching strategy (such as pushing 3 variation questions, playing 2 minutes of animated explanation, or starting peer discussion), and immediately sends the strategy to the teaching object in a form that the teaching object can understand (such as text, voice, or interactive interface), realizing a closed loop from "decision" to "execution".

[0050] Optionally, in response to the teaching intervention decision information being the editing or rejection of candidate teaching strategies, the learning information of the teaching subjects, candidate teaching strategies, and teaching intervention decision information are constructed into a teaching instruction-response pair; based on the teaching instruction-response pair, the strategy generation model used to generate candidate teaching strategies is fine-tuned through supervised learning.

[0051] The teaching strategy processing method provided in this disclosure first receives learning information of the learners; second, it generates candidate teaching strategies based on the learners' profile model; third, it detects whether the candidate teaching strategies have a risk of deviation based on the learning information and / or the strategy information of the candidate teaching strategies; then, in response to the detection of a risk of deviation in the candidate teaching strategies, it generates and sends a root cause report of the deviation risk of the candidate teaching strategies to the education terminal; finally, it receives teaching intervention decision information sent by the education terminal, and generates and sends a target teaching strategy to the learners based on the teaching intervention decision information. Thus, by identifying and correcting teaching strategy deviations in real time, it ensures that the optimal teaching plan is accurately matched for each learner, significantly improving teaching effectiveness and personalized experience.

[0052] In some optional implementations of this disclosure, the method further includes: in response to detecting that the candidate teaching strategy does not have the risk of teaching strategy deviation, taking the candidate teaching strategy as the target teaching strategy and sending the target teaching strategy to the teaching object.

[0053] In this optional implementation, after confirming that the candidate teaching strategy has no risk of deviation after testing, the execution entity running on the teaching strategy processing method immediately locks it as the target teaching strategy, and pushes the target teaching strategy in the form of a structured data packet to the corresponding teaching object in real time through the established communication link via the terminal device. At the same time, the strategy distribution log is recorded and the strategy reception confirmation mechanism of the teaching object is triggered to ensure the accurate transmission of the teaching strategy and the traceability of subsequent execution.

[0054] In some optional implementations of this disclosure, the aforementioned detection of whether a candidate teaching strategy has a teaching strategy bias risk based on student learning information and / or candidate teaching strategy information includes: calculating the recommendation probability of recommending the candidate teaching strategy to a specific gender group based on student learning information and strategy information, and / or the correlation probability between the scores of learning resources in the candidate teaching strategy and the socioeconomic background quantitative scores of the learners; obtaining a fairness risk value based on the recommendation probability and / or correlation probability; calculating the occurrence probability of knowledge points in the candidate teaching strategy and the difficulty value of the candidate teaching strategy based on strategy information; obtaining a pedagogical risk value based on the occurrence probability and difficulty value; calculating the length of the consecutive failure sequence of the learners' historical answer sequences corresponding to the candidate teaching strategy based on student learning information and strategy information; and detecting whether the candidate teaching strategy has a teaching strategy bias risk based on the fairness risk value, the pedagogical risk value, and the length of the consecutive failure sequence.

[0055] In this optional implementation method, such as Figure 2 As shown, the risks of deviations in teaching strategies include: fairness risk, pedagogical risk, and emotional risk; in Figure 2In this process, after obtaining candidate teaching strategies, fairness risk assessment, pedagogical risk assessment, and affective risk assessment can be conducted on the candidate teaching strategies respectively. Figure 2 The risk aggregation and judgment process includes: fairness risk assessment of candidate teaching strategies to obtain fairness risk value; pedagogical risk assessment of candidate teaching strategies to obtain pedagogical risk value; and affective risk assessment of candidate teaching strategies to obtain the length of the continuous failure sequence.

[0056] In this optional implementation, the strategy information is information that describes the candidate teaching strategies, and may include the type of the candidate teaching strategy, the strategy object, and the content description.

[0057] In this optional implementation, the occurrence probability reflects the probability of knowledge points appearing in the candidate teaching strategies, and the difficulty value reflects the ease with which the learners can recognize the candidate teaching strategies. The length of the consecutive failure sequence of the learners' historical answer sequences corresponding to the candidate teaching strategies reflects the learners' accumulated frustration.

[0058] In this optional implementation, a deep neural network is obtained, comprising: an input layer, a shared underlying network, and a task-specific pyramidal network. The input layer encodes and concatenates all relevant information (learning information and / or strategy information of candidate teaching strategies) to form a high-dimensional input vector. Specifically, the learning information includes: static features of the learners (such as gender, socioeconomic background scores), dynamic knowledge status, historical behavioral statistics, and historical answer sequences. Encoding the learning information includes: one-hot or embedding encoding of the static features of the learners; normalization of the numerical features such as the dynamic knowledge status and historical behavioral statistics of the learners; and transformation of the historical answer sequences of the learners into a fixed-length context vector using a recurrent neural network or a Transformer encoder. The strategy information includes: knowledge points, resource identifiers, and metadata involved in the candidate teaching strategies. Encoding the strategy information includes: embedding encoding of the category features such as knowledge points and resource identifiers involved in the candidate teaching strategies; and normalization of the metadata.

[0059] In this optional implementation, the core task of the shared network is to learn a general, information-rich joint feature representation h_shared from the original input. This joint feature representation h_shared captures the complex, non-linear interactions between learning and policy, which are the underlying information needed to assess all types of risk.

[0060] In this optional implementation, the task-specific tower network includes: a fairness risk tower, a pedagogical risk tower, and an affective risk tower. The fairness risk tower employs at least two fully connected layers (e.g., at least two layers with Softmax activation) to output the probability of correlation between estimated resource scores and learners' socioeconomic scores within the estimated strategy. The pedagogical risk tower employs at least two fully connected layers (layers with Softmax activation and linearly activated neurons). The number of Softmax-activated layers is the same as the number of knowledge points in the candidate teaching strategies. The Softmax-activated layers output the probability of occurrence of knowledge points in the candidate teaching strategies, and the linearly activated neurons output the difficulty value of the candidate teaching strategies. The affective risk tower employs multiple fully connected layers (e.g., neurons with ReLU activation) to output the length of consecutive failure sequences in the historical answer sequences corresponding to the candidate teaching strategies.

[0061] In this optional implementation, the execution entity running on the teaching strategy processing method first performs three calculations based on the learner's gender, socioeconomic background, and other learning information, as well as the learning resources, knowledge points, and difficulty information involved in the candidate teaching strategies: First, it estimates the probability that the strategy will be recommended to a certain gender group, and estimates the correlation probability between the resource score within the strategy and the learner's socioeconomic quantitative score, and combines the two probabilities to obtain a fairness risk value; Second, it calculates the pedagogical risk value by statistically analyzing the probability of each knowledge point appearing in the strategy and combining it with the overall difficulty value; Third, it retrieves the learner's historical answer sequence and calculates the length of the consecutive failure sequence on the questions corresponding to the strategy; Finally, it inputs the fairness risk value, the pedagogical risk value, and the consecutive failure length into a weighted or threshold model. If the output exceeds a preset threshold, it is determined that the candidate teaching strategy has a teaching strategy bias risk.

[0062] Optionally, the above calculation of the probability of occurrence of knowledge points in candidate teaching strategies and the difficulty value of candidate teaching strategies based on strategy information includes: calculating the probability of occurrence of knowledge points in candidate teaching strategies by using the knowledge point entropy algorithm, specifically, the knowledge point entropy algorithm can adopt the formula shown in equation (1); calculating the difficulty value by using the difficulty z-Score algorithm, specifically, the difficulty algorithm can adopt the formula shown in equation (2).

[0063] (1)

[0064] In equation (1), These are knowledge points in the candidate teaching strategies. The lower the entropy value, the higher the repetition rate.

[0065] (2)

[0066] In equation (2), It is the average difficulty of the strategy. It represents the mean and standard deviation of the difficulty level of students' history answers.

[0067] Optionally, the above-mentioned method of obtaining the teaching method risk value based on the probability of occurrence and the difficulty value includes: weighted summation of the probability of occurrence and the difficulty value to obtain the teaching method risk value.

[0068] Optionally, the above calculation of the consecutive failure sequence length of the historical answer sequence corresponding to the candidate teaching strategy for the teaching object based on the learning information and strategy information includes: obtaining the consecutive failure sequence length by the length calculation method shown in equation (3).

[0069] (3)

[0070] In equation (3), It is a sequence of students' answers. This indicates an incorrect answer.

[0071] In some optional implementations of this disclosure, the above-mentioned calculation of the recommendation probability of recommending candidate teaching strategies to a specific gender group based on learning information and strategy information, and / or the correlation probability between the scores of learning resources in the candidate teaching strategies and the quantitative scores of the socioeconomic background of the teaching subjects; obtaining the fairness risk value based on the recommendation probability and / or correlation probability includes: determining a first number of strategy identifiers for each gender group and each gender group within the specific gender group based on learning information; determining the target strategy identifier of the candidate teaching strategies based on strategy information, and selecting a second number of target strategy identifiers from all strategy identifiers; subtracting the second number from the first number for each gender group to obtain the recommendation probability; and / or dividing the covariance between the resource scores of learning resources in the candidate teaching strategies and the quantitative scores of the socioeconomic background of the teaching subjects by the standard deviation of the resource scores and the standard deviation of the quantitative scores of the socioeconomic background of the teaching subjects to obtain the correlation probability; and weighting and summing the recommendation probability and the correlation probability to obtain the fairness risk value.

[0072] In this optional implementation, the recommendation probability is the difference in the probability of recommending the corresponding strategy among multiple different gender groups, which is used to characterize the difference in population equality, i.e. gender bias. The recommendation probability can also be calculated by the difference in the probability of recommending candidate teaching strategies among various given gender groups in a specific gender group. Specifically, the formula for calculating the recommendation probability is shown in equation (4).

[0073] (4)

[0074] In equation (1), It represents the probability of recommending the candidate teaching strategy within a given gender group. It is the probability of recommending the candidate teaching strategy in another given gender group, and DPD represents the difference in probabilities between the two different genders.

[0075] In this optional implementation, the correlation probability is used to characterize the correlation between the quality of learning resources and the background. The calculation of the correlation probability is shown in Equation (5).

[0076] (5)

[0077] In equation (5), Q represents the relevant probability, Q is the resource score, and B is the quantitative score of the socioeconomic background of the teaching object.

[0078] In this optional implementation, the executing entity first calculates the total number of strategy identifiers for each gender group within a specific gender group based on the learning information, and records this as the first quantity. Then, it identifies the target strategy identifier for the current candidate teaching strategy from the strategy information and counts the number of instances of this target identifier in the strategy list for that gender group, recording this as the second quantity. The difference between the second quantity and the first quantity for each gender group is used as the recommendation probability for recommending the strategy to that gender group. Simultaneously, a Peelman correlation is performed between the resource scores of each learning resource within the strategy and the socioeconomic background score of the learners, i.e., the covariance of the two is divided by the standard deviation of the resource scores and the standard deviation of the background score to obtain the correlation probability. Finally, the recommendation probability and the correlation probability are weighted and summed according to preset weights to output a fairness risk value; a higher value indicates a greater recommendation bias for that gender group.

[0079] In some optional implementations of this disclosure, the above-mentioned response to detecting that a candidate teaching strategy has a teaching strategy deviation risk, generating and sending a deviation risk root cause report of the candidate teaching strategy to the education terminal includes: in response to detecting that the teaching strategy deviation risk is obtained through the fairness risk value, using the Shapley additive interpretation algorithm to predict the contribution of the characteristics in the candidate teaching strategy, and obtaining a deviation risk root cause report including the contribution information of each characteristic; in response to detecting that the teaching strategy deviation risk is obtained through the pedagogical risk value, using the counterfactual interpretation algorithm to predict the cognitive optimization strategy of the candidate teaching strategy, and obtaining a deviation risk root cause report including the cognitive optimization strategy.

[0080] In this optional implementation, the SHAP (SHapley Additive exPlanations) value Representation of features The contribution to the output of the explanation model. Its mathematical form is shown in equation (6).

[0081] (6)

[0082] In equation (6) It is an explanatory model. It simplifies the input.

[0083] In a specific example, a risk source report for bias could be: “The main contributing features of the system’s recommended physics exercises are historical physics scores (+0.7) and gender = male (+0.25). The latter may introduce undesirable bias.”

[0084] In this optional implementation, a counterfactual explanation can be used to solve an optimization problem, such as by using the optimization algorithm shown in equation (7) to find the original student profile. The closest possible, but one that allows the model to output a "safe" policy. portrait .

[0085] (7)

[0086] Through the portrait Analyzing the contributing factors yields optimization strategies. A specific example of the root cause report of bias risk is: "If the student's recent knowledge mastery rate increases from 70% to 85%, or if the recommended question P3 is replaced with a less difficult question, the cognitive overload risk will disappear." In this example, "increasing the student's recent knowledge mastery rate from 70% to 85%" and "replacing the recommended question P3 with a less difficult question" are both cognitive optimization strategies.

[0087] like Figure 2 As shown, after detecting the risk of teaching strategy deviation in candidate teaching strategies through risk aggregation and judgment, the XAI interpretation generator is used to interpret the bias and cognitive overload of the features to obtain the contribution information and / or cognitive optimization strategies of the corresponding features. The contribution information and / or cognitive optimization strategies can be added to the root cause report of deviation risk.

[0088] In this optional implementation, when a bias risk is detected in a candidate teaching strategy, a personalized bias risk root cause report is automatically generated and pushed according to the risk type: If the bias is triggered by the fairness risk value, the SHAP (Shapley Additive Explanation) algorithm is called to quantify the marginal contribution of each feature in the strategy (such as student gender, region, prior performance, etc.), and generate a feature explanation report sorted by contribution, visually presenting to the education terminal "to what extent do these features increase unfair scores"; if the bias is triggered by the pedagogical risk value, the counterfactual explanation algorithm is used instead, and by simulating comparisons such as "if cognitive path X in the strategy is replaced with path Y, the risk will decrease from R1 to R2", a feasible cognitive optimization strategy (such as adjusting knowledge granularity, remedying prerequisite concepts, increasing the frequency of interactive feedback) is automatically derived, and this optimization strategy is sent to the teacher or system operator in real time as the core content of the root cause report, realizing a closed loop of "detection-source tracing-improvement".

[0089] Optionally, when a candidate teaching strategy is detected to have no teaching strategy deviation risk, the result of not having teaching strategy deviation risk can be directly added to the deviation risk root cause report, thereby providing feedback on the status of the candidate teaching strategy to the learners.

[0090] In some optional implementations of this disclosure, the above-mentioned receiving teaching intervention decision information sent by the educational terminal and generating and sending a target teaching strategy to the teaching subject based on the teaching intervention decision information includes: receiving teaching intervention decision information sent by the educational terminal; detecting whether the teaching intervention decision information is a cognitive optimization strategy in the deviation risk root cause report; in response to detecting that the teaching intervention decision information is a cognitive optimization strategy, using the cognitive optimization strategy as the target teaching strategy and sending the target teaching strategy to the teaching subject; in response to detecting that the teaching intervention decision information is not a cognitive optimization strategy, fine-tuning the strategy generation model based on the teaching intervention decision information, generating the target teaching strategy using the fine-tuned strategy generation model, and sending the target teaching strategy to the teaching subject.

[0091] In this optional implementation, the strategy generation model is a pre-trained model that generates strategies. By using teaching intervention decision information as training data and fine-tuning the strategy generation model, the prediction accuracy of the strategy generation model can be improved.

[0092] In this optional implementation, the policy generation model described above is a personalized policy network. Policy Network It will receive student profile data s as input, and then the policy generation model will output a candidate teaching strategy.

[0093] In this optional implementation, the above-mentioned fine-tuning strategy generation model based on teaching intervention decision information includes:

[0094] The first step involves constructing a teaching instruction dataset based on teaching intervention decision information. When the implementing entity of the teaching strategy processing method detects a teacher's intervention decision as "not approved" (i.e., rejection or editing), a structured instruction fine-tuning data point is automatically constructed. The teaching instruction data in the dataset is preference data; every teacher intervention (such as "approval," "rejection," or "editing") is recorded. For example, a single "rejection" operation constitutes a preference pair. The rejected strategy is a negative example.

[0095] For example, the teaching intervention decision information includes: a student profile model (P_student), describing the student's complete state at the time of the decision; the AI's initial strategy (S_cand): the strategy rejected by the teacher; and the teacher's intervention decision information (D_teacher), which is {action: "reject"} if it is a rejection, and {action: "edit", final_strategy: S_final} if it is an edit, where S_final is the teacher's final edited teaching strategy. The implementing entity translates the above teaching intervention decision information into a high-quality "instruction-response" pair. The instruction part details the entire context of the decision, simulating the teacher's thought process during the decision-making process. The response part is the teacher's final decision, formatted as the model's ideal output; when the teacher selects "edit", the response is the structured representation of the teacher's edited final strategy S_final; when the teacher selects "reject" but does not provide a new strategy, the response can be a special marker or explanatory text, such as {"rejection_reason": "This strategy does not meet the current teaching objectives; the focus should be on conceptual understanding rather than calculation."}.

[0096] The second step is Supervised Fine-Tuning (SFT), which aggregates all instruction-response data points collected within a period (e.g., one day or one week) into a fine-tuning dataset. The optimization objective of the policy generation model is to minimize the difference between its generated content and the "gold standard" response provided by the teacher. This is achieved by calculating the cross-entropy loss Loss = -Σ log P(Response | Instruction). Cross-entropy directly tells the model: "When encountering a similar instruction again, this response should be generated directly."

[0097] Step 3: Generate the target teaching strategy. After fine-tuning, the implementing entity obtains a fine-tuned strategy generation model, Model_finetuned. When teacher intervention occurs, the system can choose to immediately use Model_finetuned to generate a new target teaching strategy for the current student based on business logic.

[0098] In this optional implementation, the system first receives teaching intervention decision information sent by the educational terminal and checks it to determine whether it has adopted the cognitive optimization strategy provided in the deviation risk root cause report. If the check result is satisfactory, the cognitive optimization strategy is directly sent to the students as the target teaching strategy. If the check result is unsatisfactory, the system will fine-tune the strategy generation model based on the teaching intervention decision information, and regenerate the target teaching strategy using the fine-tuned model before sending it to the students, thereby realizing personalized and dynamically adjusted teaching strategy generation and delivery.

[0099] In some optional implementations of this disclosure, the above-mentioned response to detecting that the teaching intervention decision information does not adopt the cognitive optimization strategy, fine-tuning the strategy generation model based on the teaching intervention decision information, generating a target teaching strategy using the fine-tuned strategy generation model, and sending the target teaching strategy to the teaching object includes: in response to detecting that the teaching intervention decision information does not adopt the cognitive optimization strategy, adding the teaching intervention decision information to the preference dataset of the educational terminal; updating the internal parameters of the strategy model based on the preference dataset using a reinforcement learning algorithm, while ensuring that the deviation between the fine-tuned teaching strategy and the original teaching strategy is within a preset range.

[0100] In this optional implementation, the objective function of the strategy model is used to maximize the score of the reward model. The reward model scores the teaching strategy based on the preference dataset. Its goal is to predict the probability that a strategy will be accepted by the teacher. The input of the reward model is a strategy S, and the output is a scalar reward. The loss function of the reward model is shown in Equation (8).

[0101] (8)

[0102] In equation (8), It is a dataset of collected teacher preferences. It is a strategy that is recognized by teachers. This is a strategy that has not been recognized by teachers.

[0103] In this optional implementation, updating the internal parameters of the policy model using a reinforcement learning algorithm based on the preference dataset is a way to fine-tune the policy generation model, which uses a reward model. As the evaluator, the policy network is updated using a reinforcement learning algorithm. The internal parameters of the function are such that this fine-tuning step occurs asynchronously (e.g., when the system is idle). Its objective function... Driven strategy generation model Adjust its parameters so that the strategies it generates in the future can obtain rewards from the model. A higher score. Specifically, a reinforcement learning algorithm is used to optimize the policy network. The objective function is shown in equation (9).

[0104] (9)

[0105] In equation (9), the first term Used to drive policy networks The strategy generates a policy that can achieve a high score in the reward model (i.e., aligns with teacher preferences). The second term is a KL divergence penalty term to prevent the fine-tuned policy network from being penalized. Compared with the original strategy If the deviation is too large, maintain the stability of the model performance.

[0106] Through this closed loop, the implementing entities on which the teaching strategy processing methods operate will gradually learn to anticipate and proactively avoid strategy types that may be blocked by the teaching strategy deviation risk detection engine or rejected by teachers, making the entire system increasingly intelligent and rule-abiding.

[0107] In this optional implementation, if the system detects that the teaching intervention decision is "not through cognitive optimization strategy", the decision information is immediately added to the corresponding teaching object's preference dataset. Then, using the preference dataset as a feedback signal, the strategy generation model is fine-tuned using a reinforcement learning algorithm to obtain an updated strategy model. Finally, the fine-tuned model outputs a new target teaching strategy in real time and pushes it to the teaching object.

[0108] The following is combined Figure 4 The specific implementation steps of the teaching strategy processing method disclosed herein are illustrated in one scenario example as follows:

[0109] Target Audience: Student: Zhang San, a second-year high school girl majoring in science. She excels in mathematics but struggles with understanding abstract concepts such as Lenz's law in the newly introduced physics chapter on "Electromagnetic Induction." Teacher: Mr. Wang, Zhang San's physics teacher, with extensive teaching experience. System Environment: The school's personalized smart learning platform, which integrates this publicly available system. Time: Wednesday evening; the system is planning Zhang San's physics reinforcement exercises for the evening.

[0110] The detailed process is as follows: The data layer aggregates and preprocesses data through the data aggregation and preprocessing modules to build or update student profile models. In this data layer, the latest student profile data for Zhang San is first aggregated. Key information is as follows: student_id: "Zhang San 2025"; demographics: { gender: "female"}; historical_scores: { math: 95, physics: 82}; knowledge_state (based on knowledge graph): { ..., kpid_faraday_law: 0.6, kpid_lenz_law: 0.35, ...} (Faraday's law mastery is acceptable, Lenz's law mastery is very low); recent_performance_stats: { difficulty_mean (μ_hist): 0.60, difficulty_std (σ_hist): 0.12} (The average difficulty of recent practice problems is 0.6, with a standard deviation of 0.12.) The personalized strategy generation module in the AI ​​decision-making layer generates candidate teaching strategies. This module (a deep reinforcement learning model) makes a decision after receiving Zhang San's profile model. This personalized strategy model may have learned a latent association in its training data: students who are "good at math" and "male" are able to handle highly difficult, computationally intensive physics problems. Therefore, the personalized strategy model generates a candidate teaching strategy. Target: "Consolidate the 'Electromagnetic Induction' chapter"; Content: [P_01, P_02, P_03, P_04, P_05] (a practice set containing 5 questions); Content Analysis: P_01, P_02: Basic concept judgment questions (Difficulty: 0.4); P_03, P_04, P_05: Comprehensive calculation problems involving complex circuits and calculus operations, requiring less intuitive understanding of Lenz's law but higher mathematical ability (Average difficulty: 0.9); Overall Strategy Difficulty: (Calculation error, should be (0.4*2+0.9*3) / 5=3.5 / 5=0.7) (Correction:) ) (Correction: — Corrected calculation: This calculation is incorrect; it should be: Correction: The overall average difficulty of the strategy is... Corrected calculation: Correction: The overall average difficulty of the strategy is Corrected calculation: (0.4*2 +0.9*3) / 5 = (0.8 + 2.7) / 5 = 3.5 / 5 = 0.7. (Corrected and clarified calculation: The average difficulty of this practice set is...) )

[0111] Step 3: Real-time monitoring module for teaching strategy deviation risk. In this module... The data is sent to the ethical risk monitoring and XAI engine, triggering a parallel risk assessment: pedagogical risks, such as cognitive overload assessment, with the difficulty Z-Score as the metric.

[0112] ,result: (Preset threshold). Cognitive overload risk was not triggered.

[0113] Fairness bias and gender bias assessment, such as population equality difference (DPD), estimate the probability of recommending a strategy of this type (i.e., "a physics practice set containing more than 50% high-difficulty calculation problems") across different gender groups by querying a pre-trained meta-model or performing a fast Monte Carlo simulation.

[0114] ; ; ,result: (Preset threshold) The risk of gender bias is triggered!

[0115] Step 4: Generate XAI explanation. Due to the risk of gender bias being triggered, the executing entity calls the associated XAI method—SHAP analyzer. The SHAP analyzer operates on the policy generation model to calculate the contribution of each input feature in this decision, and outputs: a structured root cause report of bias risk.

[0116] { "risk_type": "Gender Bias", "explanation": { "feature_contributions": [ {"feature": "knowledge_state.kpid_lenz_law", "value": 0.35,"shap_value": +0.45}, {"feature": "historical_scores.math", "value": 95, "shap_value": +0.38}, {"feature": "demographics.gender", "value": "Male", "shap_value": +0.22}, ... ], "summary": "Decision-making is significantly positively influenced by the 'gender' characteristic."}}

[0117] Step 5: The teacher intervention interface module presents the root cause report of deviation risk. Teacher Wang receives an instant push notification from the system on her teacher's interface.

[0118] Step Six: Teacher Decision-Making and Intervention. After seeing the alarm, Teacher Wang, based on her understanding of Zhang San, quickly made a judgment: "The system is correct. What Zhang San lacks now is not calculation ability, but an understanding of Lenz's Law's core idea of ​​'resisting change.' These difficult questions will get him bogged down in the calculation details, bypassing the real difficulties." So, Teacher Wang clicked the button.

[0119] The system displayed a question basket showing the original five questions. Teacher Wang removed the three difficult questions (P_03, P_04, and P_05) and selected three new questions from the system's recommended list of "related concept questions": P_06: a graph analysis question to determine the change in magnetic flux based on experimental phenomena (difficulty: 0.5); P_07: a qualitative exercise to determine the direction of induced current (difficulty: 0.55); P_08: a simple semi-quantitative calculation question (difficulty: 0.6). This selected information will be used to generate teaching intervention decision information. Based on this information, the implementing entity will ultimately formulate a target-oriented teaching strategy. This teaching strategy It was saved.

[0120] Step 7: Strategy Implementation and Closed-Loop Feedback. The implementing entity will use the more targeted target-oriented teaching strategies edited by Teacher Wang. It was distributed to student Zhang San.

[0121] Further reference Figure 4As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a teaching strategy processing device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0122] like Figure 4 As shown, the teaching strategy processing device 400 provided in this embodiment includes: a receiving unit 401, a strategy generation unit 402, a detection unit 403, a report generation unit 404, and a sending unit 405. The receiving unit 401 can be configured to receive learning information and a profile model of the learning subjects. The strategy generation unit 402 can be configured to generate candidate teaching strategies based on the profile model of the learning subjects. The detection unit 403 can be configured to detect whether a candidate teaching strategy has a teaching strategy deviation risk based on the learning information and / or the strategy information of the candidate teaching strategies. The report generation unit 404 can be configured to generate and send a deviation risk root cause report of the candidate teaching strategy to the educational terminal in response to the detection of a teaching strategy deviation risk. The sending unit 405 can be configured to receive teaching intervention decision information sent by the educational terminal, and generate and send a target teaching strategy to the learning subjects based on the teaching intervention decision information.

[0123] In this embodiment, the specific processing of the receiving unit 401, the strategy generation unit 402, the detection unit 403, the report generation unit 404, and the sending unit 405 in the teaching strategy processing device 400, and the resulting technical effects, can be found in reference to [reference needed]. Figure 1 The relevant descriptions of steps 101, 102, 103, 104, and 105 in the corresponding embodiments will not be repeated here.

[0124] In some embodiments of this disclosure, the teaching strategy processing device 400 further includes a calibration unit (not shown in the figure), which is configured to: in response to detecting that a candidate teaching strategy does not have the risk of teaching strategy deviation, select the candidate teaching strategy as the target teaching strategy and send the target teaching strategy to the teaching object.

[0125] In some embodiments of this disclosure, the detection unit 403 is further configured to: calculate the recommendation probability of recommending candidate teaching strategies to a specific gender group based on learning information and strategy information, and / or the correlation probability between the scores of learning resources in the candidate teaching strategies and the socioeconomic background quantitative scores of the teaching subjects; obtain a fairness risk value based on the recommendation probability and / or correlation probability; calculate the occurrence probability of knowledge points in the candidate teaching strategies and the difficulty value of the candidate teaching strategies based on strategy information; obtain a pedagogical risk value based on the occurrence probability and difficulty value; calculate the length of the consecutive failure sequence of the historical answer sequence corresponding to the candidate teaching strategies for the teaching subjects based on learning information and strategy information; and detect whether the candidate teaching strategies have a teaching strategy deviation risk based on the fairness risk value, the pedagogical risk value, and the length of the consecutive failure sequence.

[0126] In some embodiments of this disclosure, the detection unit 403 is further configured to: determine a first number of strategy identifiers for each gender group and all strategy identifiers for each gender group within a specific gender group based on learning information; determine a target strategy identifier for a candidate teaching strategy based on strategy information, and select a second number of strategy identifiers belonging to the target strategy identifier from all strategy identifiers; calculate the difference between the second number and the first number for each gender group to obtain a recommendation probability; and / or divide the covariance between the resource score of the learning resource and the socioeconomic background score of the teaching object in the candidate teaching strategy by the standard deviation of the resource score and the standard deviation of the socioeconomic background score of the teaching object to obtain a correlation probability; and perform a weighted summation of the recommendation probability and the correlation probability to obtain a fairness risk value.

[0127] In some embodiments of this disclosure, the report generation unit 404 is configured to: in response to detecting that the teaching strategy deviation risk is obtained through the fairness risk value, use the Shapley additive interpretation algorithm to predict the contribution of features in the candidate teaching strategy, and obtain a deviation risk root source report including contribution information of each feature; in response to detecting that the teaching strategy deviation risk is obtained through the pedagogy risk value, use the counterfactual interpretation algorithm to predict the cognitive optimization strategy of the candidate teaching strategy, and obtain a deviation risk root source report including the cognitive optimization strategy.

[0128] In some embodiments of this disclosure, the sending unit 405 is configured to: receive teaching intervention decision information sent by an educational terminal; detect whether the teaching intervention decision information is a cognitive optimization strategy in the root cause report of deviation risk; in response to detecting that the teaching intervention decision information is a cognitive optimization strategy, use the cognitive optimization strategy as the target teaching strategy and send the target teaching strategy to the students; in response to detecting that the teaching intervention decision information is not a cognitive optimization strategy, fine-tune the strategy generation model based on the teaching intervention decision information, generate the target teaching strategy using the fine-tuned strategy generation model, and send the target teaching strategy to the students.

[0129] In some embodiments of this disclosure, the sending unit 405 is further configured to: in response to detecting that the teaching intervention decision information is not a cognitive optimization strategy, add the teaching intervention decision information to the preference dataset of the educational terminal; based on the preference dataset, update the internal parameters of the strategy model through a reinforcement learning algorithm, wherein the objective function of the strategy model is used to maximize the score of the reward model, while ensuring that the deviation between the fine-tuned teaching strategy and the teaching strategy before fine-tuning is within a preset range; and the reward model scores the teaching strategy based on the preference dataset.

[0130] The teaching strategy processing apparatus provided in the embodiments of this disclosure firstly receives learning information of the students; secondly, a strategy generation unit 402 generates candidate teaching strategies based on the student profile model; thirdly, a detection unit 403 detects whether the candidate teaching strategies have a risk of deviation based on the learning information and / or the strategy information of the candidate teaching strategies; then, a report generation unit 404 generates and sends a root cause report of the deviation risk of the candidate teaching strategies to the education terminal in response to the detection of a risk of deviation in the candidate teaching strategies; finally, a sending unit 405 receives teaching intervention decision information sent by the education terminal, and generates and sends a target teaching strategy to the students based on the teaching intervention decision information. Thus, through a closed loop of "receiving-generating-detecting-reporting-intervention," teaching strategies are corrected in real time, enabling teachers to intervene precisely and students to obtain personalized, unbiased, and efficient learning paths.

[0131] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0132] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0133] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their patterns are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0134] like Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0135] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0136] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the teaching strategy processing method. For example, in some embodiments, the teaching strategy processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the teaching strategy processing method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the teaching strategy processing method by any other suitable means (e.g., by means of firmware).

[0137] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0138] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable instructional processing device, such that when executed by the processor or controller, the patterns / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0139] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0141] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0142] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0143] The foregoing description of specific exemplary embodiments of this disclosure is for illustrative and explanatory purposes. These descriptions are not intended to limit this disclosure to the precise forms disclosed, and it will be apparent that many changes and variations can be made in accordance with the foregoing teachings. The exemplary embodiments were chosen and described in order to explain the specific principles of this disclosure and their practical application, thereby enabling those skilled in the art to implement and utilize various different exemplary embodiments of this disclosure, as well as various different choices and variations. The scope of this disclosure is intended to be defined by the claims and their equivalents.

Claims

1. A method for handling teaching strategies, the method comprising: Receive learning information and profile models of teaching students; Based on the profiling model, candidate teaching strategies are generated; Detecting whether a candidate teaching strategy has a teaching strategy bias risk based on the student learning information and / or the strategy information of the candidate teaching strategy includes: calculating the recommendation probability of recommending the candidate teaching strategy to a specific gender group based on the student learning information and the strategy information, and / or the correlation probability between the scores of learning resources in the candidate teaching strategy and the socioeconomic background quantitative score of the student; obtaining a fairness risk value based on the recommendation probability and / or the correlation probability; calculating the occurrence probability of knowledge points in the candidate teaching strategy and the difficulty value of the candidate teaching strategy based on the strategy information; obtaining a pedagogical risk value based on the occurrence probability and the difficulty value; calculating the length of a consecutive failure sequence of the student's historical answer sequence corresponding to the candidate teaching strategy based on the student learning information and the strategy information; and detecting whether the candidate teaching strategy has a teaching strategy bias risk based on the fairness risk value, the pedagogical risk value, and the length of the consecutive failure sequence. In response to the detection that the candidate teaching strategy has a teaching strategy deviation risk, generating and sending a deviation risk root cause report of the candidate teaching strategy to the education terminal includes: in response to the detection that the teaching strategy deviation risk is obtained through fairness risk value, using the Shapley additive interpretation algorithm to predict the contribution of features in the candidate teaching strategy, obtaining a deviation risk root cause report including contribution information of each feature; in response to the detection that the teaching strategy deviation risk is obtained through pedagogical risk value, using the counterfactual interpretation algorithm to predict the cognitive optimization strategy of the candidate teaching strategy, obtaining a deviation risk root cause report including the cognitive optimization strategy; The system receives teaching intervention decision information sent by the educational terminal, and generates and sends target teaching strategies to the teaching subjects based on the teaching intervention decision information.

2. The method according to claim 1, wherein, The method further includes: In response to the detection that the candidate teaching strategy does not have the risk of teaching strategy deviation, the candidate teaching strategy is selected as the target teaching strategy and sent to the teaching object.

3. The method according to claim 1, wherein, Based on the learning information and the strategy information, the probability of recommending the candidate teaching strategy to a specific gender group is calculated, and / or the correlation probability between the scores of learning resources in the candidate teaching strategy and the socioeconomic background quantitative score of the teaching object is calculated. Based on the recommended probability and / or the relevant probability, the fairness risk value is obtained as follows: Based on the learning information, determine the first number of each gender group and all strategy identifiers for each gender group in a specific gender group; Based on the strategy information, the target strategy identifier of the candidate teaching strategy is determined, and a second number of strategy identifiers belonging to the target strategy identifier are selected from all strategy identifiers; The recommendation probability is obtained by subtracting the second number from the first number for each gender group. And / or divide the covariance between the resource score of the learning resources in the candidate teaching strategy and the quantitative score of the socioeconomic background of the teaching object by the standard deviation of the resource score and the standard deviation of the quantitative score of the socioeconomic background of the teaching object to obtain the relevant probability; The fairness risk value is obtained by weighting and summing the recommendation probability and the relevant probability.

4. The method according to claim 1 or 2, wherein, The step of receiving teaching intervention decision information sent by the educational terminal, and generating and sending target teaching strategies to the teaching subjects based on the teaching intervention decision information, includes: Receive teaching intervention decision information sent by the educational terminal; The test is to determine whether the teaching intervention decision information is based on the cognitive optimization strategies in the deviation risk root cause report; In response to detecting that the teaching intervention decision information is to use the cognitive optimization strategy, the cognitive optimization strategy is used as the target teaching strategy, and the target teaching strategy is sent to the teaching object; In response to the detection that the teaching intervention decision information does not follow the cognitive optimization strategy, the strategy generation model is fine-tuned based on the teaching intervention decision information. The target teaching strategy is generated using the fine-tuned strategy generation model and sent to the teaching object.

5. The method according to claim 4, wherein, The step of responding to the detection that the teaching intervention decision information indicates a rejection of the cognitive optimization strategy, fine-tuning the strategy generation model based on the teaching intervention decision information, generating a target teaching strategy using the fine-tuned strategy generation model, and sending the target teaching strategy to the learner includes: In response to detecting that the teaching intervention decision information is not to follow the cognitive optimization strategy, the teaching intervention decision information is added to the preference dataset of the educational terminal; Based on the preference dataset, the internal parameters of the policy generation model are updated using a reinforcement learning algorithm. The objective function of the policy generation model is used to maximize the score of the reward model, while ensuring that the deviation between the fine-tuned teaching strategy and the original teaching strategy is within a preset range. The reward model scores the teaching strategy based on the preference dataset.

6. A teaching strategy processing device, the device comprising: The receiving unit is configured to receive learning information and profile models of the teaching subjects; The strategy generation unit is configured to generate candidate teaching strategies based on the profile model; The detection unit is configured to detect whether the candidate teaching strategy has a teaching strategy deviation risk based on the learning information and / or the strategy information of the candidate teaching strategy; the detection unit is further configured to: calculate the recommendation probability of recommending the candidate teaching strategy to a specific gender group based on the learning information and the strategy information, and / or the correlation probability between the score of the learning resources in the candidate teaching strategy and the socioeconomic background quantitative score of the teaching object; Based on the recommended probability and / or the relevant probability, a fairness risk value is obtained; Based on the strategy information, calculate the probability of occurrence of knowledge points in the candidate teaching strategies and the difficulty value of the candidate teaching strategies; Based on the probability of occurrence and the difficulty value, the teaching method risk value is obtained; Based on the learning information and the strategy information, calculate the length of the consecutive failure sequence of the historical answer sequence corresponding to the candidate teaching strategy for the teaching object; Based on the fairness risk value, the teaching methodology risk value, and the length of the consecutive failure sequence, the candidate teaching strategy is detected to have a teaching strategy bias risk. The report generation unit is configured to generate and send a root cause report of the deviation risk of the candidate teaching strategy to the education terminal in response to detecting that the candidate teaching strategy has a teaching strategy deviation risk; the report generation unit is further configured to: in response to detecting that the teaching strategy deviation risk is obtained through fairness risk value, use the Shapley additive interpretation algorithm to predict the contribution of the characteristics in the candidate teaching strategy, and obtain a root cause report of the deviation risk including the contribution information of each feature; In response to the detection of the deviation risk of the teaching strategy obtained through the teaching methodology risk value, the counterfactual interpretation algorithm is used to predict the cognitive optimization strategy of the candidate teaching strategy, and a deviation risk root cause report including the cognitive optimization strategy is obtained. The sending unit is configured to receive teaching intervention decision information sent by the education terminal, and based on the teaching intervention decision information, generate and send the target teaching strategy to the teaching object.

7. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method according to any one of claims 1-5.

8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.