A method, apparatus, equipment, and medium for generating online teaching strategies.
By acquiring and modeling user learning behavior sequences, the system generates optimal teaching strategy parameters, solving the problems of inflexible adjustment and high cost of personalized design in existing gamified teaching. This achieves precise matching of personalized teaching strategies, enhancing user learning interest and task persistence.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU HAILIANG DIGITAL TECH CO LTD
- Filing Date
- 2026-03-30
- Publication Date
- 2026-06-26
AI Technical Summary
Existing gamified teaching technologies cannot be flexibly adjusted according to the actual situation in the teaching process, cannot accurately perceive changes in the user's learning status, resulting in significant differences in incentive effects, easily leading to reward fatigue, and the cost of personalized design is high, making it difficult to meet the long-term personalized teaching needs.
By acquiring user learning behavior sequences, performing feature collection and modeling, and generating optimal teaching strategy parameters, including learning behavior feature vectors and state modeling, the teaching strategy is dynamically adjusted to match the user's real-time learning state.
It achieves precise matching of personalized teaching strategies, avoids reward fatigue, enhances users' learning interest and task persistence, reduces the cost of personalized design, and provides implementable and cyclical teaching support.
Smart Images

Figure CN122288937A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent education technology, and more specifically, to an online teaching strategy generation method, apparatus, device, and medium. Background Technology
[0002] In various educational scenarios, including smart learning terminals, online education, and classroom teaching, gamified learning has become an important means to enhance user learning interest, participation, and task completion rates, and is widely used in various user-generated teaching products and activities. In existing technologies, common implementations of gamified learning mainly involve integrating various gamified elements into the teaching process. These elements guide user participation through fixed game rules and incentive mechanisms, including setting up point rewards, badge unlocking, level challenges, leaderboard competitions, and animated feedback. This transforms abstract learning tasks into engaging gamified tasks, attempting to compensate for the monotony of traditional teaching and thus motivate users to learn. However, existing gamified learning systems often rely on human experience to pre-set game rules and teaching mechanisms, embedding game elements as independent display or incentive components into the teaching process. Core parameters such as reward trigger frequency, task difficulty, and feedback format are determined manually and then applied to all users' learning processes, providing basic gamified learning support.
[0003] However, existing gamified teaching technologies and systems have many shortcomings in practical applications, making it difficult to meet the long-term and effective personalized teaching needs of users in educational scenarios. Specific deficiencies include: First, game rules and reward mechanisms are mostly fixed configurations preset by humans, lacking dynamic adaptability and unable to be flexibly adjusted according to the actual situation during the teaching process; second, they fail to accurately perceive and respond to the learning status, motivational changes, and behavioral differences of different users, adopting a one-size-fits-all teaching model, resulting in significant differences in the incentive effect of the same gamification mechanism on different users; third, long-term use of fixed game rules can easily lead to reward fatigue among users, resulting in a gradual decline in learning participation and failing to achieve long-term incentive effects; fourth, some users may even experience negative incentive effects under incentive mechanisms that do not match their learning status, which is detrimental to improving learning outcomes; fifth, game elements are usually only used as display or incentive components, without deep involvement in teaching decision modeling, failing to provide effective support for personalized teaching; sixth, personalized gamification design requires a large amount of manual configuration, which is time-consuming and labor-intensive, resulting in high costs for large-scale application. Therefore, there is an urgent need for an intelligent teaching technology solution that can solve the above-mentioned deficiencies. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of the prior art by providing an online teaching strategy generation method, apparatus, device, and medium. This allows for the precise capture of user behavior feedback and status changes during the learning process, matching the optimal teaching parameters from a preset strategy space, and quickly implementing them. This avoids reward fatigue and engagement decline caused by fixed rules, while enhancing user learning interest through personalized strategy adjustments.
[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, embodiments of this application provide an online teaching strategy generation method, including: Obtain the user's learning behavior sequence under the current teaching strategy parameters in the online learning scenario; Feature acquisition is performed on the learning behavior sequence to obtain the user's learning behavior feature vector; Based on the learning behavior sequence and the learning behavior feature vector, the user's learning state is modeled to obtain the user's learning state, which includes: an instantaneous learning state vector; Based on the real-time learning state vector, the target teaching strategy parameters are determined from the preset strategy space; Update the current teaching strategy parameters of the online learning scenario to the target teaching strategy parameters.
[0006] In an optional implementation, the step of modeling the user's learning state based on the learning behavior sequence and the learning behavior feature vector to obtain the user's learning state includes: Based on the learning behavior sequence, determine the user's initial learning state vector; Based on the learning behavior feature vector, a preset mapping function is used to determine the user's learning state increment; Based on the learning state increment, the initial learning state vector is smoothed to obtain the instantaneous learning state vector.
[0007] In an optional implementation, determining the user's learning state increment based on the learning behavior feature vector using a preset mapping function includes: Based on the preset weight matrix and the learning behavior feature vector, the learning state increment is obtained using a preset mapping function.
[0008] In an optional implementation, the method further includes: The preset weight matrix is obtained using the user's historical learning data.
[0009] In an optional implementation, determining the target teaching strategy parameters from a preset strategy space based on the real-time learning state vector includes: Based on the real-time learning state vector, the user's target learning state category is determined from the preset learning state categories; Based on the target learning state category, the target teaching strategy parameters corresponding to the target learning state category are determined from multiple sets of teaching strategy parameters in the preset strategy space.
[0010] In an optional implementation, the method further includes: Obtain learning effectiveness indicators for the teaching strategy parameters of each group within multiple learning cycles; Based on the learning effectiveness indicators of the teaching strategy parameters described in each group, the priority of using the teaching strategy parameters in multiple groups is adjusted.
[0011] In an optional implementation, the method further includes: The user's learning performance is evaluated under the current teaching strategy parameters and the target teaching strategy parameters respectively, to obtain the user's current learning performance index and target learning performance index; Based on the current learning outcome indicators and the target learning outcome indicators, the feasibility of the target teaching strategy parameters is determined.
[0012] Secondly, embodiments of this application also provide an online teaching strategy generation device, the device comprising: The acquisition module is used to acquire the user's learning behavior sequence under the current teaching strategy parameters in the online learning scenario; The acquisition module is used to acquire features from the learning behavior sequence to obtain the user's learning behavior feature vector; The modeling module is used to model the user's learning state based on the learning behavior sequence and the learning behavior feature vector to obtain the user's learning state, which includes: an instantaneous learning state vector; The determination module is used to determine the target teaching strategy parameters from the preset strategy space based on the real-time learning state vector; The update module is used to update the current teaching strategy parameters of the online learning scenario to the target teaching strategy parameters.
[0013] Thirdly, embodiments of this application also provide a computer device, including: a processor, a storage medium, and a bus, wherein the storage medium stores program instructions executable by the processor, and when the computer device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the online teaching strategy generation method as described in any of the first aspects.
[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the online teaching strategy generation method as described in any of the first aspects.
[0015] The beneficial effects of this application are: This application provides an online teaching strategy generation method, apparatus, device, and medium. The method includes: acquiring a user's learning behavior sequence under current teaching strategy parameters in an online learning scenario; collecting features from the learning behavior sequence to obtain a user's learning behavior feature vector; modeling the user's learning state based on the learning behavior sequence and the learning behavior feature vector to obtain the user's learning state, which includes an immediate learning state vector; determining target teaching strategy parameters from a preset strategy space based on the immediate learning state vector; and updating the current teaching strategy parameters of the online learning scenario to the target teaching strategy parameters. This method can accurately capture user behavior feedback and state changes during the learning process, match the optimally adapted teaching parameters from the preset strategy space, and quickly implement them. This avoids reward fatigue and engagement decay caused by fixed rules, while improving user learning interest, task persistence, and completion quality through personalized strategy adjustments. Simultaneously, it accumulates real and effective behavioral and effect data for the long-term evolution of subsequent strategies, providing a feasible and cyclical execution path for personalized teaching support in intelligent education systems. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is one of the flowcharts illustrating an online teaching strategy generation method provided in an embodiment of this application; Figure 2 A second schematic flowchart illustrating an online teaching strategy generation method provided in this application embodiment; Figure 3 A third flowchart illustrating an online teaching strategy generation method provided in this application embodiment; Figure 4 A flowchart illustrating an online teaching strategy generation method provided in this application embodiment is shown in Figure 4. Figure 5The fifth flowchart illustrates an online teaching strategy generation method provided in this application embodiment; Figure 6 A schematic diagram of the functional modules of an online teaching strategy generation device provided in an embodiment of this application; Figure 7 This is a schematic diagram of a computer device provided in an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0019] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0020] In the description of this application, it should be noted that if the terms "upper", "lower", etc. appear to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship that the product of this application is usually placed in, it is only for the convenience of describing this application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.
[0021] Furthermore, the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Additionally, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0022] It should be noted that, where there is no conflict, the features in the embodiments of this application can be combined with each other.
[0023] The online teaching strategy generation method provided in this application will be explained in detail below with reference to the accompanying drawings and specific examples. The online teaching strategy generation method provided in this application can also be implemented by a computer device by running algorithms or software. The computer device can be, for example, a server or a terminal, and the terminal can be a user's computer. Figure 1 This is one of the flowcharts illustrating an online teaching strategy generation method provided in an embodiment of this application; such as Figure 1 As shown, the method includes: S101. Obtain the user's learning behavior sequence under the current teaching strategy parameters in the online learning scenario.
[0024] In this embodiment, the online learning scenario provides users with an online learning environment, enabling them to learn and answer questions online, such as in online education for children, smart learning terminals, and classroom teaching. The online learning scenario employs gamified teaching strategies to enhance users' learning interest, participation, and task completion rates. These gamified strategies reward users upon completion of learning levels, and these rewards may include points, badges, level completion systems, leaderboards, and animated feedback.
[0025] For example, taking an online math challenge learning platform for children as an example, the core learning content of this platform covers mixed arithmetic operations and simple word problems. The teaching format adopts a gamified mode with level progression, point rewards, and multiple forms of feedback. The platform's data collection module is embedded in the learning page and the question-answering interaction component to capture the user's entire learning behavior in real time under the current teaching strategy parameters. The initial default current teaching strategy parameters are: question difficulty level 2 (corresponding to two-digit addition and subtraction mixed operations), one reward triggered for every 3 questions completed, a prompt triggered for 2 consecutive errors, 5 points deducted for 3 consecutive errors, and feedback in the form of static text, such as "You answered correctly."
[0026] Taking a user's learning session as an example, the complete sequence of behaviors is recorded in chronological order: At 10:02:03, the user enters the answer interface for question 1; at 10:02:15, the user submits answer 35, which takes 12 seconds and is correct; at 10:02:20, the user enters the answer interface for question 2; at 10:02:45, the user submits incorrect answer 18, which takes 25 seconds; then at 10:02:50, the user chooses to retry; at 10:03:10, the user submits correct answer 28, which takes 20 seconds; subsequently, the user completes questions 3 through 6 in sequence, and finally exits to answer question 7 at 10:08:30. A total of 6 questions are completed in this learning session, with 4 questions remaining unanswered. All the above behaviors are organized into a complete learning behavior sequence in chronological order.
[0027] S102. Collect features from the learning behavior sequence to obtain the user's learning behavior feature vector.
[0028] The feature collection is based on a 10-minute sliding time window for the current round of learning, covering the complete answering cycle of 10 questions. Core behavioral features are extracted from the collected learning behavior sequences and quantified. The learning behavior features include: average task completion time b1, change in accuracy b2, number of active attempts b3, abandonment / retry ratio b4, and behavioral change before and after reward triggering b5. The learning behavior feature vector is represented as: B_t=[b1,b2,b3,b4,b5].
[0029] Specifically, the average task completion time is calculated by dividing the total time spent answering a single question (including retry) by the number of questions completed. In the above example, this is (12+25+20+...+30) seconds divided by 6 questions, resulting in 22.5 seconds per question. The change in accuracy is calculated by subtracting the accuracy rate of the first half from the accuracy rate of the second half, i.e., (2 / 3) - (3 / 3) = -0.33. The number of active attempts is calculated by counting the number of retry attempts when no prompt is triggered and the number of attempts to skip simple questions. Number of retryes per question: Here, the user only actively retried question 2 once and did not skip any questions, so it is 1 time; Abandon / Retry Ratio: This is the number of unfinished questions when learning was interrupted divided by the total number of retries, i.e., 4 unanswered questions divided by 1 retry, the result is 4.0; Change in behavior before and after reward trigger: This is the time taken to answer the first question after the reward trigger minus the time taken to answer the last question before the reward trigger, i.e., the time taken for question 3 (after the reward) is 18 seconds minus the time taken for question 2 (before the reward) is 20 seconds, the result is -2 seconds. The above 5 features are combined in a fixed order to form the learning behavior feature vector B_t=[22.5,-0.33,1.0,4.0,-2.0].
[0030] S103. Based on the learning behavior sequence and learning behavior feature vector, model the user's learning state to obtain the user's learning state.
[0031] The learning state includes: the instantaneous learning state vector.
[0032] Specifically, based on the learning behavior sequence and learning behavior feature vector, the user's learning state is modeled to determine the user's learning state, which also includes the learning state increment and the learning state stability score.
[0033] S104. Determine the target teaching strategy parameters from the preset strategy space based on the real-time learning state vector.
[0034] The preset strategy space stores multiple sets of teaching strategy parameters, which include strategy parameters in five dimensions: question difficulty level, reward trigger frequency, prompt trigger conditions, failure penalty parameters, and feedback presentation method. Based on the real-time learning state vector, the corresponding target teaching strategy parameters can be determined from the preset strategy space.
[0035] S105. Update the current teaching strategy parameters of the online learning scenario to the target teaching strategy parameters.
[0036] Specifically, the platform's strategy execution module is responsible for updating and implementing parameters. Updates are scheduled before the user begins their next learning round or when they re-enter the platform after an interruption in the current round. First, the strategy module receives the target parameters and stores them in the user's dedicated strategy configuration library, overwriting the original initial parameters. When the user re-enters the math challenge page, the system automatically loads the target teaching strategy parameters. For example: the question generation module generates questions at difficulty level 1; the reward module triggers a reward every 3 questions completed (e.g., a points reward animation pops up when questions 3, 6, and 9 are completed); the prompt module immediately displays a prompt, such as "Pay attention to the order of operations," after the user makes one consecutive mistake; the penalty module disables the points deduction function; and the feedback module calls animation and voice components, playing a "Great job!" voice message and star animation when the user answers correctly, and a "Think again!" voice message and encouraging animation when the user answers incorrectly. Simultaneously, a detailed parameter update log is recorded, including the user ID, update time, original parameters, target parameters, and trigger reason (frustration risk state), facilitating subsequent tracking and optimization.
[0037] In summary, this application provides an online teaching strategy generation method. The method includes: acquiring a user's learning behavior sequence under current teaching strategy parameters in an online learning scenario; collecting features from the learning behavior sequence to obtain a user's learning behavior feature vector; modeling the user's learning state based on the learning behavior sequence and the learning behavior feature vector to obtain the user's learning state, which includes an immediate learning state vector; determining target teaching strategy parameters from a preset strategy space based on the immediate learning state vector; and updating the current teaching strategy parameters of the online learning scenario to the target teaching strategy parameters. This method can accurately capture user behavior feedback and state changes during the learning process, match the optimally suited teaching parameters from the preset strategy space, and quickly implement them. This avoids reward fatigue and engagement decay caused by fixed rules, while improving user learning interest, task persistence, and completion quality through personalized strategy adjustments. Simultaneously, it accumulates real and effective behavioral and effect data for the long-term evolution of subsequent strategies, providing a feasible and cyclical execution path for personalized teaching support in intelligent education systems.
[0038] This application also provides another possible implementation of the online teaching strategy generation method. Figure 2 This is a second flowchart illustrating an online teaching strategy generation method provided in an embodiment of this application. Figure 2 As shown, based on the learning behavior sequence and learning behavior feature vector, the user's learning state is modeled to obtain the user's learning state, including: S201. Determine the user's initial learning state vector based on the learning behavior sequence.
[0039] In this embodiment, the learning state includes multiple dimensions such as learning interest level, task challenge tolerance, reward sensitivity, frustration recovery ability, and learning persistence tendency. The user's initial learning state vector is determined based on the learning behavior sequence at the current moment and historical moments.
[0040] For example, to obtain the user's behavioral sequence of three rounds of learning, the core behavioral data of each round of learning is collected first: Round 1: completion rate 90%, accuracy rate at level 2 difficulty 75%, speed improvement rate after reward 35%, retry rate for consecutive errors 40%, average completion of 9 questions; Round 2: completion rate 80%, accuracy rate at level 2 difficulty 65%, speed improvement rate after reward 28%, retry rate for consecutive errors 55%, average completion of 8 questions; Round 3: completion rate 70%, accuracy rate at level 2 difficulty 70%, speed improvement rate after reward 27%, retry rate for consecutive errors 55%, average completion of 7 questions.
[0041] The above data is then standardized and mapped to the interval [0,1]: the learning interest level is the average completion rate over multiple rounds, i.e., (90%+80%+70%) / 3=80%, corresponding to 0.8; the task challenge tolerance is the average accuracy rate over multiple rounds at difficulty level 2, i.e., (75%+65%+70%) / 3=70%, corresponding to 0.7; the reward sensitivity is the average speed improvement rate after reward, i.e., (35%+28%+27%) / 3=30%, corresponding to 0.3; the frustration recovery ability is the average retry rate of consecutive errors, i.e., (40%+55%+55%) / 3=50%, corresponding to 0.5; the learning persistence tendency is the ratio of the average number of completed questions to the total number of questions in a single round, i.e., (9+8+7) / 30=80%, corresponding to 0.8. Finally, the initial learning state vector S0=[0.8,0.7,0.3,0.5,0.8] is formed.
[0042] S202. Based on the learning behavior feature vector, a preset mapping function is used to determine the user's learning state increment.
[0043] Optionally, the learning state increment can be obtained by using a preset mapping function based on a preset weight matrix and the learning behavior feature vector.
[0044] Specifically, the preset mapping function is the linear mapping function ΔS_t=W. B_t, where W is a preset weight matrix, is used to linearly map the learning behavior feature vector B_t to obtain the learning state increment ΔS_t.
[0045] Optionally, a preset weight matrix can be obtained using the user's historical learning data.
[0046] Specifically, the process of obtaining the weight matrix is as follows: the data source is the historical learning data of multiple users on the platform, which includes multiple rounds of learning behavior and corresponding manually labeled state changes; a linear regression model is used for training, with the behavior feature vector B as input and the manually labeled state change ΔS as output, to fit and obtain the weight matrix W.
[0047] S203. Based on the learning state increment, smooth the initial learning state vector to obtain the instantaneous learning state vector.
[0048] The core of smoothing is determining the smoothing coefficient α and performing smoothing calculations. The value of α is related to the stability of the strategy, specifically: Strategy stability score = number of consecutive rounds the strategy was used × average performance satisfaction (1-5 points). For example, when the stability score is ≥10, α=0.8 (high stability, weakening incremental impact); when 5≤stability score<10, α=0.6 (medium stability); and when the stability score<5, α=0.4 (low stability, strengthening incremental impact). In this scenario, if the initial strategy is used consecutively for 3 rounds, the average performance satisfaction = 3 points, and the stability score = 3×3=9, therefore α=0.6 is chosen.
[0049] Using the smoothing formula S_t=α S0+(1-α) The calculation is performed by substituting the initial learning state vector S0, the state increment vector ΔS_t, and the smoothing coefficient α, to calculate the learning state in each dimension, and finally obtain the smoothed real-time learning state vector.
[0050] The method provided in this application, through a refined learning state modeling process including initial state initialization, state increment calculation, and smoothing, effectively solves the problem of traditional teaching systems' difficulty in accurately depicting users' dynamic learning states. Its advantages lie in: determining the initial state based on historical learning behavior sequences ensures fundamental modeling accuracy; converting behavioral features into state increments through a preset mapping function and weight matrix achieves a quantitative correlation between behavior and state; and the weighted smoothing mechanism filters out interference from short-term behavioral noise on state judgment, improving the stability and interpretability of the real-time learning state vector. The final output of an accurate real-time learning state vector provides a core basis for the precise matching of subsequent teaching strategies, ensuring that strategy adjustments align with the user's actual learning state and individual differences, avoiding negative incentive effects caused by blind adjustments.
[0051] This application also provides another possible implementation of the online teaching strategy generation method. Figure 3 This is a third flowchart illustrating an online teaching strategy generation method provided in an embodiment of this application. Figure 3As shown, based on the real-time learning state vector, the target teaching strategy parameters are determined from the preset strategy space, including: S301. Based on the real-time learning state vector, determine the user's target learning state category from the preset learning state categories.
[0052] In this embodiment, the preset learning state categories include task adaptation state, frustration risk state, reward sensitive state, and normal state. The judgment rules for each category combine the core dimension and auxiliary dimension thresholds: the task adaptation state requires that the change in the accuracy of the core dimension be greater than or equal to 0.85 and the average task completion time be less than or equal to the preset time threshold; the frustration risk state requires that the number of active attempts in the core dimension be greater than or equal to 3 and the abandonment / retry ratio be greater than or equal to 2; the reward sensitive state requires that the change in behavior before and after the reward is triggered in the core dimension be less than or equal to -20%; the normal state does not meet any of the above category conditions.
[0053] Then, by combining the real-time learning state vector and the corresponding learning behavior feature vector, the corresponding target learning state category is determined.
[0054] S302. Based on the target learning state category, determine the target teaching strategy parameters corresponding to the target learning state category from multiple sets of teaching strategy parameters in the preset strategy space.
[0055] In the preset strategy space, the parameter combination designed for the failure risk state is to reduce the difficulty level of the question by one level, lower the prompt trigger threshold so that the prompt appears earlier, and not deduct points after failure; the parameter combination designed for the task adaptation state is to increase the difficulty level of the question by one level, reduce the reward trigger frequency, and keep the feedback method unchanged; the parameter combination designed for the reward sensitive state is to increase the reward trigger frequency and enable animation and voice feedback.
[0056] Then, based on the target learning state category, the target teaching strategy parameters corresponding to the target learning state category are determined.
[0057] The method provided in this application achieves precise alignment between learning states and teaching strategies through state category determination and strategy parameter matching logic, effectively improving the pertinence and effectiveness of strategy adjustments. Its value lies in the pre-defined clear rules for determining learning state categories, allowing users' dynamic states to be clearly categorized and avoiding strategy adaptation bias caused by state ambiguity. This process overcomes the limitations of traditional manual strategy configuration. Through automated matching of states and strategies, it improves the efficiency of personalized teaching and ensures that strategy adjustments directly address users' core learning needs, thereby enhancing learning participation and effectiveness.
[0058] This application also provides another possible implementation of the online teaching strategy generation method. Figure 4This is a fourth flowchart illustrating an online teaching strategy generation method provided in an embodiment of this application. Figure 4 As shown, the method also includes: S401. Obtain learning effectiveness indicators for each group of teaching strategy parameters within multiple learning cycles.
[0059] S402. Adjust the priority of using multiple sets of teaching strategy parameters based on the learning effect indicators of each set of teaching strategy parameters.
[0060] In this embodiment, a learning cycle is defined as one learning cycle containing three rounds of learning tasks (10 questions per round), with a long-term statistical period of 7 days, covering four complete learning cycles for a total of 12 rounds. The statistical learning effectiveness indicators include single learning session duration, number of consecutive completed levels, learning interruption rate, accuracy improvement, and willingness to re-participate. The definitions and statistical methods for each indicator are as follows: Single learning session duration is the total duration of each round of learning from start to finish (or interruption), calculated as the average of 12 rounds; Number of consecutive completed levels is the number of levels passed consecutively without interruption, calculated as the average of 12 rounds; Learning interruption rate is the percentage calculated by dividing the number of interrupted rounds by the total number of rounds; Accuracy improvement is the absolute difference between the final accuracy rate and the initial accuracy rate; Willingness to re-participate is the percentage calculated by dividing the number of times actively re-entering the learning environment by the total number of available attempts.
[0061] The statistics cover multiple sets of strategy parameters used by users within 7 days. Examples: Strategy parameter A corresponds to an average single learning session duration of 8.5 minutes, an average number of consecutive levels of 6.2, a learning interruption rate of 41.7% (interruptions in 5 out of 12 rounds), an accuracy improvement of 5%, and a repeat participation intention of 66.7% (active participation using 8 out of 12 available attempts); Strategy parameter B corresponds to an average single learning session duration of 10.2 minutes, an average number of consecutive levels of 7.8, a learning interruption rate of 16.7% (interruptions in 2 out of 12 rounds), an accuracy improvement of 12%, and a repeat participation intention of 83.3% (active participation using 10 out of 12 available attempts); Strategy parameter C corresponds to an average single learning session duration of 9.8 minutes, an average number of consecutive levels of 8.5, a learning interruption rate of 8.3% (interruptions in 1 out of 12 rounds), an accuracy improvement of 15%, and a repeat participation intention of 91.7% (active participation using 11 out of 12 available attempts).
[0062] Based on the learning effectiveness indicators of each group of teaching strategy parameters, the priority of using multiple groups of teaching strategy parameters is adjusted. Specifically, when the learning interruption rate corresponding to a certain parameter combination is consistently higher than a preset threshold, the priority of using that combination is automatically reduced; when the learning duration and number of consecutive levels corresponding to a certain parameter combination are significantly increased, the selection probability of that combination is increased, thereby realizing the adjustment of the priority of using multiple groups of teaching strategy parameters.
[0063] The method provided in this application achieves continuous evolution and optimization of teaching strategies through long-term learning effectiveness index statistics and strategy priority adjustment, solving the problem of traditional systems lacking long-term effectiveness evaluation and strategy iteration capabilities. Its core benefit lies in using effectiveness data from multiple learning cycles as a basis, avoiding judgment bias caused by short-term effects; and objectively evaluating and prioritizing multiple sets of strategy parameters through a weighted scoring method, which strengthens high-fitness strategies and eliminates low-effectiveness strategies, ensuring that teaching strategies always evolve in a direction that better aligns with children's long-term learning needs. This long-term optimization mechanism not only avoids reward fatigue and engagement decline caused by fixed strategies but also reduces the manual configuration cost of personalized teaching, providing intelligent education systems with sustainable optimization and scalable strategy iteration capabilities.
[0064] This application also provides another possible implementation of the online teaching strategy generation method. Figure 5 This is the fifth flowchart illustrating an online teaching strategy generation method provided in an embodiment of this application. Figure 5 As shown, the method also includes: S501. Evaluate the user's learning effectiveness under the current teaching strategy parameters and the target teaching strategy parameters respectively, and obtain the user's current learning effectiveness index and target learning effectiveness index.
[0065] In this embodiment, a before-and-after comparison design is used to evaluate the effect. The evaluation period is a single round of learning (10 questions). The same user uses the current teaching strategy parameters and the target teaching strategy parameters in two adjacent rounds of learning respectively. For example, the fourth round of learning uses the current teaching strategy parameter A0 (difficulty level 2, reward frequency once every 3 questions, prompt condition two consecutive errors, deduction of 5 points, static feedback), and the fifth round of learning uses the target teaching strategy parameter A3 (difficulty level 1, reward frequency once every 2 questions, prompt condition one consecutive error, no penalty, animated feedback).
[0066] The evaluation results show that in the fourth round of learning corresponding to the current parameter A0, the user completed 6 questions, with a response time of 7.2 minutes and an accuracy rate of 66.7% (4 / 6). There was an interruption (4 questions were not completed), and the user actively retried once. The feedback satisfaction was "No" (the "Like" button was not clicked). In the fifth round of learning corresponding to the target parameter A3, the user completed 9 questions, with a response time of 9.5 minutes and an accuracy rate of 88.9% (8 / 9). There were no interruptions, and the user actively retried 3 times. The feedback satisfaction was "Yes" (the "Like" button was clicked). By comparing the two sets of parameters, the core learning effect indicators and their changes are clearly obtained.
[0067] S502. Based on the current learning outcome indicators and the target learning outcome indicators, determine the feasibility of the target teaching strategy parameters.
[0068] Specifically, multi-dimensional feasibility assessment criteria were established: the pass standard for core effectiveness is ≥8 questions completed in a single round or an accuracy rate ≥80%, while the excellent standard is ≥9 questions completed in a single round and an accuracy rate ≥85%; the pass standard for sustainability is no interruptions or an interruption rate ≤10%, while the excellent standard is no interruptions and a response time ≥9 minutes (effective time); the pass standard for participation is ≥2 voluntary retry attempts or a "yes" feedback satisfaction rating, while the excellent standard is ≥3 voluntary retry attempts and a "yes" feedback satisfaction rating. Based on the evaluation results, the learning effect corresponding to target parameter A3 is as follows: 9 questions completed with an accuracy rate of 88.9%, meeting the excellent standard for core effectiveness; no interruptions and a response time of 9.5 minutes, meeting the excellent standard for sustainability; 3 voluntary retry attempts and a "yes" feedback satisfaction rating, meeting the excellent standard for participation. Furthermore, no negative indicators (such as excessively long response time, i.e., response time ≥15 minutes, decreased accuracy rate, etc.) appeared during the implementation of the target teaching strategy parameter, therefore, target teaching strategy parameter A3 is determined to be fully feasible. The following additional suggestion is given: retain this parameter combination as a long-term strategy and include it in the high-priority candidate set for subsequent evolutionary optimization, so as to continue to play its role in optimizing learning effects.
[0069] The method provided in this application, through a process of before-and-after comparative effect evaluation and feasibility determination, provides scientific verification and risk control for the implementation of the target teaching strategy, effectively improving the reliability and safety of personalized teaching. Its advantage lies in that, through comparative evaluation of adjacent rounds for the same user, it can accurately quantify the improvement (or difference) in effect between the target strategy and the current strategy, avoiding the one-sidedness of single-strategy evaluation; the multi-dimensional feasibility determination criteria not only focus on learning effectiveness but also take into account the user's learning experience and interest, ensuring that strategy adjustments do not cause negative problems (such as excessive fatigue or decreased interest). This process not only provides a decision-making basis for the implementation of the target strategy but also accumulates effect feedback data for subsequent strategy optimization, ensuring that the teaching strategy always aligns with the user's learning needs and continuously improves learning effectiveness and experience.
[0070] The following will continue to explain the online teaching strategy generation device and computer equipment provided in any of the above embodiments of this application. The specific implementation process and the resulting technical effects are the same as those in the corresponding method embodiments. For the sake of brevity, parts not mentioned in this embodiment can be referred to the corresponding content in the method embodiments.
[0071] Figure 6 This is a schematic diagram of the functional modules of an online teaching strategy generation device provided in an embodiment of this application. Figure 6 As shown, the online teaching strategy generation device 100 includes: The acquisition module 110 is used to acquire the user's learning behavior sequence under the current teaching strategy parameters in the online learning scenario; The acquisition module 120 is used to acquire features from the learning behavior sequence to obtain the user's learning behavior feature vector; Modeling module 130 is used to model the user's learning state based on the learning behavior sequence and the learning behavior feature vector to obtain the user's learning state, which includes: an instantaneous learning state vector. The determination module 140 is used to determine the target teaching strategy parameters from the preset strategy space based on the real-time learning state vector; The update module 150 is used to update the current teaching strategy parameters in the online learning scenario to the target teaching strategy parameters.
[0072] Optionally, the modeling module 130 is also used to determine the user's initial learning state vector based on the learning behavior sequence; determine the user's learning state increment based on the learning behavior feature vector using a preset mapping function; and smooth the initial learning state vector based on the learning state increment to obtain the instantaneous learning state vector.
[0073] Optionally, the determining module 140 is also used to obtain the learning state increment by using a preset mapping function based on the preset weight matrix and the learning behavior feature vector.
[0074] Optionally, the acquisition module 110 is also used to acquire a preset weight matrix using the user's historical learning data.
[0075] Optionally, the determining module 140 is further configured to determine the user's target learning state category from a preset learning state category based on the real-time learning state vector; and to determine the target teaching strategy parameters corresponding to the target learning state category from multiple sets of teaching strategy parameters in a preset strategy space based on the target learning state category.
[0076] Optionally, the device further includes: The acquisition module 110 is also used to acquire learning effect indicators of teaching strategy parameters for each group within multiple learning cycles; The adjustment module is used to adjust the priority of multiple sets of teaching strategy parameters based on the learning effect indicators of each set of teaching strategy parameters.
[0077] Optionally, the determining module 140 is also used to evaluate the user's learning effect under the current teaching strategy parameters and the target teaching strategy parameters respectively, to obtain the user's current learning effect index and the target learning effect index; and to determine the executability of the target teaching strategy parameters based on the current learning effect index and the target learning effect index.
[0078] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
[0079] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more microprocessors, or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).
[0080] Figure 7 This is a schematic diagram of a computer device provided in an embodiment of this application. This computer device can be used for generating online teaching strategies. Figure 7 As shown, the computer device includes: a processor 210, a storage medium 220, and a bus 230.
[0081] Storage medium 220 stores machine-readable instructions executable by processor 210. When the computer device is running, processor 210 communicates with storage medium 220 via bus 230, and processor 210 executes the machine-readable instructions to perform the steps of the above method embodiment. The specific implementation and technical effects are similar, and will not be described again here.
[0082] Optionally, this application also provides a storage medium 220, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the above-described method embodiments. The specific implementation and technical effects are similar, and will not be repeated here.
[0083] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0084] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0085] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0086] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0087] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for generating online teaching strategies, characterized in that, include: Obtain the user's learning behavior sequence under the current teaching strategy parameters in the online learning scenario; Feature acquisition is performed on the learning behavior sequence to obtain the user's learning behavior feature vector; Based on the learning behavior sequence and the learning behavior feature vector, the user's learning state is modeled to obtain the user's learning state, which includes: an instantaneous learning state vector; Based on the real-time learning state vector, the target teaching strategy parameters are determined from the preset strategy space; Update the current teaching strategy parameters of the online learning scenario to the target teaching strategy parameters.
2. The method according to claim 1, characterized in that, The step of modeling the user's learning state based on the learning behavior sequence and the learning behavior feature vector to obtain the user's learning state includes: Based on the learning behavior sequence, determine the user's initial learning state vector; Based on the learning behavior feature vector, a preset mapping function is used to determine the user's learning state increment; Based on the learning state increment, the initial learning state vector is smoothed to obtain the instantaneous learning state vector.
3. The method according to claim 2, characterized in that, The step of determining the user's learning state increment based on the learning behavior feature vector and using a preset mapping function includes: Based on the preset weight matrix and the learning behavior feature vector, the learning state increment is obtained using a preset mapping function.
4. The method according to claim 3, characterized in that, The method further includes: The preset weight matrix is obtained using the user's historical learning data.
5. The method according to claim 1, characterized in that, The step of determining the target teaching strategy parameters from the preset strategy space based on the real-time learning state vector includes: Based on the real-time learning state vector, the user's target learning state category is determined from the preset learning state categories; Based on the target learning state category, the target teaching strategy parameters corresponding to the target learning state category are determined from multiple sets of teaching strategy parameters in the preset strategy space.
6. The method according to claim 5, characterized in that, The method further includes: Obtain learning effectiveness indicators for the teaching strategy parameters of each group within multiple learning cycles; Based on the learning effectiveness indicators of the teaching strategy parameters described in each group, the priority of using the teaching strategy parameters in multiple groups is adjusted.
7. The method according to claim 1, characterized in that, The method further includes: The user's learning performance is evaluated under the current teaching strategy parameters and the target teaching strategy parameters respectively, to obtain the user's current learning performance index and target learning performance index; Based on the current learning outcome indicators and the target learning outcome indicators, the feasibility of the target teaching strategy parameters is determined.
8. An online teaching strategy generation device, characterized in that, The device includes: The acquisition module is used to acquire the user's learning behavior sequence under the current teaching strategy parameters in the online learning scenario; The acquisition module is used to acquire features from the learning behavior sequence to obtain the user's learning behavior feature vector; The modeling module is used to model the user's learning state based on the learning behavior sequence and the learning behavior feature vector to obtain the user's learning state, which includes: an instantaneous learning state vector; The determination module is used to determine the target teaching strategy parameters from the preset strategy space based on the real-time learning state vector; The update module is used to update the current teaching strategy parameters of the online learning scenario to the target teaching strategy parameters.
9. A computer device, characterized in that, include: The computer device includes a processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the computer device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the online teaching strategy generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, performs the steps of the online teaching strategy generation method as described in any one of claims 1 to 7.