A reinforcement learning course intervention method for learning motivation decline

By acquiring user learning behavior data to generate motivation status information, and intervening based on personalized incentive strategies, and making adaptive adjustments, the predictability and personalization of learning motivation decline in online learning platforms are solved, achieving efficient maintenance and optimization of learning motivation.

CN121010486BActive Publication Date: 2026-01-23HUNAN ANNA INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511541798.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-01-23
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing online learning platforms lack predictability and initiative when facing declining learning motivation, and general incentives cannot be personalized, resulting in poor intervention effects and potentially negative impacts on users.

Method used

By acquiring user learning behavior data, generating motivational state information, intervening based on personalized incentive strategies, and making adaptive adjustments based on feedback, a personalized and adaptive reinforcement learning curriculum intervention system is constructed.

Benefits of technology

It enables proactive intervention against declining learning motivation, improves the targeting and effectiveness of intervention, avoids waste of resources, ensures that the system can be continuously optimized and adapted to dynamic changes in users, and enhances learning willingness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010486B_ABST
    Figure CN121010486B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of reinforcement learning course intervention, and specifically discloses a reinforcement learning course intervention method for learning motivation recession, which comprises behavior data set acquisition, motivation state information generation, individualized motivation strategy generation, implementation effect feedback acquisition and motivation strategy adaptive adjustment; the present application generates motivation state information containing current motivation level and trend prediction by collecting and processing user learning behavior data in real time, and then generates individualized motivation strategy; after executing the strategy, implementation effect feedback is acquired by monitoring user response behavior, and the motivation analysis process and the motivation strategy generation process are dynamically optimized according to the feedback result by using an adaptive adjustment model; the present application realizes the transformation from passive remedy to active intervention, significantly improves the pertinence and acceptance of motivation measures, and exhibits strong robustness and intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of reinforcement learning curriculum intervention technology, and more specifically, to a reinforcement learning curriculum intervention method for addressing declining learning motivation. Background Technology

[0002] With the deep integration of artificial intelligence and educational technology, intelligent learning systems have become an important component of modern education. These systems aim to optimize the learning experience and outcomes through technological means. Learning motivation, as the core intrinsic factor driving learning behavior, is crucial for maintaining and stimulating learners' sustained engagement and success. Therefore, developing intelligent technologies that can effectively monitor and intervene in the decline of learning motivation is of great significance for improving the overall effectiveness of online education platforms.

[0003] Currently, existing online learning platforms typically employ relatively fixed technical solutions to address the problem of declining learning motivation. For example, some systems have built-in trigger mechanisms based on preset rules. When certain simple behavioral indicators of a user fall below a threshold, such as several consecutive days of inactivity or failing a quiz, they automatically push uniform reminder messages or encouraging remarks. Other platforms introduce generic gamification elements, such as points, badges, and leaderboard systems, attempting to maintain user engagement through external rewards. The configuration of these incentives is essentially the same for all users.

[0004] However, the aforementioned existing technical solutions have significant technical shortcomings in practical applications. First, interventions based on fixed rules are often reactive, only triggering when motivational decline has already produced noticeable behavioral consequences, lacking predictability and proactivity. Second, standardized incentive content ignores the vast differences in personality, preferences, and needs among learners, leading to varying intervention effects and potentially negatively impacting some users. Finally, the intervention logic of these systems is static, unable to self-adjust and optimize based on actual intervention results, and cannot learn from user interactions, thus failing to continuously improve their intervention effectiveness. Summary of the Invention

[0005] In view of this, in order to solve the problems mentioned in the background, a reinforcement learning curriculum intervention method for learning motivation decline is proposed.

[0006] The objective of this invention can be achieved through the following technical solution: This invention provides a reinforcement learning course intervention method for learning motivation decline, including the following steps: S1, obtaining behavioral data set: obtaining the user's learning behavior data and processing the learning behavior data to obtain a behavioral data set.

[0007] S2. Motivational State Information Generation: Perform motivational analysis on the behavioral data set to generate motivational state information that includes the current motivational level and predictions of motivational trends.

[0008] S3. Personalized Incentive Strategy Generation: Generate personalized incentive strategies based on motivational state information.

[0009] S4. Obtaining Implementation Results Feedback: Implement personalized incentive strategies and monitor user response behavior to obtain implementation results feedback.

[0010] S5. Adaptive Adjustment of Incentive Strategies: Based on feedback on implementation results, the motivation analysis process and the generation process of personalized incentive strategies are adaptively adjusted.

[0011] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: (1) By constructing motivational state information that includes motivational trend prediction, the present invention enables the intervention system not only to respond to the already occurred motivational decline, but also to provide early warning and intervene in potential motivational decline risks. This forward-looking intervention mechanism changes the previous delayed response mode, realizes the transformation from passive remediation to active prevention, and can stabilize and enhance learners' learning intentions earlier, preventing motivational problems from worsening.

[0012] (2) This invention dynamically prioritizes the incentive library by combining the user's historical incentive preferences with the current learning context, and finely adjusts the intensity and timing of incentive strategies based on the user's immediate motivation level. This deeply personalized design ensures that each intervention is highly compatible with the user's individual characteristics and situation, greatly improving the pertinence and acceptability of incentive measures, and avoiding the waste of resources and user resentment caused by one-size-fits-all interventions.

[0013] (3) This invention establishes a complete closed-loop learning system from strategy implementation, effect monitoring, quantitative feedback to adaptive adjustment of the model and strategy. This system can autonomously learn and accumulate experience from the success or failure of each intervention, continuously optimizing the accuracy of its motivation monitoring model and the effectiveness of its strategy generation rules. This self-evolutionary capability enables the system to maintain a high level of intervention over a long period of time and adapt to the dynamic changes in user motivation patterns, demonstrating strong robustness and intelligence. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1This is a schematic diagram of the method steps of the present invention.

[0016] Figure 2 This is a diagram of the intervention system architecture for learning motivation decline according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figure 1 and Figure 2 This invention provides a reinforcement learning course intervention method for learning motivation decline, including: S1, acquiring behavioral data set: acquiring users' learning behavior data and processing the learning behavior data to obtain a behavioral data set.

[0019] In a specific embodiment of the present invention, the steps of acquiring user learning behavior data and processing the learning behavior data to obtain a behavior data set include: collecting user interaction events on the learning platform in real time to obtain raw behavior data.

[0020] It should be noted that, in order to generate a structured behavioral data set for motivation analysis, this method first deploys front-end and back-end data acquisition agents on the learning platform to capture various user interaction events in real time, forming raw behavioral data. These interaction events include, but are not limited to, page browsing, video playback control such as play, pause, drag, content commenting, forum posting, quiz submission, etc. Each piece of raw behavioral data records the user identifier, event timestamp, event type, and specific event content.

[0021] The obtained raw behavioral data is filtered and denoised to generate clean behavioral data.

[0022] It should be noted that the system then initiates a filtering and denoising process to preprocess the raw behavioral data to generate clean behavioral data. This process first filters out non-learning-related data, such as administrator operation logs, system heartbeat packets, and suspicious bot access records. Next, it handles data anomalies, such as filling in missing key timestamps, correcting out-of-order events caused by network latency, and identifying and smoothing out extreme data points, such as excessively long learning session records caused by users forgetting to close their browsers.

[0023] The generated cleaning behavior data is then subjected to feature extraction and aggregation to form a behavior data set.

[0024] It's important to note that after obtaining the cleaning behavior data, the next step is feature extraction and aggregation. The core of this stage is transforming the discrete event stream into a quantifiable feature vector that characterizes the user's learning state. Within a preset time window, such as a week, the system aggregates and extracts multi-dimensional features from each user's cleaning behavior data. These features can be categorized into several aspects, including learning engagement, learning achievement, social connectivity, and behavioral consistency. For example, learning engagement features might include login frequency, average daily study time, and number of times key learning resources are accessed; learning achievement features might include average quiz scores and on-time homework completion rates; and social connectivity features might include the number of forum posts and received replies. To ensure numerical comparability of different features and eliminate the influence of unit of measurement, all extracted raw feature values ​​undergo min-max standardization, mapping them to a uniform interval. After these steps, a multi-dimensional feature vector is generated for each user within each time window. All these vectors constitute the final set of behavioral data used as input for the subsequent motivation monitoring model.

[0025] This invention, through the aforementioned refined data processing flow, transforms raw, fragmented, and noisy user interaction logs into a highly structured feature set with high information density that accurately reflects the user's learning state. This process not only provides a high-quality data foundation for subsequent motivational level assessment and trend prediction, significantly improving the accuracy and reliability of the motivation monitoring model, but also, through the construction of multi-dimensional features, enables the system to more comprehensively and profoundly understand the motivational changes behind user behavior, laying a solid technical foundation for truly personalized and adaptive incentive strategy interventions. Ultimately, it ensures the effectiveness of the entire intervention system, enabling it to take action based on accurate insights into the user's state, rather than on coarse or incomplete data.

[0026] S2. Motivational State Information Generation: Perform motivational analysis on the behavioral data set to generate motivational state information that includes the current motivational level and predictions of motivational trends.

[0027] In a specific embodiment of the present invention, the specific steps of performing motivational analysis on the behavioral data set to generate motivational state information including the current motivational level and motivational trend prediction include: inputting the behavioral data set into a motivational monitoring model for evaluating the motivational level, and calculating the current motivational level.

[0028] In a specific embodiment of the present invention, the motivation monitoring model for evaluating motivation level is obtained through the following steps: acquiring historical learning behavior data and its corresponding motivation labels, and constructing a training dataset.

[0029] It's important to note that to obtain a motivation monitoring model capable of accurately assessing motivation levels, a training dataset containing supervised information is first required. This process begins with collecting a large amount of anonymized historical learning behavior data and processing it into a series of standardized behavioral feature vectors according to step S1. The key is assigning accurate motivational labels to this historical behavioral data. These motivational labels can be obtained in various ways, such as inviting educational psychology experts to evaluate users' behavioral sequences and provide quantitative motivational scores, or directly collecting users' self-reported data at specific time points using standardized psychological scales embedded within the learning platform, such as the Academic Motivation Scale (AMS). Finally, each behavioral feature vector is paired with its corresponding motivational label, forming data pairs in the form of behavioral feature vector and motivational level label. All these data pairs together constitute the training dataset.

[0030] The neural network model is iteratively trained using a constructed training dataset, enabling the model to learn the mapping relationship between behavioral data and motivation levels.

[0031] The predictive accuracy of the neural network model was verified, and the trained neural network model was used as a motivation monitoring model for assessing motivation levels.

[0032] It should be noted that, next, the neural network model will be iteratively trained using the constructed training dataset. This neural network model, such as a multilayer perceptron with multiple hidden layers, is designed to learn the complex mapping relationship between behavioral features and motivation levels. During training, the model receives a behavioral feature vector as input and outputs a predicted motivation level value. The system uses a loss function to quantify the difference between this predicted value and the true motivation label, with the goal of minimizing the total loss across the entire training dataset. This optimization process can be expressed as: ,in, This represents the optimal set of internal weights obtained after training the neural network model. This indicates the search for weights that minimize the objective function. ; This represents summing the loss over all samples in the training dataset; This represents a loss function, such as the mean squared error function, used to calculate the difference between the predicted and the true values; The neural network model is based on the input behavioral feature vector. and current weight The generated predicted values ​​of motivation level; Is with The corresponding real motivation tags, This represents the input behavior feature vector number in the training dataset. The internal weights of the model are continuously adjusted through backpropagation and gradient descent optimizers. The training continues until the model's performance on the validation set reaches the preset convergence criterion. After training, the model's prediction accuracy is finally validated by calculating the mean absolute error (MAE) between the predicted values ​​and the true labels. This MAE is then compared to a preset MAE threshold. If the MAE is less than the threshold, the validation passes; otherwise, it fails. The MAE threshold can be set to 0.1 based on historical data comparison and expert judgment. Once validated, the trained neural network model with its weights fixed becomes the motivation monitoring model used in the system for real-time evaluation of motivation levels.

[0033] It should also be noted that the specific operation for calculating the current motivation level is as follows: The user behavior feature vector within a certain time window is used as input to the motivation monitoring model. The model performs a series of forward propagation calculations and ultimately outputs a standardized continuous value, which represents the current motivation level. This value is mapped to a preset range, such as 0 to 1; a higher value indicates a stronger current learning motivation. This calculation process can be represented as: ,in, Representing the user at a specific point in time The current level of motivation is a dimensionless evaluation value; Representing the user at a specific point in time The feature vector in the corresponding behavioral data set contains multi-dimensional quantitative indicators such as learning engagement and learning achievement. This represents the nonlinear mapping function performed by the motivation monitoring model.

[0034] Analyze the changing patterns of behavioral data sets over time to generate motivation trend predictions.

[0035] It should be noted that the specific process of generating motivation trend prediction includes: 1) Arranging the current motivation levels at each time point in chronological order to form a motivation level time series; 2) Calculating the time mean and the average motivation level in the motivation level time series. Then, calculating the difference between the time value and the time mean at each point, and the difference between the motivation level value and the average motivation level, multiplying these two differences and summing all products to obtain the covariance. At the same time, calculating the square of the difference between the time value and the time mean at each point, and summing all squared values ​​to obtain the variance. Finally, the slope is obtained by dividing the covariance by the variance; 3) The sign of the slope indicates the direction of the trend. A positive slope means that the motivation level increases over time, and a negative slope means that the motivation level decreases over time. The magnitude of the absolute value of the slope indicates the strength of the motivation trend. The larger the absolute value, the more obvious the motivation trend; the smaller the absolute value, the gentler the motivation trend.

[0036] The current motivation level calculated by integration and the generated motivation trend prediction together constitute motivation state information.

[0037] It should be noted that the system integrates the calculated current motivation level with the generated motivation trend prediction, encapsulating them together into a structured data object, namely motivational state information. This information not only includes a real-time snapshot of the user's motivation but also predictive insights into its dynamic changes.

[0038] This invention combines static motivation level assessment with dynamic trend prediction to construct a more comprehensive and forward-looking user motivation profile. It overcomes the judgment lag problem that can result from relying solely on assessments at a single point in time, enabling early identification of potential learning motivation decline risks, even when the user's current motivation level is still within an acceptable range but a downward trend is already evident. This profound insight into motivational states provides a precise basis for subsequent intervention strategy generation, making the timing of incentive measures more forward-looking and proactive. This allows for more effective prevention rather than merely remediation of learning motivation decline, significantly enhancing the intelligence and effectiveness of the entire intervention system.

[0039] S3. Personalized Incentive Strategy Generation: Generate personalized incentive strategies based on motivational state information.

[0040] In a specific embodiment of the present invention, the specific steps of generating a personalized incentive strategy based on motivational state information include: parsing the motivational trend prediction in the motivational state information and determining the motivational decline risk level.

[0041] It should be noted that the specific method for determining the motivation decline risk level is as follows: extract the motivation trend intensity from the motivation trend prediction in the motivation state information, compare the motivation trend intensity with the motivation trend intensity range corresponding to each motivation decline risk level stored in the database, and if the motivation trend intensity is within the motivation trend intensity range corresponding to a certain motivation decline risk level, then the motivation decline risk level is taken as the final motivation trend intensity. The motivation decline risk level includes three levels: high risk, medium risk, and low risk.

[0042] Based on the determined motivation decline risk level, match the various incentive types corresponding to the current motivation decline risk level from the incentive library that stores various incentive types corresponding to each motivation decline risk level, and then match and select an initial incentive type from the various incentive types.

[0043] In a specific embodiment of the present invention, the specific steps of matching and selecting an initial incentive type from multiple incentive types include: obtaining the user's historical incentive response records to form user preference data.

[0044] Assess user preference data and, in conjunction with the user's current learning context, prioritize various incentive types corresponding to the current level of motivational decline risk.

[0045] From the incentive types sorted by priority, select the one with the highest priority as the initial incentive type.

[0046] It's important to note that the system first constructs user preference data for each user by continuously tracking and recording user responses to all historically implemented incentive strategies. Specifically, the system records each incentive event, including the type, intensity, and timing of the incentive, as well as the changes in key metrics within the user behavior data set over a period after the incentive is pushed. It then calculates the difference between the learning duration and the set reference learning duration, and compares this difference to the set reference learning duration to obtain the learning duration effectiveness. Similarly, it calculates the difference between the task completion rate and the set reference task completion rate, and compares this difference to the set reference task completion rate to obtain the task completion rate effectiveness. Finally, it adds the learning duration effectiveness to the task completion rate effectiveness to obtain the results for each incentive category. The system first assesses the historical validity of each incentive type. Secondly, it analyzes the user's current learning context, obtaining the difference between the user's most recent test score and a set reference test score. This difference is then compared to the set reference test score to obtain the test score fit. Next, the system obtains the difference between the user's forum activity level and a set reference activity level, and this difference is compared to the set reference activity level to obtain the activity fit. Finally, the system adds the test score fit and the activity fit to obtain the context fit for each incentive type. Finally, a weighted summation model is used to calculate the final priority of each incentive type. This model can be expressed as: ,in, It is an incentive type Ultimately, this is the final priority, which is a comprehensive evaluation value; Representing the Types of incentives Indicates the type number of the incentive. ; User preferences for incentive types are obtained from user preference data. Historical validity; The type of motivation is derived from the current learning context. Contextual fit; and These are preset weighting coefficients, representing the importance of historical preferences and the current context in decision-making, and satisfying the following conditions: After calculating the priority of all available stimulus types, the system sorts them from highest to lowest priority and automatically selects the stimulus type with the highest priority in the sorted list as the initial stimulus type to be passed to subsequent steps for dynamic adjustment.

[0047] In a specific embodiment of the present invention, in the reinforcement learning course intervention method for learning motivation decline, the historical preference weight can be 0.6 and the current context weight can be 0.4. The reason for this is that: historical preferences reflect the user's long-term behavior pattern and have a stable guiding role in the acceptance of incentive types, so they are given a higher weight; while the current context directly affects immediate needs, but it has dynamic fluctuations, so its short-term impact needs to be balanced by a relatively lower weight.

[0048] By combining the current motivation level in the motivation state information, the intensity and timing parameters of the initial incentive type are dynamically adjusted to generate a personalized incentive strategy.

[0049] It's important to note that, finally, the system enters a dynamic adjustment phase, refining the selected initial incentive type to generate the final personalized incentive strategy. The core of this phase is to dynamically set two key parameters for the incentive type—intensity and timing—based on the current motivation level from the motivation status information. The intensity parameter determines the strength of the incentive, such as the rarity of the badge or the level of detail in the suggested content; the timing parameter determines the specific time the incentive is pushed out, such as immediately or upon the user's next login. The generation process of the intensity and timing parameters can be represented by the following mapping function: , ,in, For strength parameters, Timing parameters; The initial incentive type is selected based on the risk level; This is the user's current motivation level, a value obtained from motivation status information; and These are two dynamically adjusting functions that output optimal parameter values ​​based on different incentive types and the user's current motivation level. For example, for the same initial incentive type, when the user's current motivation level is high, the intensity parameter might be set to grant a more challenging incentive; while when their motivation level is low, it might be adjusted to grant a more easily accessible incentive to rebuild confidence. Ultimately, the complete scheme, composed of the initial incentive type and its dynamically adjusted intensity and timing parameters, constitutes the output personalized incentive strategy.

[0050] This invention employs a hierarchical, progressive strategy generation mechanism to bridge the gap between general and highly personalized interventions. First, macro-level strategy type matching is performed based on the risk of motivational decline, ensuring the correctness of the intervention direction. Then, micro-level parameter optimization is conducted based on the user's current motivational level, guaranteeing the appropriateness of the intervention intensity. This dual personalization design allows interventions to address motivational decline issues of varying urgency while also aligning with the user's current emotional and cognitive state, achieving a shift from passive response to proactive prevention. Furthermore, the intensity and timing of interventions are more precise and humane. This not only significantly improves the efficiency of incentive resource utilization and avoids excessive interference with users, but also, by providing appropriate external support, more effectively stimulates and maintains users' intrinsic learning motivation.

[0051] S4. Obtaining Implementation Results Feedback: Implement personalized incentive strategies and monitor user response behavior to obtain implementation results feedback.

[0052] In a specific embodiment of the present invention, the specific steps of implementing the personalized incentive strategy and monitoring the user's response behavior include: pushing the personalized incentive strategy to the user through the user interface.

[0053] After personalized incentive strategies are pushed out, users' login frequency, learning task completion rate, and interaction depth are tracked to obtain a set of observed behavioral indicators.

[0054] It should be noted that after the personalized incentive strategy is successfully pushed out, the system immediately initiates a behavior monitoring program within a preset time window to track and quantify the user's response behavior. Within this window, the system continuously collects user interaction events and calculates a set of predefined observation behavior metrics. Specifically, the login frequency metric is obtained by counting the number of unique login sessions during this period; the learning task completion rate metric is derived by querying the course database and calculating the ratio of the number of specified learning tasks completed by the user within this time period to the total number of tasks that should be completed; and the interaction depth metric is a comprehensive measure aimed at evaluating the quality of user participation in community discussions rather than just the quantity, and its calculation method is as follows: ,in, The final score representing the depth of interaction is a dimensionless comprehensive score. The number of posts initiated by users within the monitoring window; The number of replies posted by the user; The average character length of all content posted by the user; , and These represent the number of posts, the number of replies, and the character length to be referenced, respectively. , and These are pre-set weighting coefficients used to adjust the relative importance of different interactive behaviors in the depth assessment. After the monitoring window ends, the system integrates the three specific values ​​calculated—login frequency, learning task completion rate, and interaction depth—into a set, namely, a set of observed user behavior indicators.

[0055] In one specific embodiment of the present invention, =0.4、 =0.4、 =0.2, the value is based on the fact that users actively initiating posts and posting replies can more directly reflect the user's enthusiasm and depth of interaction in community discussions, so it is given a higher weight; while the average character length of the content can initially measure the level of detail of the content, it has a smaller impact on the depth of interaction than the former two, so it is given a lower weight. This setting can more reasonably and comprehensively evaluate the quality of users' participation in community discussions.

[0056] In a specific embodiment of the present invention, the specific steps for obtaining feedback on the implementation effect include: comparing a set of observed behavioral indicators of the user with a behavioral baseline stored in the database for defining the expected effect, and calculating the relative change of each observed behavioral indicator.

[0057] It should be noted that, to obtain feedback on the implementation effect for system adaptive adjustments, this method first compares a set of observed user behavior metrics with a behavioral baseline stored in the database to calculate the relative change of each observed behavior metric. This behavioral baseline is established independently for each user and is specifically defined as the values ​​of the user's login frequency, learning task completion rate, and interaction depth metrics within a time window of the same length prior to the current personalized incentive strategy push. The system calculates the relative change of each metric using the following formula: ,in, Representing the The relative change of an observed behavioral indicator is a dimensionless ratio. The first one obtained within the monitoring window after the incentive is implemented. The actual value of an observed behavioral indicator, such as the monitored login frequency; It is extracted from the database and corresponds to the user's behavioral baseline before the incentive implementation. The value of each indicator, Indicates the observed behavior indicator number, This calculation is performed separately for three metrics: login frequency, learning task completion rate, and interaction depth, resulting in a set of three relative changes.

[0058] The relative changes of each observed behavioral indicator are weighted and summed with their corresponding percentage weights to obtain a comprehensive incentive effect score, which is then used as the core content for feedback on the implementation effect.

[0059] In a specific embodiment of the present invention, when calculating the comprehensive incentive effect score, the typical weight values ​​for login frequency, learning task completion rate, and interaction depth indicators can be set to 0.3, 0.5, and 0.2, respectively. The basis for these values ​​is that the completion of learning tasks directly reflects the user's progress and achievements in the learning course, and is the most critical factor in evaluating the incentive effect, so it is given the highest weight. Login frequency reflects the user's activity level in accessing the system, which is an important aspect of the incentive effect, so it is given the second highest weight. Although the interaction depth indicator can reflect the user's participation in community discussions, its impact on the overall incentive effect is relatively smaller than that of learning task completion and login behavior, so it is given the lowest weight, thus reasonably measuring the incentive implementation effect.

[0060] This invention, through the introduction of individualized behavioral baselines for comparison and the use of weighted aggregation, successfully transforms multi-dimensional, fragmented behavioral changes after incentives into a standardized, single-dimensional comprehensive incentive effect score. This not only makes the evaluation of incentive effects more objective and accurate, effectively filtering out differences in individual user behavioral habits, but also allows the evaluation criteria to be flexibly aligned with specific teaching objectives through a configurable weight system. The resulting clear and quantifiable feedback signal provides an ideal input for the subsequent adaptive adjustment model, enabling it to clearly determine the success or failure and extent of the previous intervention. This is a crucial step in achieving closed-loop optimization and continuous learning of the entire intervention system.

[0061] S5. Adaptive Adjustment of Incentive Strategies: Based on feedback on implementation results, the motivation analysis process and the generation process of personalized incentive strategies are adaptively adjusted.

[0062] In a specific embodiment of the present invention, the specific steps for adaptively adjusting the motivation analysis process and the generation process of personalized incentive strategies include: inputting the implementation effect feedback into the adaptive adjustment model used to calculate the parameter update amount to obtain the parameter adjustment amount.

[0063] It should be noted that, to achieve self-evolution and continuous optimization of the intervention method, this method first uses the implementation effect feedback generated in the previous stage—that is, the comprehensive incentive effect score—as the core input signal, feeding it into an adaptive adjustment model. This model is the core of the entire system's learning capability; its function is to transform a scalar effect score into specific parameter adjustment instructions for multiple modules within the system. The model calculates a set of parameter adjustment amounts based on the sign and magnitude of the score. A positive score indicates that the previous intervention was effective, and the model will calculate a positive adjustment amount to strengthen the relevant decision; conversely, a negative score will generate a negative adjustment amount to weaken the relevant decision. The calculation of the parameter adjustment amount can be expressed as: ,in, The representative parameter adjustment amount is a basic increment used to update system parameters; It is a preset learning rate used to control the step size of each adjustment, so as to avoid excessive oscillation of the system due to a single feedback. It is a comprehensive score of the incentive effect of the input.

[0064] In a specific embodiment of the present invention, the preset learning rate can be 0.1. The basis for this value is that experiments have shown that if the learning rate is too large, it may cause the system to fluctuate violently during the adjustment process and make it difficult to converge to a stable state; if the value is too small, the adjustment process will be too slow and affect the system's response efficiency to feedback. The value of 0.1 can control the risk of oscillation while ensuring that the system adjusts according to feedback at a relatively appropriate speed, achieving a good balance between stability and response speed.

[0065] The internal weights of the motivation monitoring model used to assess motivation levels are updated by adjusting the applied parameters.

[0066] Adjust the application parameters and synchronously update the selection weights of the corresponding incentive strategies and the generation rules of personalized incentive strategies in the incentive library.

[0067] It should be noted that after obtaining the parameter adjustment amount, the system will simultaneously update the motivation analysis process and the personalized incentive strategy generation process. First, at the motivation analysis level, although the adjustment amount is not directly used to modify the internal weights of the motivation monitoring model in real time, the system will store the complete record of this intervention, including the behavioral data set before the intervention, the predicted motivational state information, and the final implementation effect feedback, as a new high-quality training sample in the database. When a certain number of new samples are accumulated, the system will trigger incremental training or periodic retraining of the motivation monitoring model. In this way, feedback is indirectly but stably applied to update its internal weights, thereby improving the accuracy of future motivation assessments. Second, at the personalized incentive strategy generation level, the parameter adjustment amount is directly used to update two core components. First, the selection weights of the corresponding incentive strategies in the incentive library are updated. Specifically, for the incentive type executed this time, its selection weights will be updated according to the following rules: ,in, This is the updated selection weight; This refers to the selection weight of the incentive type before execution. This mechanism increases the probability that effective incentive types will be selected in the future, and vice versa. Second, the generation rules for personalized incentive strategies are updated synchronously. This means that parameter adjustments are used to fine-tune the dynamic adjustment function that determines the intensity and timing of incentives. For example, if applying a specific intensity of incentive is effective at a certain current motivation level, the system will strengthen the tendency to choose that intensity at that motivation level, thereby making the parameter configuration of the strategy more refined and effective.

[0068] This invention, through the introduction of a closed-loop adaptive adjustment mechanism based on implementation effect feedback, endows the entire intervention system with the ability to dynamically learn and self-optimize. It transforms every interaction with the user into a learning opportunity, enabling the system to learn from both successful and unsuccessful interventions. This continuous iterative optimization ensures that the system's two core capabilities—the ability to accurately perceive the user's state and the ability to make effective intervention decisions based on this perception—both continuously improve over time. Ultimately, this method not only executes preset rules but also evolves intervention patterns that truly adapt to each user's unique and dynamically changing psychological needs, thereby maintaining the efficiency and appeal of intervention measures in the long term.

[0069] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A reinforcement learning curriculum intervention method for addressing declining learning motivation, characterized in that, Includes the following steps: S1. Behavioral Data Set Acquisition: Acquire user learning behavior data and process the learning behavior data to obtain a behavioral data set; S2. Motivational State Information Generation: Perform motivational analysis on the behavioral data set to generate motivational state information that includes the current motivational level and motivational trend prediction. The specific steps for performing motivational analysis on the behavioral data set to generate motivational state information including the current motivational level and motivational trend prediction include: inputting the behavioral data set into a motivational monitoring model for evaluating motivational levels, and calculating the current motivational level; Analyze the changing patterns of behavioral data sets over time to generate motivation trend predictions; The current motivation level calculated by integration and the generated motivation trend prediction together constitute motivation state information; S3. Personalized Incentive Strategy Generation: Generate personalized incentive strategies based on motivational state information; The specific steps for generating personalized incentive strategies based on motivational state information include: analyzing motivational trend predictions in motivational state information and determining the motivational decline risk level. Based on the determined motivation decline risk level, match the current motivation decline risk level with the incentive library that stores the various incentive types corresponding to each motivation decline risk level, and then match and select an initial incentive type from the various incentive types. By combining the current motivation level in the motivation state information, the intensity and timing parameters of the initial incentive type are dynamically adjusted to generate a personalized incentive strategy; S4. Obtaining Implementation Results Feedback: Implement personalized incentive strategies and monitor user response behavior to obtain implementation results feedback; S5. Adaptive Adjustment of Incentive Strategies: Based on feedback on implementation results, the motivation analysis process and the generation process of personalized incentive strategies are adaptively adjusted. The specific steps for adaptively adjusting the motivation analysis process and the generation process of personalized incentive strategies include: inputting the implementation effect feedback into the adaptive adjustment model used to calculate the parameter update amount to obtain the parameter adjustment amount; The internal weights of the motivation monitoring model used to assess motivation levels are updated by adjusting the applied parameters. Adjust the application parameters and synchronously update the selection weights of the corresponding incentive strategies and the generation rules of personalized incentive strategies in the incentive library.

2. The reinforcement learning curriculum intervention method for learning motivation decline according to claim 1, characterized in that: The specific steps for acquiring user learning behavior data and processing the learning behavior data to obtain a behavior data set include: Real-time collection of user interaction events on the learning platform to obtain raw behavioral data; The obtained raw behavioral data is filtered and denoised to generate clean behavioral data. The generated cleaning behavior data is then subjected to feature extraction and aggregation to form a behavior data set.

3. The reinforcement learning curriculum intervention method for learning motivation decline according to claim 1, characterized in that: The motivation monitoring model used to assess motivation levels is obtained through the following steps: Obtain historical learning behavior data and their corresponding motivation labels to construct a training dataset; The neural network model is iteratively trained using the constructed training dataset, enabling the neural network model to learn the mapping relationship between behavioral data and motivation levels. The predictive accuracy of the neural network model was verified, and the trained neural network model was used as a motivation monitoring model for assessing motivation levels.

4. The reinforcement learning curriculum intervention method for learning motivation decline according to claim 1, characterized in that: The specific steps for matching and selecting an initial incentive type from multiple incentive types include: Obtain users' historical incentive response records to form user preference data; Assess user preference data and, in conjunction with the user's current learning context, prioritize various incentive types corresponding to the current level of motivational decline risk. From the incentive types sorted by priority, select the one with the highest priority as the initial incentive type.

5. A reinforcement learning curriculum intervention method for addressing declining learning motivation according to claim 1, characterized in that: The specific steps for implementing personalized incentive strategies and monitoring user response behavior include: Personalized incentive strategies are pushed to users through the user interface; After personalized incentive strategies are pushed out, users' login frequency, learning task completion rate, and interaction depth are tracked to obtain a set of observed behavioral indicators.

6. A reinforcement learning curriculum intervention method for addressing declining learning motivation according to claim 5, characterized in that: The specific steps for obtaining feedback on the implementation effect include: The user's set of observed behavioral metrics are compared with the behavioral baseline stored in the database to define the expected effect, and the relative change of each observed behavioral metric is calculated. The relative changes of each observed behavioral indicator are weighted and summed with their corresponding percentage weights to obtain a comprehensive incentive effect score, which is then used as the core content for feedback on the implementation effect.

Citation Information

Patent Citations

  • Elderly health situation monitoring and early warning system and device based on digital twinborn

    CN119673464A

  • Book intelligent recommendation method and system based on user portrait

    CN120372099A