Intelligent Learning Path Planning System and Method for English Vocabulary Based on Reinforcement Learning

Through the intelligent learning path planning system of English vocabulary based on reinforcement learning, combined with learner characteristics, equipment performance and environmental monitoring data, the learning path is dynamically adjusted, which solves the problems of rigid learning paths and equipment interference, and improves learning efficiency and resource utilization.

CN119904005BActive Publication Date: 2025-07-18LONGYAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510382623.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In the prior art, the learning effect of learners is reflected only by the test results, and cannot accurately reflect the learners' real learning situation, resulting in insufficient dynamic adjustment of personalized learning paths and reducing learning efficiency.

Method used

The English vocabulary intelligent learning path planning system based on reinforcement learning is adopted. By obtaining learner individual feature information, collecting learning equipment performance data and learning environment monitoring data in real time, building a multi-dimensional credibility assessment model, and dynamically adjusting the learning path, solving the problem of rigid learning paths and disconnection between task allocation and user capabilities.

Benefits of technology

It realizes personalized planning of learning paths, improves the targetedness and time resource utilization of vocabulary exercises, ensures the achievement rate of learning goals, and maintains the consistency and stability of learning path planning in a complex hardware environment, and reduces the misjudgment rate caused by device interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904005B_ABST
    Figure CN119904005B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent learning path planning system and method for English vocabulary based on reinforcement learning, which relates to the technical field of educational learning, and includes: a preliminary learning path planning module, a learning device feedback module, a learning effect analysis module, and a learning path update module. By providing the intelligent learning path planning system and method for English vocabulary based on reinforcement learning, the present invention obtains the basic vocabulary test results, planned learning time and goals of learners, and combines dynamic vocabulary matching with learning starting point setting to achieve personalized planning of the learning path. The reinforcement learning model is used to dynamically divide vocabulary tasks, and the vocabulary difficulty gradient, usage frequency and multi-modal practice methods are integrated into the phased tasks, solving the problems of rigid learning paths and disconnection between task assignment and user capabilities in traditional methods, thereby significantly improving the pertinence of vocabulary practice and the utilization rate of time resources while ensuring the achievement rate of learning goals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of education and learning, and specifically to an intelligent learning path planning system and method for English vocabulary based on reinforcement learning. Background Art

[0002] Driven by the acceleration of globalization and the innovation of educational concepts, English vocabulary learning is undergoing a paradigm transformation from rote memorization to ability cultivation. Cognitive science research has revealed the deep mechanisms of human memory and understanding, emphasizing the improvement of vocabulary activation rate through semantic network construction, contextual synesthesia, and distributed practice, rather than simply relying on repeated memorization. At the same time, through a systematic grading system, real context simulation, and task-based output design, it helps learners achieve a natural transformation from knowledge accumulation to communicative ability in a dynamic language environment, and ultimately form a sustainable autonomous learning ecosystem.

[0003] For example, the invention patent with the publication number CN112149994B is an English personal ability tracking learning system based on statistical analysis, including an English personal ability test and analysis module, a learning path planning module, a module for matching students' abilities with courses, a tutoring resource module, a student ability tracking module, and a real question testing module; this application is based on the statistical analysis of students' exam samples, starting from an English ability test, and oriented towards path planning for setting learning goals for students, tracking and evaluating the learning process and progress of students at each step, calculating the progress value of students and the distance from the learning goal, and continuously correcting the accuracy of the evaluated progress value, ability value, and deduced learning path with real question testing; the system has made detailed label descriptions for courses and tutoring resources, and the data generated during students' learning will all enter the student personal database in the system, reflecting the daily learning progress.

[0004] For example, the invention patent with the publication number CN112734142B is a resource learning path planning method and device based on deep learning, related to artificial intelligence technology, including first collecting the learning data click records of virtual objects in real time and storing them in the virtual object database; then obtaining target data according to screening conditions in the virtual object database to form a sample set for model training to obtain a prediction model; finally, if the initial input features of the current course sub-trajectory information are received, input them into the prediction model for operation to obtain an output result.

[0005] However, in the process of implementing the technical solution of the present application, it is found that the above technologies have at least the following technical problems: the learning effect of current learners is only reflected by test results. However, learners will be interfered by external factors during actual learning, resulting in the test results being unable to accurately reflect the true learning situation of learners, thereby affecting the dynamic adjustment of personalized learning paths and reducing the learning efficiency of learners. Summary of the Invention

[0006] In view of the deficiencies of the prior art, the present invention provides an intelligent learning path planning system and method for English vocabulary based on reinforcement learning, which can effectively solve the problems involved in the above-mentioned background art.

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: In the first aspect of the present invention, an intelligent learning path planning system for English vocabulary based on reinforcement learning is provided, including: a preliminary learning path planning module, which is used to obtain the individual characteristics information of the learner and conduct a preliminary planning of the vocabulary learning path. The preliminary planning of the vocabulary learning path includes vocabulary library matching, learning starting point setting, and learning task division; a learning device feedback module, which is used to collect the performance data of the learner's learning device in real time, quantify the interference degree of the learning device performance data on the learning effect, obtain a learning device interference quantification index, and judge whether the learning device performance will interfere with the learning effect and give feedback; a learning effect analysis module, which is used to monitor the learning environment of the learner, collect the learner's interaction data in real time, and evaluate the reference value of the detection result of the learner's vocabulary learning effect in combination with the learning device interference quantification index; a learning path update module, which is used to obtain the vocabulary learning effect, judge whether the vocabulary learning task is completed according to the reference value of the vocabulary learning effect detection result, and update the learning path according to the completion situation of the vocabulary learning task.

[0008] As a further method, the specific process of conducting the preliminary planning of the vocabulary learning path is as follows: The individual characteristics information of the learner includes the learner's basic vocabulary test results, set learning time, and learning goal; the learning vocabulary library is matched according to the learner's set learning goal; the learning starting point of the learner is set according to the learner's basic vocabulary test results; the learning tasks include each vocabulary to be learned and each vocabulary practice method.

[0009] As a further method, the specific process of quantifying the interference degree of the learning device performance data on the learning effect is as follows: The learning device performance data includes CPU occupancy rate, memory occupancy rate, and vocabulary refresh response delay; for the CPU occupancy rate and memory occupancy rate, the maximum theoretical occupancy rate is set as the reference value according to the device hardware specifications, and the actual occupancy rate is converted into a ratio relative to the reference value. For the vocabulary refresh response delay, a dynamic reference range is set based on the lowest response delay and the highest response delay of the learning device response in the historical learning task, the vocabulary refresh response delay is mapped to the standardized value within the reference range, and weighted processing is performed according to the preset device interference weight coefficient to generate a learning device interference quantification index, and the learning device interference quantification index is used to quantitatively evaluate the interference degree of the learning device performance on the learner's learning effect.

[0010] As a further method, to determine whether the performance of the learning device will interfere with the learning effect and provide feedback, the specific process is as follows: Obtain the preset threshold of the learning device interference quantification index from the vocabulary learning database; Compare the learning device interference quantification index with the learning device interference quantification index threshold. If the learning device interference quantification index is less than the learning device interference quantification index threshold, mark the learning device interference as negligible interference and do not include it in the evaluation of the reference value of the vocabulary learning effect detection result; If the learning device interference quantification index is not less than the learning device interference quantification index threshold, mark the learning device interference as significant interference, push a learning device interference prompt to the learner, and include the learning device interference in the evaluation of the reference value of the vocabulary learning effect detection result.

[0011] As a further method, to evaluate the reference value of the vocabulary learning effect detection result of the learner, the specific process is as follows: Process the learner interaction data to obtain the comprehensive interaction anomaly evaluation parameters of each vocabulary of the learner; By integrating the learning environment monitoring data, the learning device interference quantification index, and the comprehensive interaction anomaly evaluation parameters of each vocabulary of the learner, construct a multi-dimensional credibility evaluation model for each vocabulary to obtain the multi-dimensional credibility evaluation value of each vocabulary. The multi-dimensional credibility evaluation value of the vocabulary is used to quantitatively evaluate the reference value of the vocabulary learning effect detection result; The learning environment monitoring data includes the light intensity, noise decibel value, and device tilt angle at each monitoring time point.

[0012] As a further method, to process the learner interaction data to obtain the comprehensive interaction anomaly evaluation parameters of each vocabulary of the learner, the specific process is as follows: Based on the interaction data of all users in the historical tasks, establish a group benchmark data set of vocabulary learning duration, practice type time-consuming ratio, and interaction frequency; For the interaction data of the current learner, respectively process to obtain the deviation ratio of each vocabulary learning duration from the benchmark duration, the deviation amplitude of the time-consuming ratio of each practice type from the benchmark ratio, and the difference coefficient of the interaction frequency of each vocabulary learning process from the benchmark frequency; Use dynamic weighting to aggregate the above deviations to generate the comprehensive interaction anomaly evaluation parameters of each vocabulary of the learner. The comprehensive interaction anomaly evaluation parameters of each vocabulary of the learner are used to characterize the overall deviation degree of the current learner's interaction behavior from the group standard; The learner interaction data includes the learning duration of each vocabulary, the time-consuming ratio of each practice type of each vocabulary, and the interaction frequency of each vocabulary learning process.

[0013] As a further method, to determine whether the vocabulary learning task is completed and update the learning path according to the completion status of the vocabulary learning task, the specific process is as follows: The vocabulary learning effect includes the correct rate of each vocabulary practice; match the multi-dimensional credibility evaluation values of each vocabulary with the vocabulary practice correct rate thresholds corresponding to the preset multi-dimensional credibility evaluation value intervals of each vocabulary in the vocabulary learning database, and compare the correct rate of each vocabulary practice with the correct rate thresholds of each vocabulary; if the correct rate of a certain vocabulary practice is not less than the correct rate threshold of this vocabulary, mark this vocabulary as an effectively learned vocabulary, otherwise, mark this vocabulary as an ineffectively learned vocabulary, and count the number of ineffectively learned vocabulary; based on the number of ineffectively learned vocabulary and the multi-dimensional credibility evaluation values of each vocabulary, process to obtain the completion degree of the vocabulary learning task, and update the learning path according to the completion degree of the vocabulary learning task.

[0014] As a further method, updating the learning path according to the completion degree of the vocabulary learning task includes: comparing the completion degree of the vocabulary learning task with the preset vocabulary learning task completion degree threshold in the vocabulary learning database. If the completion degree of the vocabulary learning task is less than the vocabulary learning task completion degree threshold, mark the vocabulary learning task as an unfinished learning task, and the reinforcement learning model obtains the unfinished learning task and re-performs the practice task allocation.

[0015] As a further method, updating the learning path according to the completion degree of the vocabulary learning task further includes: if the completion degree of the vocabulary learning task is not less than the vocabulary learning task completion degree threshold, the reinforcement learning model allocates a review task for the vocabulary learning task, and obtains the ineffectively learned vocabulary in this vocabulary learning task, and re-performs the practice task allocation for the ineffectively learned vocabulary.

[0016] The second aspect of the present invention provides an intelligent learning path planning method for English vocabulary based on reinforcement learning, including: S1. Obtain the individual characteristic information of the learner and conduct a preliminary planning of the vocabulary learning path. The preliminary planning of the vocabulary learning path includes word bank matching, learning starting point setting, and learning task division; S2. Real-time collect the performance data of the learner's learning device, and quantitatively process the interference degree of the learning device performance data on the learning effect to obtain a learning device interference quantization index, and judge whether the learning device performance will interfere with the learning effect and give feedback; S3. Monitor the learning environment of the learner, and real-time collect the learner interaction data, and evaluate the reference value of the vocabulary learning effect detection result of the learner in combination with the learning device interference quantization index; S4. Obtain the vocabulary learning effect, and judge whether the vocabulary learning task is completed according to the reference value of the vocabulary learning effect detection result, and update the learning path according to the completion status of the vocabulary learning task.

[0017] Compared with the prior art, the embodiments of the present invention at least have the following beneficial effects:

[0018] (1) The present invention provides an intelligent learning path planning system and method for English vocabulary based on reinforcement learning. By obtaining the basic vocabulary test results, planned learning time, and goals of learners, and combining dynamic vocabulary database matching with learning starting point setting, personalized planning of the learning path is achieved. The reinforcement learning model is used to dynamically divide vocabulary tasks, integrating vocabulary difficulty gradients, usage frequencies, and multimodal practice methods into phased tasks, solving the problems of rigid learning paths and disconnection between task allocation and user capabilities in traditional methods. Thus, while ensuring the achievement rate of learning goals, the pertinence of vocabulary practice and the utilization rate of time resources are significantly improved.

[0019] (2) The present invention analyzes learning device interference and compares it with a preset threshold to trigger a hierarchical response strategy, filtering out low-confidence learning effect data caused by deteriorated device performance at the data source, achieving precise perception and dynamic compensation of device interference. This helps to maintain the consistency and stability of learning path planning in complex hardware environments. Especially in scenarios of high-concurrency tasks or low-performance devices, it can help reduce the misjudgment rate caused by device lag, improving the reliability and anti-interference fault tolerance ability of personalized learning path updates.

[0020] (3) The present invention constructs a multi-dimensional credibility evaluation model by integrating learning environment monitoring data, learning device interference quantification indicators, and comprehensive interaction anomaly evaluation parameters of each vocabulary of learners, solving the limitations of traditional solutions that rely solely on user behavior or device indicators. At the same time, it improves the accuracy of judging the reference value of learning effects, especially in complex scenarios where device and environmental interference are not directly reflected in the explicit behavior of users, ensuring the robustness and anti-interference of dynamic updates of the learning path.

[0021] Of course, it is not necessary for any product implementing the present invention to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the connection of system modules of the present invention.

[0023] Figure 2 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0025] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0026] Referring to Figure 1 As shown, the first aspect of the present invention provides an intelligent learning path planning system for English vocabulary based on reinforcement learning, including: a preliminary learning path planning module, configured to obtain individual learner characteristic information and perform a preliminary planning of the vocabulary learning path, where the preliminary planning of the vocabulary learning path includes word bank matching, learning starting point setting, and learning task division; a learning device feedback module, configured to collect in real time the performance data of the learner's learning device, quantify the degree of interference of the learning device performance data on the learning effect, obtain a learning device interference quantification index, determine whether the learning device performance will interfere with the learning effect and give feedback; a learning effect analysis module, configured to monitor the learner's learning environment, collect in real time the learner's interaction data, and evaluate the reference value of the detection result of the learner's vocabulary learning effect in combination with the learning device interference quantification index; a learning path update module, configured to obtain the vocabulary learning effect, judge whether the vocabulary learning task is completed according to the reference value of the detection result of the vocabulary learning effect, and update the learning path according to the completion situation of the vocabulary learning task. A vocabulary learning database is used to store data during the vocabulary learning process.

[0027] Specifically, for the preliminary planning of the vocabulary learning path, the specific process is as follows: the individual learner characteristic information includes the learner's basic vocabulary test result, the set learning time, and the learning goal.

[0028] Among them, the learner's basic vocabulary test result includes the vocabulary size, vocabulary difficulty distribution, and common error types, and the learner can be pre-tested for vocabulary, and the test results can be automatically obtained through the vocabulary test system. The learning time refers to the time period planned by the learner for vocabulary learning, and the learning goal includes the vocabulary size and mastery level expected by the learner.

[0029] A learning word bank is matched according to the learning goal set by the learner.

[0030] It should be explained that the learning word bank refers to a vocabulary set covering different topics and different difficulties. The learning word bank can be automatically matched according to the learner's basic vocabulary test result and learning goal, or can be manually selected by the learner.

[0031] Set the learning starting point of the learner according to the test results of the learner's basic vocabulary.

[0032] The learning starting point refers to the starting point of the personalized learning path determined according to the learner's current vocabulary mastery level. By obtaining the learner's vocabulary size and vocabulary difficulty distribution through the basic vocabulary test, and according to the vocabulary size and difficulty distribution, select the vocabulary set of the initial learning tasks that matches the learner's basic vocabulary test results from the preset learning word bank.

[0033] The learning time set by the learner includes multiple learning time periods. Mark the vocabulary that needs to be mastered in each learning time period as each learning task. For a single vocabulary, it may be practiced multiple times in the same learning task, and each practice corresponds to a vocabulary practice method, such as listening, speaking, reading, etc. The learning task division refers to allocating the vocabulary to be learned and different vocabulary practice methods in each learning time period, and performing the learning task allocation according to the reinforcement learning model. For example, according to factors such as the difficulty level, usage frequency, and relevance between vocabularies of the vocabulary, associate the vocabularies to the same learning task, and then according to the user's current learning stage and goals, allocate the corresponding learning tasks in sequence.

[0034] The reinforcement learning model is a vocabulary learning method based on the reinforcement learning theory, which regards vocabulary learning as a process in which an agent interacts with vocabulary learning tasks and related situations. Reinforcement Learning (RL) has been widely used in the field of artificial intelligence, involving robot control, game intelligence, etc. Common reinforcement learning algorithms include Q-learning, Deep Q-Network, Policy Gradient Method, and Actor-Critic Method. The agent completes the vocabulary learning task by adopting different vocabulary practice methods, and adjusts the vocabulary practice method in a timely manner according to the feedback of the vocabulary learning task, so as to optimize the vocabulary learning effect.

[0035] The learning task includes each vocabulary to be learned and each vocabulary practice method.

[0036] In a specific embodiment, the present invention realizes the personalized planning of the learning path by obtaining the test results of the learner's basic vocabulary, planned learning time and goals, combining dynamic word bank matching and learning starting point setting. Based on the learner's vocabulary size, difficulty distribution and error type characteristics, adaptively screen the initial learning word bank that matches the current level, and according to the preset learning goals and time constraints, use the reinforcement learning model to dynamically divide the vocabulary tasks, and integrate the vocabulary difficulty gradient, usage frequency and multi-modal practice methods into the stage tasks, solving the problems of rigid learning paths and disjointed task allocation and user capabilities in traditional methods, thus significantly improving the pertinence of vocabulary practice and the utilization rate of time resources while ensuring the achievement rate of learning goals.

[0037] It should be understood that the present invention evaluates the learning effect of vocabulary and the reference value of the learning effect detection result based on the multiple practice data of vocabulary. The multiple practices of vocabulary are not necessarily continuous in the monitoring time period.

[0038] Specifically, the interference degree of the learning device performance data on the learning effect is quantified. The specific process is as follows: The learning device performance data includes the CPU occupancy rate, the memory occupancy rate, and the vocabulary refresh response latency.

[0039] In this embodiment, the CPU occupancy rate, the memory occupancy rate, and the vocabulary refresh response latency can be directly obtained through the corresponding interfaces or tools provided by the device operating system. The vocabulary refresh response latency refers to the average time for the device to respond to the vocabulary display when the learner interacts with the device.

[0040] For the CPU occupancy rate and the memory occupancy rate, the maximum theoretical occupancy rate is set as the reference value according to the device hardware specifications, and the actual occupancy rate is converted into a ratio relative to the reference value. For the vocabulary refresh response latency, a dynamic reference range is set based on the lowest response latency and the highest response latency of the learning device response in the historical learning tasks. The vocabulary refresh response latency is mapped to the standardized value within the reference range and weighted according to the preset device interference weight coefficient to generate a learning device interference quantization index. The learning device interference quantization index is used to quantitatively evaluate the interference degree of the learning device performance on the learner's learning effect.

[0041] In a specific embodiment, the acquisition method of the learning device interference quantization index is as follows:

[0042] ;

[0043] ;

[0044] In the formula, represents the learning device interference quantization index, represents the device interference weight coefficient corresponding to the preset CPU occupancy rate, represents the device interference weight coefficient corresponding to the preset memory occupancy rate, represents the device interference weight coefficient corresponding to the preset vocabulary refresh response latency, represents the current CPU occupancy rate, represents the maximum CPU occupancy rate of the device, represents the current memory occupancy rate, represents the maximum memory occupancy rate of the device, represents the standardized value of the vocabulary refresh response latency, Indicates the vocabulary refresh response delay. YCmin is the lowest response delay in historical tasks, and YCmax is the highest response delay in historical tasks. Among them, 、 and take values in the range of 0 to 1, reflecting the importance of different types of performance data interfering with the learning effect. The device interference weight coefficient can be obtained by matching the specific device type. For example, a mapping relationship is established by corresponding the device type with each device interference weight coefficient and stored in the vocabulary learning database. The weight coefficient suitable for the learning device is selected from the vocabulary learning database through the mapping relationship.

[0045] The learning device interference quantization index is obtained by jointly processing the CPU occupancy rate, memory occupancy rate, and vocabulary refresh response delay. The greater the CPU occupancy rate, memory occupancy rate, and vocabulary refresh response delay, the higher the degree of interference of the learning device on the learner's learning effect. The CPU occupancy rate, memory occupancy rate, and vocabulary refresh response delay affect each other. A high CPU occupancy rate will lead to a scheduling delay of the vocabulary refresh instruction, insufficient memory resources will exacerbate the vocabulary loading lag, and the cumulative effect of the refresh delay will further inversely push up the resource consumption of the CPU and memory, forming a positive feedback loop of device performance degradation.

[0046] The present invention analyzes the CPU occupancy rate, memory occupancy rate, and vocabulary refresh response delay, maps the three to a unified interference evaluation dimension, and combines the preset weight coefficient to quantify the comprehensive interference degree of device performance on the learning process, thereby breaking through the limitations of traditional single-index evaluation. This solution can adapt to the learning scenarios of devices with different hardware configurations, and while ensuring the accuracy of the learning path planning, significantly reduce the interference fault tolerance threshold of device performance fluctuations on the learning effect evaluation, providing a data basis for the reliable update of the subsequent learning path.

[0047] Specifically, it is determined whether the learning device performance will interfere with the learning effect and feedback is given. The specific process is as follows: Obtain the preset learning device interference quantization index threshold from the vocabulary learning database; compare the learning device interference quantization index with the learning device interference quantization index threshold. If the learning device interference quantization index is less than the learning device interference quantization index threshold, the learning device interference is marked as negligible interference and is not included in the evaluation of the reference value of the vocabulary learning effect detection result.

[0048] If the learning device interference quantization index is not less than the learning device interference quantization index threshold, the learning device interference is marked as significant interference, and a learning device interference prompt is pushed to the learner, and the learning device interference is included in the evaluation of the reference value of the vocabulary learning effect detection result.

[0049] In a specific embodiment, by analyzing the interference of the learning device and comparing it with a preset threshold, a hierarchical response strategy is triggered to filter out low-confidence learning effect data caused by device performance degradation at the data source, achieving precise perception and dynamic compensation of device interference, which helps to maintain the consistency and stability of the learning path planning in a complex hardware environment. Especially in scenarios of high-concurrency tasks or low-performance devices, it can help reduce the misjudgment rate caused by device lag and improve the reliability and anti-interference fault tolerance ability of personalized learning path updates.

[0050] Specifically, to evaluate the reference value of the detection result of the learner's vocabulary learning effect, the specific process is as follows: Process the learner's interaction data to obtain the comprehensive interaction anomaly evaluation parameter for each vocabulary of the learner.

[0051] By fusing the learning environment monitoring data, the device interference quantization index of the learning device, and the comprehensive interaction anomaly evaluation parameter for each vocabulary of the learner, a multi-dimensional credibility evaluation model for each vocabulary is constructed to obtain the multi-dimensional credibility evaluation value for each vocabulary. The multi-dimensional credibility evaluation value of the vocabulary is used to quantitatively evaluate the reference value of the detection result of the vocabulary learning effect; the learning environment monitoring data includes the light intensity, noise decibel value, and device tilt angle at each monitoring time point.

[0052] The light intensity is collected in real time by an illuminometer, the noise decibel value can be measured by a sound level meter, and the device tilt angle is continuously tracked and recorded by an internal gyroscope. The device tilt angle refers to the degree of inclination of the learning device relative to the standard horizontal plane during use, that is, the angle between the plane where the learning device screen is located and the horizontal plane.

[0053] In a specific embodiment, the calculation formula of the multi-dimensional credibility evaluation model for vocabulary is:

[0054] ;

[0055] In the formula, represents the multi-dimensional credibility evaluation value of the i-th vocabulary, represents the credibility weight coefficient corresponding to the preset device interference quantization index of the learning device, represents the credibility weight coefficient corresponding to the preset comprehensive interaction anomaly evaluation parameter, represents the credibility weight coefficient corresponding to the preset environmental factors, represents the device interference quantization index of the learning device, represents the comprehensive interaction anomaly evaluation parameter of the i-th vocabulary, represents the light intensity of the i-th vocabulary at the t-th monitoring time point, represents the reference appropriate light intensity, represents the allowable deviation value of the light intensity, represents the noise decibel value of the i-th word at the t-th monitoring time point, represents the critical maximum noise decibel value, represents the device tilt angle of the i-th word at the t-th monitoring time point, represents the preset threshold of the learning device interference quantization index. i represents the number of each word, i = 1, 2, 3,..., n, where n represents the total number of words, t represents the number of each monitoring time point, t = 1, 2, 3,..., s, and s represents the total number of monitoring time points.

[0056] Among them, the reference appropriate light intensity and the critical maximum noise decibel value are preset in the vocabulary learning database. The reference appropriate light intensity refers to the optimal light intensity for vocabulary learning in an ideal environment, and the critical maximum noise decibel value is the maximum noise threshold to ensure that the learning effect is not disturbed. 、 and The value range of is between 0 and 1, and it is dynamically adjusted according to the actual learning environment and user behavior data. For example, a mapping relationship is established between the comprehensive interaction anomaly evaluation parameter type and each credibility weight coefficient and stored in the vocabulary learning database, and the credibility weight coefficient is selected from the vocabulary learning database through the mapping relationship.

[0057] The multi-dimensional credibility evaluation value of the word is calculated by weighted calculation of factors such as the learning device interference quantization index, the comprehensive interaction anomaly evaluation parameter, the light intensity deviation value, the noise decibel value, and the device tilt angle. Among them, the larger the learning device interference quantization index, the comprehensive interaction anomaly evaluation parameter, the light intensity deviation value, and the noise decibel value, or the greater the fluctuation degree of the device tilt angle, the lower the multi-dimensional credibility evaluation value of the word, indicating that the learning effect of the word during practice testing is worse.

[0058] The parameters in the multi-dimensional credibility evaluation value of the word are interrelated and affect each other. For example, device performance and environmental factors will directly affect the interaction behavior pattern of learners. The present invention constructs a multi-dimensional credibility evaluation model by integrating learning environment monitoring data, learning device interference quantization index, and comprehensive interaction anomaly evaluation parameters of each word of the learner, and solves the limitation of single dependence on user behavior or device indicators in the traditional scheme. Users may maintain a pseudo-normal state due to behavioral inertia. Through non-linear correlation modeling of multi-source heterogeneous data, the combined effects of device performance degradation, environmental interference, and user behavior deviation are incorporated into a unified evaluation framework, improving the determination accuracy of the reference value of the learning effect. Especially in complex scenarios where device and environmental interference are not directly reflected in the explicit behavior of users, the robustness and anti-interference ability of the dynamic update of the learning path are ensured.

[0059] Specifically, the learner interaction data is processed to obtain the comprehensive interaction anomaly evaluation parameters for each vocabulary of the learner. The specific process is as follows: Based on the interaction data of all users in historical tasks, a group benchmark data set of vocabulary learning duration, time-consuming proportion of practice types, and interaction frequency is established. The benchmark learning duration for each vocabulary, the benchmark time-consuming proportion for each practice type, and the benchmark interaction frequency for each vocabulary are obtained by taking the average of the historical data.

[0060] It should be understood that the time-consuming proportion of practice types refers to the proportion of different practice types in the total learning duration, and the vocabulary learning interaction frequency refers to the interaction frequency between the learner and the device during the learning process of a specific vocabulary.

[0061] For the interaction data of the current learner, the deviation ratio of the learning duration of each vocabulary from the benchmark duration, the deviation amplitude of the time-consuming proportion of each practice type from the benchmark proportion, and the difference coefficient of the interaction frequency of each vocabulary learning process are respectively processed; dynamic weighting is used to aggregate the above deviations to generate the comprehensive interaction anomaly evaluation parameters for each vocabulary of the learner. The comprehensive interaction anomaly evaluation parameters for each vocabulary of the learner are used to characterize the overall deviation degree of the current learner's interaction behavior from the group standard.

[0062] In a specific embodiment, the method for obtaining the comprehensive interaction anomaly evaluation parameters for each vocabulary of the learner is as follows:

[0063] ;

[0064] In the formula, represents the comprehensive interaction anomaly evaluation parameter for the i-th vocabulary, represents the interaction anomaly evaluation weight coefficient corresponding to the preset vocabulary learning duration, represents the interaction anomaly evaluation weight coefficient corresponding to the preset time-consuming proportion of practice types, represents the interaction anomaly evaluation weight coefficient corresponding to the preset interaction frequency, represents the learning duration of the i-th vocabulary of the current learner, represents the benchmark learning duration of the i-th vocabulary, represents the allowable deviation value of the learning duration of the i-th vocabulary, represents the time-consuming proportion of the r-th practice type of the i-th vocabulary, represents the benchmark time-consuming proportion of the r-th practice type of the i-th vocabulary, represents the allowable deviation value of the time-consuming proportion of the r-th practice type of the i-th vocabulary, represents the learning interaction frequency of the i-th vocabulary, represents the benchmark interaction frequency of the i-th vocabulary learning, It represents the allowable deviation value of the learning interaction frequency of the i-th vocabulary. Here, i represents the number of each vocabulary, where i = 1, 2, 3,..., n, and n represents the total number of vocabularies. r represents the number of each exercise type, where r = 1, 2, 3,..., h, and h represents the total number of exercise types.

[0065] 、 and The value range of is between 0 and 1. A mapping relationship between different vocabularies and the weight coefficients of each interaction anomaly evaluation can be established in the learning experiment of different vocabularies. For example, a mapping relationship is established by corresponding each vocabulary to the weight coefficient of each interaction anomaly evaluation and stored in the vocabulary learning database. The weight coefficient of the interaction anomaly evaluation is selected from the vocabulary learning database through the mapping relationship.

[0066] The comprehensive interaction anomaly evaluation parameter is obtained by jointly processing the vocabulary learning duration, the time-consuming proportion of the exercise type, and the vocabulary learning interaction frequency. Among them, the greater the deviation between the vocabulary learning duration, the time-consuming proportion of the exercise type, and the vocabulary learning interaction frequency and the corresponding reference values, the greater the comprehensive interaction anomaly evaluation parameter, indicating that the interaction behavior of the learner during vocabulary learning is more abnormal.

[0067] The present invention constructs a multi-dimensional anomaly evaluation parameter by dynamically associating the deviation characteristics of the vocabulary learning duration, the time-consuming proportion of the exercise type, and the interaction frequency, and accurately captures the implicit abnormal patterns in the user's behavior. The vocabulary learning duration, the time-consuming proportion of the exercise type, and the interaction frequency are interrelated and can comprehensively reflect the user's behavior anomalies. Based on the historical group data, a reference is established, and the deviation of the vocabulary learning duration, the deviation of the exercise type distribution, and the anomaly of the interaction frequency of the current user are non-linearly aggregated. This helps to distinguish the passive increase in the interaction frequency caused by equipment performance fluctuations from the active learning behavior, reduces the misjudgment rate, and provides a high-fidelity behavior analysis basis for the update of the learning path.

[0068] The learner interaction data includes the learning duration of each vocabulary, the time-consuming proportion of each vocabulary for each exercise type, and the interaction frequency during the learning process of each vocabulary.

[0069] Specifically, it is judged whether the vocabulary learning task is completed, and the learning path is updated according to the completion situation of the vocabulary learning task. The specific process is as follows: The vocabulary learning effect includes the correct rate of each vocabulary exercise; the multi-dimensional credibility evaluation values of each vocabulary are matched with the vocabulary exercise correct rate thresholds corresponding to the preset multi-dimensional credibility evaluation value intervals in the vocabulary learning database, and the correct rate of each vocabulary exercise is compared with the correct rate threshold of each vocabulary.

[0070] If the correct rate of a certain vocabulary exercise is not less than the correct rate threshold of this vocabulary, then this vocabulary is marked as an effectively learned vocabulary; otherwise, this vocabulary is marked as an ineffectively learned vocabulary, and the number of ineffectively learned vocabularies is counted.

[0071] According to the number of ineffective learned words and the multi-dimensional credibility evaluation values of each word, the completion degree of the word learning task is processed, and the learning path is updated according to the completion degree of the word learning task.

[0072] In a specific embodiment, the method for obtaining the completion degree of the word learning task is as follows:

[0073] ;

[0074] In the formula, \(W\) represents the completion degree of the word learning task, \(e\) represents the natural constant, represents the task completion degree weight coefficient corresponding to the preset number of ineffective learned words, represents the task completion degree weight coefficient corresponding to the preset multi-dimensional credibility evaluation value of the word, \(N\) represents the number of ineffective learned words, represents the multi-dimensional credibility evaluation value of the \(i\)-th word, \(i\) represents the number of each word, \(i = 1, 2, 3, \cdots, n\), and \(n\) represents the total number of words.

[0075] and The value ranges of and are between 0 and 1, and can be obtained by matching the number of words and the word difficulty in the learning task. For example, the number of words and the word difficulty are quantified, and the quantization results are mapped one-to-one with each task completion degree weight coefficient and stored in the word learning database. The task completion degree weight coefficient is selected from the word learning database through the mapping relationship.

[0076] In this embodiment, the multi-dimensional credibility evaluation value of the word and the number of ineffective learned words affect each other. The higher the multi-dimensional credibility evaluation value of the word, the lower the degree of interference of the learning task, and the number of ineffective learned words. The present invention breaks through the traditional single correct rate threshold determination method through a composite calculation model that fuses the number of ineffective learned words and the multi-dimensional credibility evaluation values of each word, and realizes the accurate quantification of the task completion quality. At the same time, it helps to reduce the misjudgment rate of the task completion degree and provides an adaptive optimization basis for the dynamic learning path.

[0077] Specifically, updating the learning path according to the completion degree of the word learning task includes: comparing the completion degree of the word learning task with the preset word learning task completion degree threshold in the word learning database. If the completion degree of the word learning task is less than the word learning task completion degree threshold, the word learning task is marked as an unfinished learning task, and the reinforcement learning model obtains the unfinished learning task and re-performs the practice task allocation.

[0078] In a specific embodiment, the present invention analyzes the overall vocabulary learning task, breaks through the limitation of the traditional solution that only provides feedback on the learning effect of a single vocabulary, and integrates the systematic influence of environmental interference factors on the learning process from a task-level global perspective. When environmental interference leads to a decline in the overall efficiency of the learning task, it is possible to identify the cumulative damage of the environment to attention and cognitive load based on the cluster analysis of the multi-dimensional credibility evaluation values of the vocabulary, rather than attributing it in isolation to insufficient individual vocabulary mastery. At the same time, analyzing the overall vocabulary learning task helps to identify ineffective learning of the task caused by environmental or device interference and update the learning path in a timely manner.

[0079] Specifically, updating the learning path according to the completion degree of the vocabulary learning task further includes: if the completion degree of the vocabulary learning task is not less than the vocabulary learning task completion degree threshold, the reinforcement learning model assigns a review task to the vocabulary learning task, and obtains the ineffective learning vocabulary in the vocabulary learning task, and re-assigns a practice task for the ineffective learning vocabulary.

[0080] Referring to Figure 2 As shown, the second aspect of the present invention provides an intelligent learning path planning method for English vocabulary based on reinforcement learning, including: S1. Obtain the individual characteristic information of the learner and conduct a preliminary planning of the vocabulary learning path. The preliminary planning of the vocabulary learning path includes vocabulary library matching, learning starting point setting, and learning task division; S2. Real-time collect the performance data of the learner's learning device, and quantitatively process the interference degree of the learning device performance data on the learning effect to obtain a learning device interference quantification index, and judge whether the learning device performance will interfere with the learning effect and give feedback; S3. Monitor the learning environment of the learner, and real-time collect the learner interaction data, and evaluate the reference value of the detection result of the learner's vocabulary learning effect in combination with the learning device interference quantification index; S4. Obtain the vocabulary learning effect, and judge whether the vocabulary learning task is completed according to the reference value of the vocabulary learning effect detection result, and update the learning path according to the completion situation of the vocabulary learning task.

[0081] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principle and practical application of the present invention, so that those skilled in the art can understand and utilize the present invention well. As long as it does not deviate from the structure of the present invention or exceed the scope defined by the present invention, it should fall within the protection scope of the present invention.

Claims

1. An intelligent learning path planning system for English vocabulary based on reinforcement learning, characterized in that: Including: A preliminary learning path planning module, which is used to obtain the individual characteristics information of the learner and conduct a preliminary planning of the vocabulary learning path. The preliminary planning of the vocabulary learning path includes word bank matching, learning starting point setting, and learning task division; A learning device feedback module, which is used to collect the performance data of the learner's learning device in real time, quantify the interference degree of the learning device performance data on the learning effect, obtain a learning device interference quantification index, and judge whether the learning device performance will interfere with the learning effect and give feedback; A learning effect analysis module, which is used to monitor the learning environment of the learner, collect the learner's interaction data in real time, and evaluate the reference value of the vocabulary learning effect detection result of the learner in combination with the learning device interference quantification index; A learning path update module, which is used to obtain the vocabulary learning effect, judge whether the vocabulary learning task is completed according to the reference value of the vocabulary learning effect detection result, and update the learning path according to the completion situation of the vocabulary learning task; The process of quantifying the interference degree of the learning device performance data on the learning effect is as follows: for the CPU occupancy rate and the memory occupancy rate, the maximum theoretical occupancy rate is set as the reference value according to the device hardware specifications, and the actual occupancy rate is converted into a ratio relative to the reference value. For the vocabulary refresh response delay, a dynamic reference range is set based on the lowest response delay and the highest response delay of the learning device response in the historical learning task, the vocabulary refresh response delay is mapped to the standardized value within the reference range, and weighted processing is performed according to the preset device interference weight coefficient to generate a learning device interference quantification index; The acquisition method of the learning device interference quantification index is: ; ; Wherein, represents the interference quantization index of the learning device, represents the device interference weight coefficient corresponding to the preset CPU occupancy rate, represents the device interference weight coefficient corresponding to the preset memory occupancy rate, represents the device interference weight coefficient corresponding to the preset vocabulary refresh response latency, represents the current CPU occupancy rate, represents the maximum CPU occupancy rate of the device, represents the current memory occupancy rate, represents the maximum memory occupancy rate of the device, represents the normalized value of the vocabulary refresh response latency, represents the vocabulary refresh response latency, is the lowest response latency in historical tasks, is the highest response latency in historical tasks; wherein, , and have a value range of 0 to 1, reflecting the importance of different types of performance data on the interference of learning effects. The device interference weight coefficient is obtained by matching the specific device type; Based on the interaction data of all users in the historical tasks, a group benchmark data set of vocabulary learning duration, practice type time-consuming ratio, and interaction frequency is established; For the interaction data of the current learner, the deviation ratio of each vocabulary learning duration from the benchmark duration, the deviation amplitude of the time-consuming ratio of each practice type from the benchmark ratio, and the difference coefficient of the interaction frequency of each vocabulary learning process from the benchmark frequency are respectively processed; Dynamic weighting is used to aggregate the deviation ratio of each vocabulary learning duration from the benchmark duration, the deviation amplitude of the time-consuming ratio of each practice type from the benchmark ratio, and the difference coefficient of the interaction frequency of each vocabulary learning process from the benchmark frequency to generate a comprehensive interaction anomaly evaluation parameter for each vocabulary of the learner; The process of evaluating the reference value of the vocabulary learning effect detection result of the learner is as follows: By fusing the learning environment monitoring data, the learning device interference quantification index, and the comprehensive interaction anomaly evaluation parameter of each vocabulary of the learner, a multi-dimensional credibility evaluation model for each vocabulary is constructed to obtain the multi-dimensional credibility evaluation value of each vocabulary.

2. The intelligent learning path planning system for English vocabulary based on reinforcement learning according to claim 1, characterized in that: The process of conducting a preliminary planning of the vocabulary learning path is as follows: The individual characteristics information of the learner includes the learner's basic vocabulary test result, the set learning time, and the learning goal; The learning word bank is matched according to the learning goal set by the learner; The learning starting point of the learner is set according to the learner's basic vocabulary test result; The learning tasks include each vocabulary to be learned and each vocabulary practice method.

3. The intelligent learning path planning system for English vocabulary based on reinforcement learning according to claim 1, characterized in that: Quantifying the degree of interference of learning device performance data on learning effect further includes: The learning device performance data includes CPU occupancy rate, memory occupancy rate, and vocabulary refresh response latency; The learning device interference quantification index is used to quantify and evaluate the degree of interference of learning device performance on the learning effect of learners.

4. The intelligent learning path planning system for English vocabulary based on reinforcement learning according to claim 1, wherein: Judging whether the learning device performance will interfere with the learning effect and giving feedback, the specific process is: Obtain the preset learning device interference quantification index threshold from the vocabulary learning database; Compare the learning device interference quantification index with the learning device interference quantification index threshold. If the learning device interference quantification index is less than the learning device interference quantification index threshold, mark the learning device interference as negligible interference and do not include it in the evaluation of the reference value of the vocabulary learning effect detection result; If the learning device interference quantification index is not less than the learning device interference quantification index threshold, mark the learning device interference as significant interference, push a learning device interference prompt to the learner, and include the learning device interference in the evaluation of the reference value of the vocabulary learning effect detection result.

5. The intelligent learning path planning system for English vocabulary based on reinforcement learning according to claim 1, wherein: Evaluating the reference value of the vocabulary learning effect detection result of the learner further includes: Processing the learner interaction data to obtain the comprehensive interaction anomaly evaluation parameter of each vocabulary of the learner; The vocabulary multi-dimensional credibility evaluation value is used to quantify and evaluate the reference value of the vocabulary learning effect detection result; The learning environment monitoring data includes the light intensity, noise decibel value, and device tilt angle at each monitoring time point.

6. The intelligent learning path planning system for English vocabulary based on reinforcement learning according to claim 5, characterized in that: Processing the learner interaction data to obtain the comprehensive interaction anomaly evaluation parameter of each vocabulary of the learner further includes: The comprehensive interaction anomaly evaluation parameter of each vocabulary of the learner is used to characterize the overall deviation degree of the current learner interaction behavior from the group standard; The learner interaction data includes the learning duration of each vocabulary, the time-consuming proportion of each practice type of each vocabulary, and the interaction frequency during the learning process of each vocabulary.

7. The intelligent learning path planning system for English vocabulary based on reinforcement learning according to claim 1, wherein: Judging whether the vocabulary learning task is completed and updating the learning path according to the completion situation of the vocabulary learning task, the specific process is: The vocabulary learning effect includes the practice correct rate of each vocabulary; Match the multi-dimensional credibility evaluation value of each vocabulary with the vocabulary practice correct rate threshold corresponding to the preset multi-dimensional credibility evaluation value interval of each vocabulary in the vocabulary learning database, and compare the practice correct rate of each vocabulary with the vocabulary practice correct rate threshold; If the practice correct rate of a certain vocabulary is not less than the vocabulary practice correct rate threshold, mark the vocabulary as an effectively learned vocabulary, otherwise, mark the vocabulary as an ineffectively learned vocabulary, and count the number of ineffectively learned vocabularies; According to the number of ineffectively learned vocabularies and the multi-dimensional credibility evaluation value of each vocabulary, process to obtain the completion degree of the vocabulary learning task, and update the learning path according to the completion degree of the vocabulary learning task.

8. The intelligent learning path planning system for English vocabulary based on reinforcement learning according to claim 1, wherein: Updating the learning path according to the completion degree of the vocabulary learning task includes: Compare the completion degree of the vocabulary learning task with the preset vocabulary learning task completion degree threshold in the vocabulary learning database. If the completion degree of the vocabulary learning task is less than the vocabulary learning task completion degree threshold, mark the vocabulary learning task as an unfinished learning task, and the reinforcement learning model obtains the unfinished learning task and reallocates the practice task.

9. The intelligent learning path planning system for English vocabulary based on reinforcement learning according to claim 1, wherein: The learning path update based on the completion degree of the vocabulary learning task further includes: If the completion degree of the vocabulary learning task is not less than the vocabulary learning task completion degree threshold, the reinforcement learning model assigns a review task to the vocabulary learning task, obtains the ineffective learning vocabulary in the vocabulary learning task, and re-assigns a practice task for the ineffective learning vocabulary.

10. A method applied to the intelligent learning path planning system for English vocabulary based on reinforcement learning according to any one of claims 1-9, characterized in that: Including: S1. Obtain the individual characteristics information of the learner, and conduct a preliminary planning of the vocabulary learning path. The preliminary planning of the vocabulary learning path includes vocabulary library matching, learning starting point setting, and learning task division; S2. Collect the performance data of the learner's learning device in real time, and quantitatively process the interference degree of the learning device performance data on the learning effect to obtain a learning device interference quantization index, and judge whether the learning device performance will interfere with the learning effect and give feedback; S3. Monitor the learning environment of the learner, collect the learner's interaction data in real time, and evaluate the reference value of the detection result of the learner's vocabulary learning effect in combination with the learning device interference quantization index; S4. Obtain the vocabulary learning effect, judge whether the vocabulary learning task is completed according to the reference value of the vocabulary learning effect detection result, and update the learning path according to the completion situation of the vocabulary learning task.

Citation Information

Patent Citations

  • A statistical analysis-based English personal ability tracking learning system

    CN112149994B

  • Deep Learning-Based Resource Learning Path Planning Method and Apparatus

    CN112734142B

  • Intelligent language education system and method

    CN118135851A

  • Teaching equipment management control system and method

    CN118333582A

  • Intelligent multimedia interactive teaching and checking system and method

    CN119379506A