Methods, systems, equipment, and media for generating ideological and political education learning paths based on reinforcement learning.

By using reinforcement learning-based methods to obtain learners' historical interaction records, analyze their cognitive states and emotional tendencies, and generate personalized ideological and political learning paths, the problem of the lack of personalization in traditional ideological and political learning paths is solved. This achieves the coordinated advancement of cognitive bias correction and emotional resonance cultivation, thereby improving the pertinence and effectiveness of education.

CN122134515APending Publication Date: 2026-06-02JIANGXI TELLHOW ANIMATION VOCATIONAL COLLEGE

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGXI TELLHOW ANIMATION VOCATIONAL COLLEGE
Filing Date
2026-02-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Traditional ideological and political education approaches lack personalization and fail to effectively combine learners' cognitive states and emotional inclinations, resulting in low knowledge absorption efficiency, insufficient emotional resonance, and difficulty in accurately correcting cognitive biases, thus failing to meet the needs of personalized ideological and political education.

Method used

By using reinforcement learning-based methods, learners' historical interaction records are obtained, cognitive state vectors and emotional tendency vectors are analyzed, learning path sequences are generated, the recommendation order of bias-related knowledge points is adjusted, and path segments are optimized to improve long-term stability and emotional resonance by combining multi-objective optimization and reinforcement learning models.

Benefits of technology

It enables the scientific generation of personalized ideological and political learning paths, accurately corrects cognitive biases, improves knowledge efficiency and emotional resonance, meets the dual needs of ideological and political education, and enhances the pertinence and effectiveness of education.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134515A_ABST
    Figure CN122134515A_ABST
Patent Text Reader

Abstract

This application relates to a method, system, device, and medium for generating ideological and political learning paths based on reinforcement learning. The method includes: acquiring learners' historical interaction records, inputting them into a sequence model for time-series modeling, and outputting a cognitive state vector and an emotional tendency vector; then inputting these two vectors into a preset deviation detection model, outputting an emotional resonance intensity value and a cognitive deviation probability. Combining these values ​​with a preset ideological and political knowledge graph, an initial learning path sequence is generated. If the cognitive deviation probability exceeds a preset threshold, the recommended order of deviation-related knowledge points in the sequence is adjusted to obtain an adjusted learning path sequence. A multi-objective optimization algorithm is used to calculate short-term path segments from this sequence, which are then input into a reinforcement learning model to calculate long-term stability scores. The path segment with the highest score is selected as the learner's personalized ideological and political learning recommendation path. This method improves the scientific rigor and adaptability of ideological and political learning paths, effectively enhancing the implementation effect of personalized ideological and political education.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent education technology, and in particular relates to a method, system, device and medium for generating ideological and political learning paths based on reinforcement learning. Background Technology

[0002] As a core component of moral education, ideological and political education relies heavily on the scientific generation of personalized learning paths to enhance its relevance and effectiveness. With the development of intelligent education technology, artificial intelligence technologies such as reinforcement learning and knowledge graphs have been gradually applied to the field of learning path planning, providing technical support for the personalized empowerment of ideological and political learning. Traditional ideological and political learning paths often adopt a standardized planning model, neglecting the differences in learners' individual cognitive states, emotional tendencies, and potential cognitive biases. This easily leads to low knowledge absorption efficiency, insufficient emotional resonance, and difficulty in accurately correcting cognitive biases. Existing learning path generation technologies based on intelligent algorithms primarily focus on optimizing the single objective of knowledge acquisition efficiency, failing to deeply integrate the correction of cognitive biases and the cultivation of emotional resonance in ideological and political learning. They also lack quantitative consideration of the long-term stability of learning paths, failing to meet the dual needs of ideological and political learning—which combines cognitive shaping and value guidance—and thus failing to satisfy the development requirements of personalized ideological and political education. Summary of the Invention

[0003] Therefore, it is necessary to provide a reinforcement learning-based method, system, device, and medium for generating ideological and political learning paths that can achieve the synergistic advancement of cognitive bias correction, knowledge efficiency improvement, and emotional resonance cultivation, thereby enhancing the scientific nature and adaptability of ideological and political learning paths.

[0004] Firstly, this application provides a method for generating ideological and political learning paths based on reinforcement learning, including: The learner's historical interaction records are obtained, and the historical interaction records are processed through a sequence model to obtain cognitive state vectors and emotional tendency vectors.

[0005] Based on the cognitive state vector and the emotional tendency vector, the intensity value of emotional resonance and the probability of cognitive deviation are obtained through the deviation detection model.

[0006] A learning path sequence is generated based on the probability of cognitive bias, the intensity of emotional resonance, and a preset knowledge graph. If the probability of cognitive bias exceeds a preset threshold, the recommended order of knowledge points related to the bias in the learning path sequence is adjusted.

[0007] The learning path sequence corresponding to the recommended order of knowledge points is optimized by multi-objective calculation to obtain short-term path segments. The short-term path segments are then input into a reinforcement learning model to calculate long-term stability scores, and the segments with the highest scores are selected as personalized ideological and political learning recommendation paths for learners.

[0008] In one embodiment, historical interaction records are processed using a sequence model to obtain cognitive state vectors and sentiment tendency vectors, including: Obtain learners' historical interaction records; these records include learners' behavior logs and feedback information.

[0009] The cognitive association information in the historical interaction records is input into the sequence model for analysis to obtain the learner's cognitive state vector; the sequence model is constructed by a deep learning algorithm.

[0010] Based on the cognitive state vector and feedback information from historical interaction records, emotional association features, behavioral feedback features, and cognitive state vectors are extracted from the feedback information. After unified dimensional processing and weighted fusion, fused emotional features are obtained.

[0011] The learner's emotional tendency vector is determined by calculating the emotional score of each sub-feature based on the fusion of emotional features and adjustment coefficients.

[0012] In one embodiment, the sentiment score is calculated using the following formula: in, Indicates the first Sentiment scores for individual characteristics, Indicates the first Normalized eigenvalues ​​of individual features This represents the weight associated with cognitive states, and its value range is... Based on the degree of influence of cognitive state vectors on emotional tendencies, This indicates the weight of feedback timeliness, with a value range of... , Indicates the first The matching coefficient of knowledge mastery corresponding to each sub-feature Indicates the first The timeliness coefficient of feedback information corresponding to each sub-feature The sentiment correction factor is determined based on the semantic sentiment intensity of the feedback text, with positive semantic correspondence values. Neutral semantics correspond to values Negative semantics correspond to values , This represents the total number of sub-features in the fused emotional features.

[0013] In one embodiment, based on the cognitive state vector and the emotional tendency vector, the emotional resonance intensity value and the probability of cognitive bias are obtained through a deviation detection model, including: The knowledge mastery, understanding depth, and cognitive logical correlation information in the cognitive state vector are compared and analyzed with the preset ideological and political knowledge point standards. Potential cognitive bias tendency dimensions and bias correlation degree are extracted as cognitive bias correlation features.

[0014] The intensity values ​​of each emotional dimension in the emotional tendency vector are extracted, and the emotional resonance intensity value, which represents the learner's emotional fit with the ideological and political content, is calculated using a weighted summation algorithm.

[0015] A pre-defined deviation detection model is used to extract features of cognitive deviation-related features and calculate deviation tendency feature values; the deviation detection model is constructed based on the standard cognitive framework of ideological and political knowledge points.

[0016] Based on the bias tendency feature value and the sentiment tendency vector, a weighted fusion algorithm is used to calculate the comprehensive bias influence factor.

[0017] The comprehensive deviation impact factor is compared with the preset cognitive deviation threshold. If it exceeds the cognitive deviation threshold, the comprehensive deviation impact factor is mapped to the corresponding cognitive deviation probability.

[0018] In one embodiment, a learning path sequence is generated based on the probability of cognitive bias, the intensity of emotional resonance, and a preset knowledge graph. If the probability of cognitive bias exceeds a preset threshold, the recommended order of knowledge points related to the bias in the learning path sequence is adjusted, including: Based on the probability of cognitive bias, the intensity of emotional resonance, and the pre-set knowledge graph, a sequence of candidate learning paths is generated; the knowledge graph includes the logical relationships between ideological and political knowledge points.

[0019] If the probability of cognitive bias exceeds the preset cognitive bias threshold, the priority order of the bias-related knowledge points in the candidate learning path sequence will be adjusted.

[0020] Based on the adjusted priority order, an optimized learning path sequence is generated.

[0021] Extract the association feature values ​​of key knowledge points in the optimized learning path sequence; the association feature values ​​include the dependency relationship feature values ​​between knowledge points and the deviation correction and adaptation feature values.

[0022] Based on the associated feature value, the probability of cognitive bias, and the intensity of emotional resonance, a dynamic weight model is constructed. The bias correction weight of each knowledge point is calculated through the dynamic weight model to obtain the recommended order of knowledge points after bias correction.

[0023] In one embodiment, the personalized ideological and political learning recommendation path for learners is obtained through the following steps: The learning path sequence corresponding to the recommended order of knowledge points is optimized by multi-objective calculation of short-term path segments; the short-term path segments are ordered sequences of knowledge points that take into account both knowledge efficiency and emotional resonance.

[0024] Input short-term path segments into a long-term growth simulator. By simulating the entire learning process of learners according to short-term path segments, output the emotional state vector and knowledge mastery vector corresponding to each knowledge point learning step.

[0025] The emotional state vector and knowledge mastery vector corresponding to each step are input into the reinforcement learning model. By iteratively calculating the cumulative value of each step's state, the long-term stability score of the short-term path segment is output.

[0026] Determine whether the long-term stability score has reached the preset stability threshold. If not, adjust the order of knowledge points in the short-term path segment or supplement bias correction knowledge points based on the state value deviation output during the reinforcement learning model iteration process, and generate the adjusted path segment.

[0027] The adjusted path segment is input into the long-term growth simulator again, and iterative calculation is performed through the reinforcement learning model until the number of iterations reaches the preset iteration threshold or the long-term stability score of a certain path segment meets the stability threshold. The path segment with the highest long-term stability score is selected from all the path segments obtained from the iterations.

[0028] Output the path segment with the highest long-term stability score as a personalized ideological and political learning recommendation path for learners.

[0029] In one embodiment, the long-term stability score is calculated using the following formula: in, This represents the long-term stability score of short-term path segments. This represents the total number of knowledge points in a short-term path segment. Indicates the first Importance weight of each knowledge point, range of values Based on the core importance of knowledge points within the ideological and political knowledge system and the priority of deviation correction, Represents the sentiment co-weight, with a value range of The percentage of the impact of emotional resonance on long-term learning stability. Indicates the first The overall strength of the emotional state vector. Indicates the first The pass rate of the knowledge mastery degree vector. This represents the time decay coefficient, with a range of values. Based on the presupposition of the long-term memory retention pattern of ideological and political knowledge.

[0030] Secondly, this application also provides a reinforcement learning-based system for generating ideological and political learning paths, the system including: The interaction record module is used to acquire learners' historical interaction records. The historical interaction records are processed through a sequence model to obtain cognitive state vectors and emotional tendency vectors.

[0031] The deviation detection module is used to analyze the emotional resonance intensity value and the probability of cognitive deviation based on the cognitive state vector and the emotional tendency vector through the deviation detection model.

[0032] The path generation module is used to generate a learning path sequence based on the probability of cognitive bias, the intensity of emotional resonance, and a preset knowledge graph. If the probability of cognitive bias exceeds a preset threshold, the recommended order of knowledge points related to the bias in the learning path sequence is adjusted.

[0033] The path optimization module is used to perform multi-objective optimization calculations on the learning path sequence corresponding to the recommended order of knowledge points to obtain short-term path segments. The short-term path segments are then input into the reinforcement learning model to calculate the long-term stability score, and the segment with the highest score is selected as the personalized ideological and political learning recommendation path for learners.

[0034] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0035] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned method.

[0036] The aforementioned reinforcement learning-based method, system, computer equipment, and storage medium for generating ideological and political learning paths acquire learners' historical interaction records, including behavioral logs and feedback information. These records are then input into a sequence model for feature extraction and analysis. The sequence model performs temporal modeling on cognitive and emotional association information within the historical interaction records, outputting the learner's cognitive state vector and emotional tendency vector. Based on these vectors, a preset deviation detection model is input. This model extracts cognitive deviation-related features from the cognitive state vector and emotional resonance features from the emotional tendency vector, performs fusion analysis, and outputs an emotional resonance intensity value and a cognitive deviation probability. Based on the cognitive deviation probability, the emotional resonance intensity value, and a preset ideological and political knowledge graph, an initial learning path sequence is generated. It is then determined whether the cognitive deviation probability exceeds a preset cognitive deviation threshold. If it does, the recommended order of knowledge points related to the cognitive deviation in the initial learning path sequence is adjusted to obtain an adjusted learning path sequence. For the learning path sequence corresponding to the adjusted knowledge point recommendation order, a multi-objective optimization algorithm is used to calculate short-term path segments that balance knowledge efficiency and emotional resonance. These short-term path segments are then input into a reinforcement learning model, which iteratively calculates the long-term stability score of the short-term path segment based on the emotional state vector and knowledge mastery vector corresponding to the short-term path segment. The path segment with the highest long-term stability score among all iterative calculations is selected as the learner's personalized ideological and political learning recommendation path. This method achieves personalized generation of ideological and political learning paths through coherent data flow processing, effectively solving the problems of uniformity and lack of specificity in traditional ideological and political learning paths. It accurately mines learners' cognitive states and emotional tendencies through sequence models, and combines these with a deviation detection model to achieve quantitative analysis of cognitive biases and emotional resonance, ensuring that the generated paths are tailored to individual learner differences. By driving the recommended order of knowledge points with the probability of cognitive bias, and combining multi-objective optimization and reinforcement learning models to consider the short-term adaptability and long-term stability of the paths, it achieves the synergistic advancement of cognitive bias correction, knowledge efficiency improvement, and emotional resonance cultivation. This enhances the scientific nature and adaptability of ideological and political learning paths, meets the dual needs of ideological and political education in both cognitive shaping and value guidance, and effectively improves the implementation effect of personalized ideological and political education. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 A flowchart illustrating the method for generating ideological and political learning paths based on reinforcement learning provided in this embodiment of the invention; Figure 2 The structural block diagram of the reinforcement learning-based ideological and political learning path generation system provided in the embodiments of the present invention is shown. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0040] In one embodiment, such as Figure 1 As shown, this application provides a method for generating ideological and political learning paths based on reinforcement learning, which may include the following steps: Step S101: Obtain the learner's historical interaction records, and process the historical interaction records using a sequence model to obtain cognitive state vectors and emotional tendency vectors.

[0041] Specifically, learners' historical interaction records are acquired, and these records are processed using a sequence model to obtain cognitive state vectors and affective tendency vectors. These historical interaction records include various behavioral logs and feedback information generated by learners during their ideological and political education. Behavioral logs cover knowledge point clicks, learning duration, and quiz performance, while feedback information includes evaluations of the learning content and questions. These historical interaction records are used as input to the sequence model, which extracts and models the temporal features of the cognitive and affective information within the records. Through temporal analysis of learners' long-term learning behavior and feedback, the model accurately identifies cognitive features such as the learner's knowledge point mastery and comprehension depth, as well as the intensity features of various dimensions of affective tendency. The final output is a cognitive state vector that comprehensively represents the learner's current cognitive level, and an affective tendency vector that accurately reflects the learner's emotional attitude towards the ideological and political content.

[0042] Step S102: Based on the cognitive state vector and the emotional tendency vector, the emotional resonance intensity value and the cognitive bias probability are obtained through the deviation detection model analysis.

[0043] Based on the cognitive state vector and the emotional tendency vector, an emotional resonance intensity value and a cognitive deviation probability are obtained through deviation detection model analysis. The cognitive state vector and the emotional tendency vector are jointly input into a pre-defined deviation detection model, which is constructed based on the standard cognitive framework of ideological and political knowledge points. First, cognitive deviation-related features that differ from the standard cognition of ideological and political knowledge points are extracted from the cognitive state vector. Then, emotional resonance features reflecting the learner's emotional fit with the ideological and political content are extracted from the emotional tendency vector. Subsequently, the two types of features are fused and analyzed, and the emotional resonance intensity value and the cognitive deviation probability are obtained through quantitative calculation. The emotional resonance intensity value characterizes the learner's emotional fit with the ideological and political learning content, while the cognitive deviation probability characterizes the likelihood that the learner's current cognitive state deviates from the standard cognition of ideological and political knowledge points.

[0044] Step S103: Generate a learning path sequence based on the probability of cognitive bias, the intensity of emotional resonance, and the preset knowledge graph. If the probability of cognitive bias exceeds the preset threshold, adjust the recommended order of knowledge points related to the bias in the learning path sequence.

[0045] Based on cognitive bias probability and emotional resonance intensity, and using a pre-set ideological and political knowledge graph as the data foundation, this study generates an initial ideological and political learning path sequence that is adapted to the learner's current cognitive and emotional state by utilizing the logical connections, hierarchical relationships, and bias correction rules of ideological and political knowledge points in the knowledge graph. Simultaneously, a pre-set cognitive bias threshold is retrieved, and the calculated cognitive bias probability is compared with this threshold. If the cognitive bias probability exceeds the pre-set threshold, the recommended order of knowledge points in the initial learning path sequence that directly correspond to the cognitive bias is adjusted, prioritizing bias correction knowledge points, and finally obtaining the adjusted ideological and political learning path sequence.

[0046] Step S104: Perform multi-objective optimization calculation on the learning path sequence corresponding to the recommended order of knowledge points to obtain short-term path segments. Input the short-term path segments into the reinforcement learning model to calculate the long-term stability score and select the segment with the highest score as the personalized ideological and political learning recommendation path for learners.

[0047] For the ideological and political learning path sequence corresponding to the adjusted knowledge point recommendation order, a multi-objective optimization algorithm is used for calculation, with the dual optimization objectives of improving knowledge learning efficiency and emotional resonance, to generate short-term ideological and political learning path segments that take both objectives into account. Then, the short-term path segments are input into a preset reinforcement learning model. The reinforcement learning model combines the emotional state vector and knowledge mastery vector corresponding to the learning process of the short-term path segments, and obtains the long-term stability score of each short-term path segment through multiple rounds of iteration calculation. This score represents the ability of the path to maintain the knowledge mastery effect and emotional resonance state during long-term learning. Finally, from all iterative calculation results, the path segment with the highest long-term stability score is selected as the personalized ideological and political learning recommendation path for the learner.

[0048] The aforementioned reinforcement learning-based method for generating ideological and political learning paths involves acquiring historical interaction records containing learner behavior logs and feedback information. These records are then input into a sequence model for time-series modeling, extracting and analyzing cognitive and emotional correlation information, and outputting cognitive state vectors and emotional tendency vectors. These two vectors are then input into a pre-defined deviation detection model to extract and fuse cognitive deviation correlation features and emotional resonance features, outputting an emotional resonance intensity value and a cognitive deviation probability. Combining these values ​​with a pre-defined ideological and political knowledge graph, an initial learning path sequence is generated. If the cognitive deviation probability exceeds a pre-defined threshold, the recommended order of deviation-related knowledge points in the sequence is adjusted, resulting in an adjusted learning path sequence. A multi-objective optimization algorithm is then applied to this sequence to obtain short-term path segments that balance knowledge efficiency and emotional resonance. These short-term path segments are input into a reinforcement learning model, and a long-term stability score is iteratively calculated using the corresponding emotional state vector and knowledge mastery vector. The path segment with the highest score is selected as the learner's personalized ideological and political learning recommendation path. This method achieves personalized generation of ideological and political learning paths through coherent data flow processing, solving the problems of uniformity and lack of specificity in traditional paths. Relying on sequence models and deviation detection models, it accurately quantifies learners' cognitive states, emotional tendencies, and cognitive biases, ensuring that path generation is tailored to individual differences. By driving path adjustment with the probability of cognitive bias, and combining multi-objective optimization and reinforcement learning models to balance the short-term adaptability and long-term stability of paths, it achieves the synergistic advancement of cognitive bias correction, knowledge efficiency improvement, and emotional resonance cultivation. This enhances the scientific nature and adaptability of ideological and political learning paths, meets the dual needs of cognitive shaping and value guidance in ideological and political education, and effectively improves the implementation effect of personalized ideological and political education.

[0049] In one embodiment, processing historical interaction records using a sequence model to obtain cognitive state vectors and emotional tendency vectors may include the following steps: Step S201: Obtain the learner's historical interaction records; the historical interaction records include the learner's behavior logs and feedback information.

[0050] Step S202: Input the cognitive association information in the historical interaction records into the sequence model for analysis to obtain the learner's cognitive state vector; the sequence model is constructed by a deep learning algorithm.

[0051] Step S203: Based on the cognitive state vector and feedback information in historical interaction records, extract the emotional association features, behavioral feedback features and cognitive state vector from the feedback information, and obtain the fused emotional features through unified dimension processing and weighted fusion.

[0052] Step S204: Calculate the sentiment score of each sub-feature based on the fused sentiment features and adjustment coefficients, and classify and determine the learner's sentiment tendency vector.

[0053] Specifically, the process involves acquiring learners' historical interaction records, which include behavioral logs and feedback information generated during ideological and political education. Behavioral logs record learners' learning behavior trajectories, while feedback information reflects their learning attitudes and feelings. Cognitive association information related to cognition is extracted from these historical interaction records and input into a sequence model constructed using deep learning algorithms for feature extraction and analysis. The sequence model performs temporal modeling and quantification of the cognitive association information, outputting a cognitive state vector that characterizes the learner's current cognitive level. Based on this cognitive state vector, and combined with the feedback information from the aforementioned historical interaction records, emotional association features and behavioral feedback features are further extracted from the feedback information. These are then incorporated into the obtained cognitive state vector, and the three types of features are processed with a unified dimension to eliminate dimensional differences between different types of features. A weighted fusion algorithm is then used to fuse the processed features to obtain a fused emotional feature. Based on this fused emotional feature, and combined with a preset adjustment coefficient, the emotional score of each sub-feature in the fused emotional feature is calculated. The emotional scores of each sub-feature are then categorized and integrated according to a preset emotional dimension, ultimately determining an emotional tendency vector that comprehensively reflects the learner's emotional inclination.

[0054] This embodiment achieves accurate generation of learner cognitive state vectors and emotional tendency vectors through clear, layered data flow processing, providing reliable data support for personalized planning of subsequent ideological and political learning paths. By extracting cognitive association information separately and processing it using a sequence model constructed with deep learning algorithms, it ensures that the cognitive state vector accurately reflects the learner's cognitive level, avoiding interference from irrelevant information. Simultaneously extracting emotional association features and behavioral feedback features and fusing them with the cognitive state vector overcomes the limitations of single emotional features. Unified dimensional processing ensures the rationality and accuracy of feature fusion, while the introduction of adjustment coefficients enhances the targeting of emotional score calculations for each sub-feature, making the final generated emotional tendency vector more closely match the learner's actual emotional state. The entire process is logically rigorous and the data flow is coherent, effectively improving the accuracy of cognitive and emotional feature extraction and avoiding feature omissions and dimensional confusion.

[0055] In one embodiment, the emotion score can be calculated using the following formula: in, Indicates the first Sentiment scores for individual characteristics, Indicates the first Normalized eigenvalues ​​of individual features This represents the weight associated with cognitive states, and its value range is... Based on the degree of influence of cognitive state vectors on emotional tendencies, This indicates the weight of feedback timeliness, with a value range of... , Indicates the first The matching coefficient of knowledge mastery corresponding to each sub-feature Indicates the first The timeliness coefficient of feedback information corresponding to each sub-feature The sentiment correction factor is determined based on the semantic sentiment intensity of the feedback text, with positive semantic correspondence values. Neutral semantics correspond to values Negative semantics correspond to values , This represents the total number of sub-features in the fused emotional features.

[0056] Preferably, The calculation was performed using the min-max normalization algorithm, and the formula is as follows: ,in For the first The original feature values ​​of each sub-feature. To integrate the original feature value set of all sub-features in the sentiment feature, , These are the minimum and maximum values ​​in the set, respectively. The calculation formula is: ,in For the cognitive state vector and the first Mastery of knowledge points corresponding to each sub-feature To determine the relevance of feedback information to this knowledge point, For all sub-features corresponding to gather, The maximum value in the set; The calculation formula is: ,in This refers to the time difference between the generation of feedback information and the corresponding learning behavior for the knowledge point. This is the timeliness benchmark coefficient (value range: 0.95-1.0). This is the aging attenuation coefficient (value range: 0.01-0.05).

[0057] This embodiment achieves precise quantification of the sentiment scores for each sub-feature. The formula organically combines the normalized values ​​of sub-features with cognitive state association weights, feedback timeliness weights, knowledge point mastery matching coefficients, feedback timeliness coefficients, and sentiment tendency correction factors. This achieves multi-dimensional linkage between cognitive state, feedback timeliness, semantic sentiment intensity, and the sub-features themselves, avoiding the limitations of single factors in sentiment score calculation. Simultaneously, through the cumulative normalization of the denominator, the sentiment scores of each sub-feature are placed within a unified quantification range, facilitating subsequent sub-feature classification and integration. The value ranges of each parameter are calibrated based on actual application scenarios, ensuring the feasibility of the formula. It accurately adapts to the quantification needs of learners' sentiment tendencies, improving the rationality and accuracy of sentiment score calculation. This provides precise data support for the subsequent determination of sentiment tendency vectors, thereby ensuring the synergistic effect of cognitive state vectors and sentiment tendency vectors.

[0058] In one embodiment, the emotional resonance intensity value and cognitive bias probability are obtained by analyzing the cognitive state vector and emotional tendency vector using a deviation detection model, which may include the following steps: Step S301: Compare and analyze the knowledge mastery, understanding depth, and cognitive logical correlation information in the cognitive state vector with the preset ideological and political knowledge point standards, and extract the potential cognitive bias tendency dimension and bias correlation degree as cognitive bias correlation features.

[0059] Step S302: Extract the intensity values ​​of each emotional dimension in the emotional tendency vector, and use a weighted summation algorithm to calculate the emotional resonance intensity value that represents the learner's emotional fit with the ideological and political content.

[0060] Preferably, a weighted summation algorithm is used to calculate the emotional resonance intensity value, which represents the learner's emotional fit with the ideological and political content. The emotional tendency vector includes multiple emotional dimensions such as the learner's interest in the ideological and political content, acceptance level, and willingness to continue learning. Each emotional dimension corresponds to a unique intensity value, which is quantified by the fused emotional features mentioned above and falls within the 0-1 range. After extracting the intensity values ​​corresponding to each emotional dimension, a preset weighted summation algorithm is used for calculation. The weights in the algorithm are pre-calibrated based on the degree of influence of each emotional dimension on the emotional fit of the ideological and political content, ensuring that the weight allocation aligns with the ideological and political learning scenario. The result obtained through weighted summation is the emotional resonance intensity value, which is used to quantify the learner's emotional fit with the ideological and political learning content.

[0061] Step S303: Use a preset deviation detection model to extract features from cognitive deviation-related features and calculate deviation tendency feature values; the deviation detection model is constructed based on the standard cognitive framework of ideological and political knowledge points.

[0062] The cognitive bias correlation features are the potential cognitive bias tendency dimensions and bias correlation degrees extracted after comparative analysis, which are used as input to the preset bias detection model. This bias detection model is built on the standard cognitive framework of ideological and political knowledge points. The framework includes the standard cognitive dimensions, cognitive logic, and bias judgment rules for each ideological and political knowledge point. The model further mines the core features of cognitive bias by performing feature extraction operations such as bias dimension matching and difference quantification on the input cognitive bias correlation features. Then, it calculates the extracted core features using a preset quantification algorithm, and finally outputs a bias tendency feature value, which is used to characterize the degree of learner's cognitive bias tendency.

[0063] Step S304: Based on the deviation tendency feature value and the sentiment tendency vector, a weighted fusion algorithm is used to calculate the comprehensive deviation influence factor.

[0064] Step S305: Compare the comprehensive deviation impact factor with the preset cognitive deviation threshold. If it exceeds the cognitive deviation threshold, map the comprehensive deviation impact factor to the corresponding cognitive deviation probability.

[0065] Specifically, the cognitive state vector is used to extract information on knowledge mastery, comprehension depth, and cognitive logical connections. These three types of information are then compared and analyzed against pre-defined standards for ideological and political knowledge points. Through difference identification and dimensional decomposition, potential cognitive bias tendencies (i.e., cognitive dimensions that deviate from the standards for ideological and political knowledge points) and bias correlations (i.e., the degree of correlation between each cognitive dimension and the bias) are extracted and used as cognitive bias correlation features. Simultaneously, the intensity values ​​corresponding to each emotional dimension in the emotional tendency vector are extracted. A pre-defined weighted summation algorithm is used to calculate the intensity values ​​of each emotional dimension. The weights of the weighted summation algorithm are based on the degree of influence of each emotional dimension on the emotional relevance of the ideological and political content. Finally, an emotional resonance intensity value that characterizes the learner's emotional relevance to the ideological and political content is obtained. A pre-defined bias detection model is used to extract and quantify the extracted cognitive bias correlation features, calculating the bias tendency feature values. This bias detection model is constructed based on the cognitive framework of the standards for ideological and political knowledge points, ensuring that feature extraction and bias calculation align with the needs of ideological and political cognitive evaluation. Based on the calculated deviation tendency feature values ​​and sentiment tendency vectors, a weighted fusion algorithm is used to perform fusion calculations to obtain a comprehensive deviation influence factor. The weights of the weighted fusion algorithm are calibrated according to the contribution of each feature to ideological and political cognitive deviation. A preset cognitive deviation threshold is retrieved, and the comprehensive deviation influence factor is numerically compared with this threshold. If the comprehensive deviation influence factor exceeds the preset cognitive deviation threshold, a preset mapping algorithm is used to map the comprehensive deviation influence factor to the corresponding cognitive deviation probability, thus completing the quantitative determination of cognitive deviation.

[0066] This embodiment achieves precise quantification of cognitive bias correlation features, emotional resonance intensity values, and cognitive bias probabilities, providing a reliable quantitative basis for subsequent learning path adjustments. By comparing the cognitive state vector with preset ideological and political knowledge point standards, the relevance of cognitive bias correlation features is ensured, avoiding interference from irrelevant biases. Simultaneously, the intensity values ​​of the emotional dimension are extracted and weighted summed, achieving a scientific quantification of emotional resonance intensity. Based on a deviation detection model constructed using a cognitive framework of ideological and political knowledge point standards, the adaptability and accuracy of deviation tendency feature value calculation are guaranteed. The application of a weighted fusion algorithm enables the synergistic consideration of deviation and emotional tendencies, allowing the comprehensive deviation influence factor to fully reflect the actual impact of cognitive bias. Through threshold comparison and factor mapping, precise determination of cognitive bias probabilities is achieved. The entire process is logically rigorous and the data flow is coherent, effectively improving the accuracy and rationality of cognitive bias and emotional resonance quantification, and overcoming the limitations of single-dimensional quantification.

[0067] In one embodiment, a learning path sequence is generated based on the probability of cognitive bias, the intensity of emotional resonance, and a preset knowledge graph. If the probability of cognitive bias exceeds a preset threshold, the recommended order of knowledge points related to the bias in the learning path sequence is adjusted. This may include the following steps: Step S401: Based on the probability of cognitive bias, the intensity of emotional resonance, and the preset knowledge graph, generate a sequence of candidate learning paths; the knowledge graph includes the logical relationships between ideological and political knowledge points.

[0068] Preferably, the knowledge graph includes logical relationships between ideological and political knowledge points, serving as the core framework for generating candidate learning path sequences. This ensures that the arrangement of knowledge points in the sequence conforms to the hierarchical and dependent relationships inherent in the ideological and political knowledge system. The cognitive bias probability is the quantitative result of learner cognitive bias obtained through the comprehensive bias influence factor mapping mentioned earlier. It is used to locate the knowledge point domains where learners have cognitive biases and to screen ideological and political knowledge points that meet the bias correction requirements. The emotional resonance intensity value is the quantified value of learners' emotional fit with ideological and political content obtained through the weighted summation algorithm mentioned earlier. This value is used to screen knowledge points that resonate with learners' emotional states and easily evoke emotional resonance. Based on the logical relationships between ideological and political knowledge points in the preset knowledge graph, the bias correction requirements corresponding to the cognitive bias probability and the emotional adaptation requirements corresponding to the emotional resonance intensity value are incorporated. Ideological and political knowledge points that match learners' cognitive and emotional states are then selected and arranged in an orderly manner according to the logical relationships between the knowledge points, ultimately generating a candidate learning path sequence.

[0069] Step S402: If the probability of cognitive bias exceeds the preset cognitive bias threshold, the priority order of the bias-related knowledge points in the candidate learning path sequence corresponding to the cognitive bias is adjusted.

[0070] Step S403: Generate an optimized learning path sequence based on the adjusted priority order.

[0071] Step S404: Extract the association feature values ​​of key knowledge points in the optimized learning path sequence; the association feature values ​​include the dependency relationship feature values ​​between knowledge points and the deviation correction and adaptation feature values.

[0072] Step S405: Based on the associated feature value, cognitive bias probability and emotional resonance intensity value, a dynamic weight model is constructed. The deviation correction weight of each knowledge point is calculated through the dynamic weight model to obtain the recommended order of knowledge points after deviation correction.

[0073] Specifically, candidate learning path sequences are generated based on cognitive bias probability, emotional resonance intensity, and a preset knowledge graph. The preset knowledge graph contains logical relationships between ideological and political knowledge points, supporting the rational generation of the path sequences. First, a preset cognitive bias threshold is retrieved, and the cognitive bias probability is compared with this threshold. If the cognitive bias probability exceeds the preset threshold, the priority order of the bias-related knowledge points directly corresponding to the cognitive bias in the candidate learning path sequences is adjusted. Based on the adjusted priority order of the bias-related knowledge points, the original candidate learning path sequences are reconstructed to generate... An optimized learning path sequence is generated. Relevant feature values ​​of key knowledge points are extracted from the optimized learning path sequence. These feature values ​​specifically include inter-knowledge point dependency features representing the logical connection between knowledge points, and deviation correction adaptation features representing the adaptability of knowledge points to cognitive bias correction. Using the extracted relevant feature values ​​as the core, and combining cognitive bias probability and emotional resonance intensity values, a dynamic weight model is constructed. This model is used to quantify each knowledge point in the optimized learning path sequence, obtaining the deviation correction weight corresponding to each knowledge point. Finally, based on the distribution of the deviation correction weights, the recommended order of knowledge points after deviation correction is determined.

[0074] This embodiment achieves full-process quantitative processing from candidate learning path sequence generation to the output of knowledge point recommendation order after deviation correction through multi-dimensional data linkage and progressive processing logic. The data flow is coherent and clearly directional, providing a precise basis for knowledge point ranking for the final generation of personalized ideological and political learning paths. The priority of deviation-related knowledge points is adjusted based on whether the probability of cognitive deviation exceeds a threshold, ensuring the pertinence of path optimization and effectively focusing on the core needs of cognitive deviation correction. By extracting two types of related feature values, the logical correlation rules of ideological and political knowledge points themselves are taken into account, as well as the actual needs of cognitive deviation correction, so that the subsequent model construction has solid feature data support. The dynamic weight model integrates multi-dimensional data such as related feature values, cognitive deviation probability, and emotional resonance intensity value to calculate deviation correction weights. This makes the knowledge point recommendation order not only fit the learner's current cognitive deviation status, but also take into account their emotional resonance state, improving the scientificity and adaptability of the knowledge point recommendation order, effectively ensuring the correction effect of subsequent learning paths on cognitive deviation, and meeting the implementation needs of personalized ideological and political education.

[0075] In one embodiment, the personalized ideological and political learning recommendation path for learners is obtained through the following steps, which may include the following steps: Step S501: Perform multi-objective optimization calculation on the learning path sequence corresponding to the recommended order of knowledge points to calculate short-term path segments; the short-term path segments are ordered knowledge point sequences that take into account both knowledge efficiency and emotional resonance.

[0076] Preferably, the learning path sequence corresponding to the recommended order of knowledge points is the sequence generated and adjusted above based on the probability of cognitive bias, the intensity of emotional resonance, and a preset knowledge graph. It includes ordered knowledge points that adapt to the learner's cognitive bias correction and align with their emotional state. A preset multi-objective optimization algorithm is used to calculate this learning path sequence, with knowledge learning efficiency and improved emotional resonance as the two core optimization objectives. Knowledge efficiency is quantified by the adaptability of knowledge point absorption difficulty and the logical coherence between knowledge points, while emotional resonance is quantified by the adaptability of knowledge points to the learner's emotional inclination. The algorithm segments and optimizes the path sequence to ultimately generate short-term path segments.

[0077] Step S502: Input the short-term path segments into the long-term growth simulator. By simulating the entire learning process of learners according to the short-term path segments, output the emotional state vector and knowledge mastery vector corresponding to each knowledge point learning step.

[0078] Furthermore, the long-term growth simulator pre-defines learner learning behavior models, knowledge absorption models, and emotional change models, enabling it to simulate the cognitive and emotional evolution patterns of learners during the learning process of different knowledge points. Using generated short-term path segments as input, the simulator simulates the entire process of ideological and political learning step-by-step according to the order of knowledge points within the segments. Simultaneously, it collects data on learners' emotional changes and knowledge mastery after each step of knowledge point learning. The simulator quantifies and maps these two types of data, outputting an emotional state vector and a knowledge mastery vector corresponding to each step of knowledge point learning. The emotional state vector represents the learner's emotional engagement and learning willingness after that step, while the knowledge mastery vector represents the cognitive depth and level of understanding of the knowledge point at that step.

[0079] Step S503: Input the emotional state vector and knowledge mastery vector corresponding to each step into the reinforcement learning model, and output the long-term stability score of the short-term path segment by iteratively calculating the cumulative value of the state at each step.

[0080] Step S504: Determine whether the long-term stability score has reached the preset stability threshold. If it has not, adjust the order of knowledge points in the short-term path segment or supplement bias correction knowledge points based on the state value deviation output during the reinforcement learning model iteration process, and generate the adjusted path segment.

[0081] Step S505: Input the adjusted path segment back into the long-term growth simulator and iterate through the reinforcement learning model until the number of iterations reaches the preset iteration threshold or the long-term stability score of a certain path segment meets the stability threshold. Select the path segment with the highest long-term stability score from all the path segments obtained from the iterations.

[0082] Step S506: Output the path segment with the highest long-term stability score as a personalized ideological and political learning recommendation path for learners.

[0083] Specifically, a multi-objective optimization algorithm is used to calculate the learning path sequence corresponding to the recommended order of knowledge points, generating short-term path segments. These short-term path segments are ordered sequences of knowledge points that balance knowledge learning efficiency and emotional resonance, ensuring that the segments conform to the learner's knowledge absorption patterns and align with their emotional state. The generated short-term path segments are input into a long-term growth simulator, which simulates the learner's entire process of learning ideological and political education according to these short-term path segments. The simulator synchronously collects the emotional changes and knowledge mastery status after each step of knowledge point learning, outputting the emotional state vector and knowledge mastery vector corresponding to each step, providing time-series data support for subsequent long-term stability analysis. The emotional state vector and knowledge mastery vector corresponding to each step are used as joint inputs into a pre-defined reinforcement learning model. The model iterates through multiple rounds of calculations on the learning state at each step, quantifying the cumulative value of each step's state, and then outputs the long-term stability score of the short-term path segment. This score characterizes the path segment's ability to maintain knowledge mastery and emotional resonance during long-term learning. A preset stability threshold is retrieved to determine whether the long-term stability score of a short-term path segment reaches this threshold. If not, based on the state value deviation output during the reinforcement learning model iteration, the order of knowledge points in the short-term path segment is adjusted, or appropriate deviation correction knowledge points are added to generate an adjusted path segment. The adjusted path segment is then input into the long-term growth simulator again, and the above process of simulation learning, vector output, and reinforcement learning iteration calculation is repeated until the number of iterations reaches a preset iteration threshold, or the long-term stability score of a path segment in a certain round meets the preset stability threshold, at which point the iteration stops. From the path segments obtained in all iterations, the path segment with the highest long-term stability score is selected and finally output as a personalized ideological and political learning recommendation path for the learner.

[0084] This embodiment provides learners with a scientifically tailored ideological and political education learning path, effectively addressing the problems of traditional paths lacking long-term stability considerations and insufficient adaptability. Short-term path segments are generated through a multi-objective optimization algorithm, ensuring that the path balances knowledge efficiency and emotional resonance, and is tailored to individual learner differences. A long-term growth simulator accurately simulates the entire learning process, and the output temporally sequenced emotional and knowledge vectors provide reliable data support for the iterative calculation of the reinforcement learning model, ensuring the accuracy of long-term stability scores. Through multiple rounds of iterative adjustments and selections, the final recommended path segments are ensured to possess optimal long-term stability, helping learners efficiently absorb ideological and political knowledge while maintaining a positive emotional resonance, and simultaneously addressing the need for cognitive bias correction. The entire process is logically rigorous and highly feasible, achieving a synergistic balance between short-term learning adaptability and long-term learning stability, thus enhancing the scientific rigor and adaptability of personalized ideological and political education learning paths.

[0085] In one embodiment, the long-term stability score can be calculated using the following formula: in, This represents the long-term stability score of short-term path segments. This represents the total number of knowledge points in a short-term path segment. Indicates the first Importance weight of each knowledge point, range of values Based on the core importance of knowledge points within the ideological and political knowledge system and the priority of deviation correction, Represents the sentiment co-weight, with a value range of The percentage of the impact of emotional resonance on long-term learning stability. Indicates the first The overall strength of the emotional state vector. Indicates the first The pass rate of the knowledge mastery degree vector. This represents the time decay coefficient, with a range of values. Based on the presupposition of the long-term memory retention pattern of ideological and political knowledge.

[0086] This embodiment achieves accurate calculation of the long-term stability score of short-term path segments, providing a reliable quantitative basis for subsequent path iteration adjustments and optimal segment selection. Furthermore, the formula parameters are highly compatible with the output data of the preceding technical process and the ideological and political learning scenario, ensuring a seamless data flow. The formula integrates multiple influencing factors, including the importance weight of knowledge points, the emotional synergy weight, and the time decay coefficient. It also uses the comprehensive strength of the emotional state vector and the achievement rate of the knowledge mastery vector output by the long-term growth simulator as core calculation bases. This approach considers the knowledge mastery effect, emotional resonance state, the core value of the knowledge points themselves, and the long-term memory decay law of ideological and political learning, avoiding the limitations of considering long-term stability from a single dimension. The value range of each parameter is preset based on the characteristics of the ideological and political knowledge system and the learner's learning patterns, ensuring the formula calculation has strong feasibility, and the calculated long-term stability score can truly reflect the long-term adaptability of the path segments.

[0087] In one embodiment, such as Figure 2 As shown, this application also provides a reinforcement learning-based ideological and political education path generation system, which may include: The interaction record module 601 is used to acquire the learner's historical interaction records and process the historical interaction records through a sequence model to obtain cognitive state vectors and emotional tendency vectors.

[0088] The deviation detection module 602 is used to analyze and obtain the emotional resonance intensity value and the probability of cognitive deviation based on the cognitive state vector and the emotional tendency vector through the deviation detection model.

[0089] The path generation module 603 is used to generate a learning path sequence based on the probability of cognitive bias, the intensity of emotional resonance, and a preset knowledge graph. If the probability of cognitive bias exceeds a preset threshold, the recommended order of knowledge points related to the bias in the learning path sequence is adjusted.

[0090] The path optimization module 604 is used to perform multi-objective optimization calculations on the learning path sequence corresponding to the recommended order of knowledge points to obtain short-term path segments. The short-term path segments are then input into the reinforcement learning model to calculate the long-term stability score, and the segment with the highest score is selected as the personalized ideological and political learning recommendation path for learners.

[0091] The aforementioned reinforcement learning-based ideological and political education learning path generation system comprises an interaction recording module, a deviation detection module, a path generation module, and a path optimization module. These modules work collaboratively to form a complete personalized ideological and political education learning path. The interaction recording module acquires the learner's historical interaction records, inputs these records into a pre-defined sequence model for temporal feature extraction and modeling analysis, and ultimately outputs a cognitive state vector representing the learner's cognitive level and an emotional tendency vector reflecting the learner's emotional attitude. The deviation detection module receives the cognitive state vector and emotional tendency vector output by the interaction recording module, inputs them into a pre-defined deviation detection model, and through the model's fusion analysis of the two types of vectors, outputs an emotional resonance intensity value to quantify emotional fit and a cognitive deviation probability to represent the possibility of cognitive deviation. The path generation module uses the cognitive deviation probability and emotional resonance intensity value output by the deviation detection module, combined with a pre-defined ideological and political education knowledge graph, as input to generate a learning path sequence adapted to the learner's current state. Simultaneously, it retrieves a pre-defined cognitive deviation threshold; if the cognitive deviation probability exceeds this threshold, it adjusts the recommended order of deviation-related knowledge points in the learning path sequence, obtaining the adjusted recommended order of knowledge points and the corresponding path sequence. The path optimization module receives the learning path sequence corresponding to the adjusted knowledge point recommendation order output by the path generation module, and uses a multi-objective optimization algorithm to calculate it to obtain short-term path segments that take into account both knowledge efficiency and emotional resonance. The short-term path segments are then input into a preset reinforcement learning model, and the model iteratively calculates the long-term stability score of each short-term path segment. Finally, the path segment with the highest long-term stability score is selected as the learner's personalized ideological and political learning recommendation path.

[0092] In this embodiment, each module has a clear division of labor and works in concert, effectively ensuring the scientific rigor and adaptability of the personalized ideological and political learning path generation. The interaction recording module provides accurate basic data support for the entire system, ensuring that the cognitive state vector and emotional tendency vector can truly reflect the individual differences of learners; the deviation detection module realizes the quantitative analysis of emotional resonance and cognitive deviation, providing a reliable quantitative basis for subsequent path generation; the path generation module drives path adjustment through the probability of cognitive deviation, ensuring that the path sequence can be specifically adapted to the learner's cognitive deviation correction needs and emotional state; the path optimization module combines multi-objective optimization and reinforcement learning models, taking into account both the short-term adaptability and long-term stability of the path, solving the problems of uniformity, insufficient targeting, and lack of long-term effect considerations in traditional ideological and political learning paths. The entire module system realizes the synergistic advancement of cognitive deviation correction, knowledge efficiency improvement, and emotional resonance cultivation, and can accurately output personalized ideological and political learning recommendation paths that meet the individual needs of learners, providing reliable technical support for the dual needs of cognitive shaping and value guidance in ideological and political education, and improving the implementation effect of personalized ideological and political education.

[0093] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0094] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the reinforcement learning-based ideological and political learning path generation method as described above.

[0095] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0096] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0097] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A method for generating ideological and political learning paths based on reinforcement learning, characterized in that, The method includes: Obtain learners’ historical interaction records, and process the historical interaction records using a sequence model to obtain cognitive state vectors and emotional tendency vectors; Based on the cognitive state vector and emotional tendency vector, the emotional resonance intensity value and cognitive bias probability are obtained through deviation detection model analysis. A learning path sequence is generated based on the cognitive bias probability, the emotional resonance intensity value, and the preset knowledge graph. If the cognitive bias probability exceeds the preset threshold, the recommended order of knowledge points related to the bias in the learning path sequence is adjusted. The learning path sequence corresponding to the recommended order of knowledge points is optimized by multi-objective calculation to obtain short-term path segments. The short-term path segments are then input into a reinforcement learning model to calculate long-term stability scores, and the segments with the highest scores are selected as personalized ideological and political learning recommendation paths for learners.

2. The method according to claim 1, characterized in that, The process of obtaining cognitive state vectors and sentiment tendency vectors from the historical interaction records through sequence modeling includes: Obtain the learner's historical interaction records; the historical interaction records include the learner's behavior logs and feedback information; The cognitive association information in the historical interaction records is input into a sequence model for analysis to obtain the learner's cognitive state vector; the sequence model is constructed using a deep learning algorithm. Based on the cognitive state vector and the feedback information in the historical interaction records, the emotional association features, behavioral feedback features and the cognitive state vector in the feedback information are extracted, and after unified dimension processing and weighted fusion, the fused emotional features are obtained. Based on the fused emotional features and adjustment coefficients, the emotional scores of each sub-feature are calculated, and the learner's emotional tendency vector is determined by classification.

3. The method according to claim 2, characterized in that, The sentiment score is calculated using the following formula: in, Indicates the first Sentiment scores for individual characteristics, Indicates the first Normalized eigenvalues ​​of individual features This represents the weight associated with cognitive states, and its value range is... Based on the degree of influence of cognitive state vectors on emotional tendencies, This indicates the weight of feedback timeliness, with a value range of... , Indicates the first The matching coefficient of knowledge mastery corresponding to each sub-feature Indicates the first The timeliness coefficient of feedback information corresponding to each sub-feature The sentiment correction factor is determined based on the semantic sentiment intensity of the feedback text, with positive semantic correspondence values. Neutral semantics correspond to values Negative semantics correspond to values , This represents the total number of sub-features in the fused emotional features.

4. The method according to claim 1, characterized in that, The step of obtaining the emotional resonance intensity value and cognitive bias probability through deviation detection model analysis based on the cognitive state vector and emotional tendency vector includes: The knowledge mastery, understanding depth, and cognitive logical correlation information in the cognitive state vector are compared and analyzed with the preset ideological and political knowledge point standards, and the potential cognitive bias tendency dimension and bias correlation degree are extracted as cognitive bias correlation features. The intensity values ​​of each emotional dimension in the emotional tendency vector are extracted, and the emotional resonance intensity value, which represents the learner's emotional fit with the ideological and political content, is calculated using a weighted summation algorithm. A pre-defined deviation detection model is used to extract features from the cognitive deviation-related features and calculate deviation tendency feature values; the deviation detection model is constructed based on the standard cognitive framework of ideological and political knowledge points. Based on the deviation tendency feature value and the sentiment tendency vector, a weighted fusion algorithm is used to calculate the comprehensive deviation influence factor; The comprehensive deviation influence factor is compared with a preset cognitive deviation threshold. If it exceeds the cognitive deviation threshold, the comprehensive deviation influence factor is mapped to the corresponding cognitive deviation probability.

5. The method according to claim 1, characterized in that, The process involves generating a learning path sequence based on the cognitive bias probability, the emotional resonance intensity value, and a preset knowledge graph. If the cognitive bias probability exceeds a preset threshold, the recommended order of knowledge points related to the bias in the learning path sequence is adjusted, including: Based on the cognitive bias probability, the emotional resonance intensity value, and the preset knowledge graph, a candidate learning path sequence is generated; the knowledge graph includes the logical relationships between ideological and political knowledge points. If the probability of the cognitive bias exceeds a preset cognitive bias threshold, the priority order of the bias-related knowledge points in the candidate learning path sequence corresponding to the cognitive bias is adjusted. Based on the adjusted priority order, an optimized learning path sequence is generated; Extract the association feature values ​​of key knowledge points in the optimized learning path sequence; the association feature values ​​include the dependency relationship feature values ​​between knowledge points and the deviation correction and adaptation feature values. Based on the associated feature value, the cognitive bias probability, and the emotional resonance intensity value, a dynamic weight model is constructed. The deviation correction weight of each knowledge point is calculated through the dynamic weight model to obtain the recommended order of knowledge points after deviation correction.

6. The method according to claim 1, characterized in that, The personalized ideological and political learning recommendation path for learners is obtained through the following steps: The learning path sequence corresponding to the recommended order of the knowledge points is optimized by multi-objective calculation of short-term path segments; the short-term path segments are ordered sequences of knowledge points that take into account both knowledge efficiency and emotional resonance. The short-term path segments are input into the long-term growth simulator. By simulating the entire learning process of learners according to the short-term path segments, the emotional state vector and knowledge mastery vector corresponding to each step of knowledge learning are output. The emotional state vector and knowledge mastery vector corresponding to each step are input into the reinforcement learning model. The cumulative value of each step's state is calculated iteratively, and the long-term stability score of the short-term path segment is output. Determine whether the long-term stability score has reached the preset stability threshold. If it has not, then based on the state value deviation output during the reinforcement learning model iteration process, adjust the order of knowledge points in the short-term path segment or supplement deviation correction knowledge points to generate the adjusted path segment. The adjusted path segment is input into the long-term growth simulator again, and the reinforcement learning model is used to iterate until the number of iterations reaches the preset iteration threshold or the long-term stability score of a certain path segment meets the stability threshold. The path segment with the highest long-term stability score is selected from all the path segments obtained by the iteration. The path segment with the highest long-term stability score is output as a personalized ideological and political learning recommendation path for learners.

7. The method according to claim 6, characterized in that, The long-term stability score is calculated using the following formula: in, This represents the long-term stability score of short-term path segments. This represents the total number of knowledge points in a short-term path segment. Indicates the first Importance weight of each knowledge point, range of values Based on the core importance of the knowledge points in the ideological and political knowledge system and the priority of deviation correction, Represents the sentiment co-weight, with a value range of The percentage of the impact of emotional resonance on long-term learning stability. Indicates the first The overall strength of the emotional state vector. Indicates the first The pass rate of the knowledge mastery degree vector. This represents the time decay coefficient, with a range of values. Based on the presupposition of the long-term memory retention pattern of ideological and political knowledge.

8. A system for generating ideological and political learning paths based on reinforcement learning, characterized in that, The system includes: The interaction recording module is used to acquire the learner's historical interaction records, and to process the historical interaction records through a sequence model to obtain cognitive state vectors and emotional tendency vectors; The deviation detection module is used to analyze the emotional resonance intensity value and the probability of cognitive deviation based on the cognitive state vector and the emotional tendency vector through the deviation detection model. The path generation module is used to generate a learning path sequence based on the cognitive bias probability, the emotional resonance intensity value, and a preset knowledge graph. If the cognitive bias probability exceeds a preset threshold, the recommended order of knowledge points related to the bias in the learning path sequence is adjusted. The path optimization module is used to perform multi-objective optimization calculations on the learning path sequence corresponding to the recommended order of knowledge points to obtain short-term path segments. The short-term path segments are then input into a reinforcement learning model to calculate long-term stability scores, and the segments with the highest scores are selected as personalized ideological and political learning recommendation paths for learners.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.