Dynamic Optimization Query Method for Educational System Data Based on Deep Reinforcement Learning

Through fuzzy cognitive graphs based on deep reinforcement learning and data query of the education system that optimizes the dynamic adjustment and adaptability of query methods in the education system, efficient and accurate data retrieval and personalized query are achieved.

CN119938702BActive Publication Date: 2025-07-25BEIJING GUANGNIAN WUXIAN SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510029201.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-07-25
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The data query methods of the existing education system lack the ability to adjust dynamically, the query efficiency is inefficient, the results are not correlated, and the targetedness of educational scenarios is lacking, making it difficult to adapt to the real-time changes in educational scenarios and user needs.

Method used

The method based on deep reinforcement learning is adopted to construct a fuzzy cognitive graph and dynamic state space, and the state values of key factor nodes are calculated through causal weight adjustment and fuzzy reasoning. Combining the reward function and reinforcement learning model to optimize the query strategy, a dynamic optimization query strategy is generated, and the causal weight and query strategy are updated through the feedback mechanism.

Benefits of technology

It realizes the efficiency, accuracy and adaptability of data query in the education system, and can significantly improve query efficiency and user satisfaction within 3 iterations, adapting to the multi-dimensional data characteristics of the education scenario and changes in user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938702B_ABST
    Figure CN119938702B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for dynamically optimizing queries of educational system data based on deep reinforcement learning. S1. Construct a structured feature vectorized data set; S2. Establish a fuzzy cognitive map; S3. Form a dynamic state space of the query logic of the educational system; S4. Comprehensively measure the optimization effect of the query logic; S5. Generate a dynamic optimization query strategy for educational system data; S6. Generate a query plan, and perform dynamic optimization queries on the target data in the educational system according to the query plan. S7. Output the dynamic optimization query results of the educational system data, and at the same time record the performance indicators of the query results, and feedback the recorded performance indicators to the fuzzy cognitive map and the deep reinforcement learning model for updating the causal relationship weights of the fuzzy cognitive map and the query strategy of the reinforcement learning model. The present invention significantly improves the efficiency, accuracy and adaptability of the dynamic query of educational system data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of educational technology, and particularly to a method for dynamically optimizing queries of educational system data based on deep reinforcement learning. Background Art

[0002] With the rapid development of artificial intelligence technology, educational informatization and intelligence have become important directions in modern education. The dynamic optimization query of data in the educational system plays a crucial role in supporting personalized learning, teaching assistance, and educational resource management. However, traditional educational data query methods are unable to cope when faced with dynamically changing educational scenarios and are difficult to meet the requirements of modern education for efficiency, accuracy, and flexibility.

[0003] Currently, most educational systems perform data retrieval based on predefined static query logics, usually relying on fixed database query statements or rules. There are obvious limitations when dealing with diverse data such as student learning records, course resources, teaching plans, and examination results. On the one hand, static query logics lack the ability to dynamically adjust and cannot adapt to the real-time changes in students' learning states, the continuous adjustment of teachers' teaching requirements, and the dynamic update of course resources in the educational system. On the other hand, they cannot be optimized according to the characteristics of data types and user needs, resulting in low query efficiency, poor relevance of results, and difficulty in meeting the needs of personalized learning and resource recommendation.

[0004] In recent years, some improved query methods have begun to attempt to introduce artificial intelligence technologies, such as query optimization tools based on traditional machine learning algorithms. Traditional machine learning algorithms improve query efficiency through the analysis of data features. However, their optimization ability is still limited by the manually defined feature sets and fixed optimization rules. Specifically, such technologies lack adaptability when dealing with complex and diverse data and cannot adjust query logics in real time to adapt to the dynamic changes in educational scenarios. In addition, existing methods are difficult to optimize specifically in combination with the specific needs of educational scenarios, which makes them show obvious deficiencies in practical applications.

[0005] In summary, the existing technologies mainly have the following problems when dealing with the dynamic optimization query of educational system data:

[0006] 1. Lack of dynamic adjustment ability: Static query logics cannot adapt to real-time changes in educational scenarios and require manual rewriting of query rules, with poor flexibility;

[0007] 2. Low query efficiency: Existing methods are difficult to optimize query efficiency for the diversity and complexity of educational data, resulting in slow retrieval speed;

[0008] 3. Insufficient result relevance: Since it is impossible to dynamically adjust the query strategy by combining user needs and data characteristics, the existing technologies often have the problem that the query results do not match the needs.

[0009] 4. Lack of pertinence in the educational scenario: The existing query methods usually aim at general optimization and lack customized support for the specific goals of the educational scenario.

[0010] The above problems directly limit the practicality and intelligence level of the dynamic optimization query of educational system data. There is an urgent need for a new method that can combine the dynamic state space, optimize the reward mechanism, and have adaptive learning ability to improve the efficiency, accuracy, and flexibility of educational system data query, and promote the further development of educational informatization. Summary of the Invention

[0011] An object of the present invention is to propose a method for dynamically optimizing the query of educational system data based on deep reinforcement learning, which significantly improves the efficiency, accuracy, and adaptability of the dynamic query of educational system data.

[0012] A method for dynamically optimizing the query of educational system data based on deep reinforcement learning according to an embodiment of the present invention includes the following steps:

[0013] S1. Collect multi-source data in the educational system, preprocess the collected multi-source data, and construct a structured feature vectorized data set;

[0014] S2. Based on the structured feature vectorized data set, determine the key factor nodes affecting the query logic, define the causal relationship and association weights between each factor node, and establish a fuzzy cognitive map;

[0015] S3. Use the fuzzy cognitive map as an inference tool, input the node state values in the educational system data, and calculate the dynamic state values of the key factor nodes through fuzzy reasoning based on the causal relationship and weights to form a dynamic state space of the educational system query logic;

[0016] S4. Design a reward function according to the query target of the educational system. The reward function takes the educational system data query efficiency, query result relevance, and user satisfaction as the core indicators, and comprehensively measures the optimization effect of the query logic;

[0017] S5. Based on the dynamic state space and the reward function, use a deep reinforcement learning model to train the educational system query logic. Take the dynamic state space as the input of the reinforcement learning model, and based on the policy optimization mechanism of the reinforcement learning model, through continuous iterative update, generate a dynamic optimization query strategy for educational system data.

[0018] S6. Apply the dynamically optimized query strategy for the educational system data generated by the deep reinforcement learning model to the query process of the educational system, dynamically adjust the query logic, generate a query plan, and perform a dynamically optimized query on the target data in the educational system according to the query plan;

[0019] S7. Output the dynamically optimized query results of the educational system data, and at the same time record the performance metrics of the query results, and feedback the recorded performance metrics to the fuzzy cognitive map and the deep reinforcement learning model for updating the causal relationship weights of the fuzzy cognitive map and the query strategy of the reinforcement learning model.

[0020] Optionally, the specific steps of S1 are as follows:

[0021] S11. Collect multi-source data in the educational system and represent the collected multi-source data as the original dataset D raw :

[0022] D raw ={D record , D course , D plan , D exam};

[0023] Among them, D record is the learning record data of students, including the time of students' learning activities, the progress of task completion, and test scores. D course is the course resource data, including course content, course difficulty, and learning objectives of the course. D plan is the teaching plan data, including teaching time arrangement, teaching modules, and teacher resource allocation. D exam is the exam result data, including exam scores, knowledge point coverage, and error analysis data;

[0024] S12. Remove redundant items and noise data from the original dataset D raw to generate a denoised dataset, interpolate and complete the missing data in the denoised dataset, use linear interpolation or mean filling method to generate a completed dataset, and normalize all data features in the completed dataset. The normalized dataset D norm ;

[0025] S13. Based on the normalized dataset D norm , extract key features and construct a feature vectorized dataset D vector :

[0026] D vector ={v1, v2, v3, …, v m};

[0027] Each feature vector is represented as:

[0028] v i = [f1, f2, f3, …, f n , i ∈ {1, 2..., m};

[0029] Among them, v i is the feature vector of a single data sample, and f1, f2, …, f n are respectively the extracted learning progress characteristics, course difficulty characteristics, teacher resource allocation intensity characteristics, and exam coverage characteristics.

[0030] Optionally, step S2 specifically includes the following steps:

[0031] S21. Based on the feature vectorized dataset D vector , determine the key factor nodes that affect the query logic in the education system. The key factor nodes include:

[0032] Student learning efficiency node N efficiency Describes the efficiency of students in completing learning tasks, extracted based on learning progress characteristics;

[0033] Course resource difficulty node N difficulty Describes the complexity of course resources, extracted based on course difficulty characteristics;

[0034] Student learning interest node N interest Describes the degree of interest of students in the course, extracted comprehensively based on learning progress characteristics and course difficulty characteristics;

[0035] Teacher query demand node N demand Describes the target demand of teachers for data query, extracted based on teacher resource allocation intensity characteristics;

[0036] Represent the node set as:

[0037] N = {N efficiency , N difficulty , N interest , N demand};

[0038] S22. Conduct a causal relationship analysis on each node in the key factor node N, and determine the causal relationship direction and association weight between the nodes based on the actual requirements of the education system query logic:

[0039] Define the causal relationship E i from node N j to node N ij . If node N i has a positive promoting effect on node N j , then the association weight w ij > 0; if there is an inhibitory effect, then w ij< 0; if there is no direct relationship, then w ij = 0;

[0040] Construct the causal relationship between nodes and its associated weight as an adjacency matrix W, and the matrix element w ij represents the weight relationship from node N i to node N j :

[0041]

[0042] where n is the total number of nodes, and w ij represents the causal relationship weight from node N i to node N j ;

[0043] S23. Based on the node set N and the adjacency matrix W, construct a fuzzy cognitive map FCM:

[0044] FCM = (N, E, W);

[0045] The state value A of each node in the fuzzy cognitive map i is initialized to the normalized value of the corresponding feature in the feature vectorized dataset D vector .

[0046] Optionally, step S3 specifically includes the following steps:

[0047] S31. Based on the fuzzy cognitive map FCM, input the normalized feature values in the feature vectorized dataset D vector as the initial state values of the nodes, and the node initial state vector A (0) is expressed as:

[0048]

[0049] where, is the initial state value of node N i , and n is the total number of nodes in the fuzzy cognitive map;

[0050] S32. Update the state value of each node N i based on the adjacency matrix W of the fuzzy cognitive map, and calculate the dynamic node state value of the fuzzy cognitive map:

[0051]

[0052] where, represents the state value of node N i after the (t + 1)-th round of reasoning, represents the state value of node N j after the t-th round of reasoning, and f(x) is the activation function;

[0053] S33. Construct a dynamic state space for the query logic of the education system based on the node state values calculated by fuzzy inference:

[0054] S dynamic = {S query , S data , S user};

[0055] Among them, S query is the current query requirement state, provided by the state value of the teacher query requirement node, and S data is the state of the data characteristics of the education system, comprehensively provided by the state values of the student learning efficiency node, the course resource difficulty node, and the student learning interest node. S user is the state of the user behavior pattern, and the set A (t+1) of the state values of all nodes in the fuzzy cognitive map provides an overall dynamic characteristic representation.

[0056] Optionally, the specific steps of S4 are as follows:

[0057] S41. Determine the query target based on the actual requirements of the education system. The query target includes three core indicators: the query efficiency of the education system data, the relevance of the query results, and the user satisfaction:

[0058] Query efficiency E efficiency , used to measure the optimization degree of the time required for data query;

[0059] Query result relevance E relevance , used to measure the matching degree between the returned data and the query target;

[0060] User satisfaction E satisfaction , used to measure the subjective satisfaction degree of users with the query results;

[0061] S42. Construct a reward function R based on the query target. The reward function synthesizes the core indicators and is used to measure the optimization effect of the query logic:

[0062] R = α1·E efficiency + α2·E relevance + α3·E satisfaction - β·T;

[0063] Among them, α1, α2, α3 are weight factors used to balance the importance of different indicators, T represents the time required for the query, and β is a time penalty coefficient used to avoid the negative impact of high time overhead on the query efficiency.

[0064] Optionally, the calculation of the query efficiency E efficiency :

[0065]

[0066] Calculation of query result relevance E relevance :

[0067]

[0068] where n is the number of data entries returned by the query, and R i is the relevance score of the i-th data to the query target, and S i is the weight factor of the i-th data;

[0069] Calculation of user satisfaction E satisfaction :

[0070]

[0071] where m is the total number of user feedbacks, and U j is the satisfaction score of the user for the j-th query result, and the value range is [0, 1].

[0072] Optionally, the S5 specifically includes the following steps:

[0073] S51. Use the constructed dynamic state space S dynamic as the input state space S of the reinforcement learning model t ;

[0074] S52. Based on the state space S t , the reinforcement learning model outputs a set of actions A t , and the action A t represents the dynamic adjustment of the query logic, the optimization of the data retrieval path, and the generation of the query plan in the education system query logic;

[0075] S53. According to the reward function R, combined with the feedback values of the actual query efficiency, query result relevance, and user satisfaction, calculate the immediate reward R t corresponding to the current action A t , and the immediate reward R t is used to evaluate the optimization effect of the current action A t :

[0076] R t = α1·E efficiency + α2·E relevance + α3·E satisfaction - β·T;

[0077] S54. Use the deep reinforcement learning algorithm to optimize the policy of the model, and use the policy update:

[0078]

[0079] where Q(St , A t ) is the current state space S t and the action A t 's value function, η is the learning rate, γ is the discount factor, which is used to balance the importance of immediate reward and future reward. S t+1 is the state at the next moment, A t+1 is the action at the next moment;

[0080] S55. The reinforcement learning model generates a dynamically optimized query strategy for the educational system data based on the optimized policy, and uses the optimized query strategy to guide the execution of the actual query logic. At the same time, by continuously collecting the dynamic data of the educational system query environment, the policy of the reinforcement learning model is iteratively updated.

[0081] The beneficial effects of the present invention are as follows:

[0082] (1) The present invention introduces a fuzzy cognitive map to construct the dynamic state space of the educational system. By dynamically adjusting the causal relationship weights and performing fuzzy inference to calculate the state values of the key factor nodes, the dynamic characteristics of the student learning state, curriculum resource changes, and teacher query requirements can be reflected in real time. Compared with the traditional fixed query logic, the dynamic state space of the present invention can flexibly adapt to the multi-dimensional data characteristics that change in real time in the educational scenario, effectively improving the self-adaptability of the query logic and avoiding the complexity of frequently manually adjusting the query rules.

[0083] (2) The present invention uses a deep reinforcement learning model with the dynamic state space as the input, and comprehensively measures the query efficiency, result relevance, and user satisfaction through a reward function to generate an optimal query strategy. Compared with the traditional rule-based query optimization method, the deep reinforcement learning model can improve the execution efficiency of the query path and the accuracy of the results through continuous policy optimization iteration. When facing complex data types and dynamic requirements, the reinforcement learning strategy can automatically adjust the query logic to achieve efficient and accurate data retrieval.

[0084] (3) The present invention feeds back the actual data of the query efficiency, result relevance, and user satisfaction to the fuzzy cognitive map and the reinforcement learning model through the feedback mechanism of the query result performance index to update the causal relationship weights of the fuzzy cognitive map and the policy of the reinforcement learning model, ensuring that the query logic can be optimized and adaptively adjusted in real time as the educational scenario and user needs change. Compared with the existing static optimization methods, the feedback mechanism of the present invention significantly improves the dynamic optimization ability of the query model. Experimental data shows that in the scenario of educational resource changes and user behavior changes, the method of the present invention can achieve significant policy optimization within 3 iterations, and both the query efficiency and user satisfaction are improved by more than 15%. Description of the Drawings

[0085] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the accompanying drawings:

[0086] Figure 1 is a flowchart of a method for dynamically optimizing queries of educational system data based on deep reinforcement learning proposed by the present invention. Detailed implementation manners

[0087] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0088] Reference Figure 1 , a method for dynamically optimizing queries of educational system data based on deep reinforcement learning, includes the following steps:

[0089] S1. Collect multi-source data in the educational system, preprocess the collected multi-source data, and construct a structured feature vectorized data set;

[0090] S2. Based on the structured feature vectorized data set, determine the key factor nodes affecting the query logic, define the causal relationships and association weights between the factor nodes, and establish a fuzzy cognitive map;

[0091] S3. Use the fuzzy cognitive map as an inference tool, input the node state values in the educational system data, and calculate the dynamic state values of the key factor nodes through fuzzy reasoning based on the causal relationships and weights to form a dynamic state space of the educational system query logic;

[0092] S4. Design a reward function according to the query target of the educational system. The reward function takes the query efficiency, relevance of query results, and user satisfaction of the educational system data as core indicators, and comprehensively measures the optimization effect of the query logic;

[0093] S5. Based on the dynamic state space and the reward function, use a deep reinforcement learning model to train the query logic of the educational system. Take the dynamic state space as the input of the reinforcement learning model, and based on the policy optimization mechanism of the reinforcement learning model, through continuous iterative update, generate a dynamic optimization query policy for the educational system data;

[0094] S6. Apply the dynamic optimization query policy for the educational system data generated by the deep reinforcement learning model to the query process of the educational system, dynamically adjust the query logic, generate a query plan, and perform dynamic optimization queries on the target data in the educational system according to the query plan;

[0095] S7. Output the dynamically optimized query results of the education system data, and record the performance metrics of the query results. Feed the recorded performance metrics back to the fuzzy cognitive map and the deep reinforcement learning model for updating the causal relationship weights of the fuzzy cognitive map and the query strategy of the reinforcement learning model.

[0096] In this embodiment, S1 specifically includes the following steps:

[0097] S11. Collect multi-source data in the education system and represent the collected multi-source data as the original data set D raw :

[0098] D raw ={D record , D course , D plan , D exam};

[0099] Among them, D record is the learning record data of students, including the time of students' learning activities, the progress of task completion, and test scores. D course is the course resource data, including course content, course difficulty, and learning objectives of the course. D plan is the teaching plan data, including teaching time arrangement, teaching modules, and teacher resource allocation. D exam is the exam result data, including exam scores, knowledge point coverage, and error analysis data;

[0100] S12. Remove redundant items and noise data from the original data set D raw , generate a denoised data set, interpolate and complete the missing data in the denoised data set, use linear interpolation or mean filling method to generate a completed data set, and normalize all data features in the completed data set. The normalized data set D norm ;

[0101] S13. Based on the normalized data set D norm , extract key features and construct a feature vectorized data set D vector :

[0102] D vector ={v1, v2, v3,..., v m};

[0103] Each feature vector is represented as:

[0104] v i =[f1, f2, f3,..., f n , i ∈ {1, 2..., m};

[0105] Among them, v iis the feature vector of a single data sample, f1, f2, …, f n are respectively the extracted learning progress characteristics, course difficulty characteristics, teacher resource allocation intensity characteristics, and exam coverage characteristics.

[0106] In this embodiment, S2 specifically includes the following steps:

[0107] S21. Based on the feature vectorized dataset D vector , determine the key factor nodes affecting the query logic in the education system. The key factor nodes include:

[0108] Student learning efficiency node N efficiency Describes the efficiency of students in completing learning tasks, extracted based on learning progress characteristics;

[0109] Course resource difficulty node N difficulty Describes the complexity of course resources, extracted based on course difficulty characteristics;

[0110] Student learning interest node N interest Describes the degree of interest of students in courses, extracted comprehensively based on learning progress characteristics and course difficulty characteristics;

[0111] Teacher query demand node N demand Describes the target demand of teachers for data query, extracted based on teacher resource allocation intensity characteristics;

[0112] Represent the node set as:

[0113] N = {N efficiency , N difficulty , N interest , N demand};

[0114] S22. Conduct a causal relationship analysis on each node in the key factor node N, and determine the causal relationship direction and association weight between nodes based on the actual requirements of the education system query logic:

[0115] Define the causal relationship E i from node N j to node N ij . If node N i has a positive promoting effect on node N j , then the association weight w ij > 0; if there is an inhibitory effect, then w ij < 0; if there is no direct relationship, then w ij = 0;

[0116] Construct the causal relationship between nodes and its association weight into an adjacency matrix W, and the matrix element w ij represents from node node N iTo node N j The weight relationship with:

[0117]

[0118] where n is the total number of nodes, w ij represents the causal relationship weight from node N i to node N j ;

[0119] S23. Based on the node set N and the adjacency matrix W, construct a fuzzy cognitive map FCM:

[0120] FCM = (N, E, W);

[0121] The state value A of each node in the fuzzy cognitive map i is initialized to the normalized value of the corresponding feature in the feature vectorized dataset D vector .

[0122] In this embodiment, S3 specifically includes the following steps:

[0123] S31. Based on the fuzzy cognitive map FCM, input the normalized feature values in the feature vectorized dataset D vector as the initial state values of the nodes. The initial state vector A of the nodes (0) is expressed as:

[0124]

[0125] where is the initial state value of node N i , and n is the total number of nodes in the fuzzy cognitive map;

[0126] S32. Update the state values of each node N i based on the adjacency matrix W of the fuzzy cognitive map, and calculate the dynamic node state values of the fuzzy cognitive map:

[0127]

[0128] where represents the state value of node N i after the (t + 1)-th round of reasoning, represents the state value of node N j after the t-th round of reasoning, and f(x) is the activation function;

[0129] S33. Based on the node state values calculated by fuzzy reasoning, construct a dynamic state space for the query logic of the education system:

[0130] S dynamic = {S query , S data,S user};

[0131] Among them, S query is the current query requirement status, provided by the status value of the teacher's query requirement node, S data is the educational system data characteristic status, provided by the comprehensive status values of the student learning efficiency node, the course resource difficulty node, and the student learning interest node, S user is the user behavior pattern status, and the set A of the status values of all nodes in the fuzzy cognitive map (t+1) provides an overall dynamic characteristic representation.

[0132] In this embodiment, S4 specifically includes the following steps:

[0133] S41. Based on the actual requirements of the educational system, determine the query target, and the query target includes three core indicators: the query efficiency of the educational system data, the relevance of the query results, and the user satisfaction:

[0134] Query efficiency E efficiency , used to measure the optimization degree of the time required for data query;

[0135] Query result relevance E relevance , used to measure the matching degree of the returned data with the query target;

[0136] User satisfaction E satisfaction , used to measure the subjective satisfaction degree of the user with the query results;

[0137] S42. Based on the query target, construct a reward function R. The reward function synthesizes the core indicators and is used to measure the query logic optimization effect:

[0138] R = α1·E efficiency +α2·E relevance +α3·E satisfaction -β·T;

[0139] Among them, α1, α2, α3 are weight factors used to balance the importance of different indicators, T represents the time required for the query, and β is the time penalty coefficient used to avoid the negative impact of high time overhead on the query efficiency.

[0140] In this embodiment, the calculation of the query efficiency E efficiency :

[0141]

[0142] The calculation of the query result relevance E relevance :

[0143]

[0144] where n is the number of data entries returned by the query, and R i is the relevance score of the i-th data to the query target, and S i is the weight factor of the i-th data;

[0145] The user satisfaction E satisfaction is calculated as follows:

[0146]

[0147] where m is the total number of user feedbacks, and U j is the satisfaction score of the user for the j-th query result, and the value range is [0, 1].

[0148] In this embodiment, S5 specifically includes the following steps:

[0149] S51. Use the constructed dynamic state space S dynamic as the input state space S t of the reinforcement learning model;

[0150] S52. Based on the state space S t , the reinforcement learning model outputs a set of actions A t , and the action A t represents the dynamic adjustment of the query logic, the optimization of the data retrieval path, and the generation of the query plan in the query logic of the education system;

[0151] S53. According to the reward function R, combined with the feedback values of the actual query efficiency, the relevance of the query result, and the user satisfaction, calculate the immediate reward R t corresponding to the current action A t , and the immediate reward R t is used to evaluate the optimization effect of the current action A t :

[0152] R t = α1·E efficiency + α2·E relevance + α3·E satisfaction - β·T;

[0153] S54. Use the deep reinforcement learning algorithm to optimize the policy of the model, and use the policy update:

[0154]

[0155] where Q(S t , A t ) is the value function of the current state space S t and the action A t , η is the learning rate, γ is the discount factor, which is used to balance the importance of the immediate reward and the future reward, and St+1 is the state for the next moment, A t+1 is the action for the next moment;

[0156] S55. The reinforcement learning model generates a dynamically optimized query strategy for the educational system data based on the optimized policy, uses the optimized query strategy to guide the execution of the actual query logic, and at the same time iteratively updates the policy of the reinforcement learning model by continuously collecting the dynamic data of the educational system query environment.

[0157] Example 1:

[0158] In the application of a certain city's education management platform, the platform is responsible for serving the teaching and learning data management of more than 100 primary and secondary schools in the whole city. One noon, a teacher in a certain school needs to quickly query the learning efficiency of a certain class of students and recommend learning resources for students with low efficiency, and at the same time analyze the matching degree between the difficulty of a newly launched course resource and the learning interest of students. The following is the specific application scenario:

[0159] At 11:30 am on December 15, 2024, a head teacher in the Seventh Middle School of a certain city hopes to understand the learning efficiency of the students in Class 3, Grade 8, make personalized resource recommendations for students with continuously low task completion rates, and at the same time the dean of teaching plans to analyze the matching situation between the difficulty of an AI course newly introduced in the school and the learning interest of students. The platform used the method of the present invention to complete multiple queries within 30 minutes and provided an accurate analysis report.

[0160] During the operation of the platform, the system dynamically collects students' learning record data, course resource data, teaching plan data, and exam result data, which involve the following detailed data:

[0161] Learning task completion records of 50 students in Class 3, Grade 8 (time span of the past two weeks), with a data volume of about 5GB;

[0162] Data of the newly introduced AI course resources, including course chapter division, resource difficulty grading, and course objective description, with a data volume of about 2GB;

[0163] Students' learning interest data, sourced from recent course selections, exam performances, and relevant survey feedback of students, with a total volume of about 1GB.

[0164] The first stage: Identifying students with low efficiency

[0165] The platform first uses fuzzy cognitive maps to reason about the student learning efficiency nodes. In the past 14 days, the learning efficiency score of student A (ID: S20231215-01) was lower than the threshold of 0.6 (score range 0-1). Their task completion time far exceeded the average (about 2 hours while the class average was 45 minutes), and the learning efficiency status value was determined to be abnormal.

[0166] Further analysis shows that the status value of student A's learning interest node is 0.4, far lower than the class average of 0.8, and their usage rate of course resources at the medium to high difficulty level is only 15%. Based on the dynamic query logic, the system automatically generated personalized learning resource recommendations, including 3 basic-level resources (in the examples: "AI Introduction Basic Tasks", "Image Processing Introduction", "Computer Processing Introduction") and 2 intermediate resources related to their interests (in the examples: "Fun Robot Programming", "Fun Computing"), and the result generation time was 12 seconds.

[0167] Phase Two: Analyze the matching degree between course resource difficulty and interest

[0168] Regarding the newly launched AI course, the dean of teaching affairs hopes to analyze the matching situation between the interest status value of students in the course and the difficulty of course resources. The platform dynamically queried the student data of the eighth grade of the whole school and found through analysis that:

[0169] The difficulty score of the course resources is 0.85, belonging to the higher difficulty level;

[0170] Among the 1200 students in the eighth grade, the number of students with an interest status value higher than 0.7 is 400, accounting for about 33%.

[0171] Further matching analysis shows that the task completion rate of students with an interest status value higher than 0.7 in the course reaches 95%, while the completion rate of students with an interest status value lower than 0.5 is only 48%. The system automatically generated a matching analysis report, and the report generation took about 18 seconds.

[0172] Phase Three: Generate performance indicators and optimization feedback

[0173] After completing the query task, the system automatically recorded the query performance indicators, including:

[0174] Query efficiency (average response time): 15 seconds;

[0175] Query result relevance (system evaluation relevance score): 92%;

[0176] User satisfaction score: 89%.

[0177] The platform fed back the above performance indicators to the fuzzy cognitive map and the reinforcement learning model:

[0178] 1. The system dynamically adjusts the causal relationship weight between the "learning efficiency node" and the "learning interest node" in the fuzzy cognitive map, from the original weight of 0.7 to 0.8, to adapt to the actual situation where students' learning efficiency and interest are strongly correlated.

[0179] 2. The reinforcement learning model updates the query strategy based on immediate rewards, and dynamically adjusts the query logic in subsequent similar queries, improving the execution efficiency of the query path.

[0180] To verify the effectiveness of the present invention, the platform simultaneously uses the traditional method and the method of the present invention to conduct a comparative test on the above scenario, and the results are shown in Table 1 below:

[0181] Table 1 Comparative test data of the traditional method and the method of the present invention for the above scenario

[0182]

[0183] This embodiment demonstrates the application effect of the method of the present invention in the actual educational scenario. By dynamically adjusting the query logic and optimizing the query strategy, efficient and accurate query results are achieved, significantly improving the query efficiency and user satisfaction, and providing an efficient solution for the dynamic optimization query of educational system data.

[0184] The present invention introduces a fuzzy cognitive map to construct the dynamic state space of the educational system. By dynamically adjusting the causal relationship weight and performing fuzzy inference to calculate the state values of key factor nodes, the dynamic characteristics of students' learning status, curriculum resource changes, and teachers' query requirements can be reflected in real time. Compared with the traditional fixed query logic, the dynamic state space of the present invention can flexibly adapt to the multi-dimensional data characteristics that change in real time in the educational scenario, effectively improving the self-adaptability of the query logic and avoiding the complexity of frequently manually adjusting query rules.

[0185] The present invention uses a deep reinforcement learning model with the dynamic state space as the input, and comprehensively measures the query efficiency, result relevance, and user satisfaction through a reward function to generate an optimal query strategy. Compared with the traditional rule-based query optimization method, the deep reinforcement learning model can improve the execution efficiency of the query path and the accuracy of the results through continuous policy optimization iterations. When facing complex data types and dynamic requirements, the reinforcement learning strategy can automatically adjust the query logic to achieve efficient and accurate data retrieval.

[0186] Through the feedback mechanism of the query result performance indicators, the present invention feeds back the actual data of query efficiency, result relevance, and user satisfaction to the fuzzy cognitive map and the reinforcement learning model to update the causal relationship weights of the fuzzy cognitive map and the strategy of the reinforcement learning model, ensuring that the query logic can be optimized in real time and adaptively adjusted with the changes in the educational scenario and user needs. Compared with the existing static optimization methods, the feedback mechanism of the present invention significantly improves the dynamic optimization ability of the query model. Experimental data shows that in the scenarios of educational resource changes and user behavior changes, the method of the present invention can achieve significant policy optimization within 3 iterations, and both the query efficiency and user satisfaction are increased by more than 15%.

[0187] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.

Claims

1. A method for dynamically optimizing queries of educational system data based on deep reinforcement learning, characterized in that, It includes the following steps: S1. Collect multi-source data in the education system, preprocess the collected multi-source data, and construct a structured feature vectorized dataset; S2. Based on the structured feature vectorized dataset, determine the key factor nodes affecting the query logic, define the causal relationships and association weights between each factor node, and establish a fuzzy cognitive map; S3. Take the fuzzy cognitive map as an inference tool, input the node state values in the education system data, and calculate the dynamic state values of the key factor nodes through fuzzy inference based on the causal relationships and weights, forming a dynamic state space of the education system query logic; S4. Design a reward function according to the query objectives of the education system. The reward function takes the query efficiency, relevance of query results, and user satisfaction of the education system data as core indicators, and comprehensively measures the optimization effect of the query logic; S5. Based on the dynamic state space and the reward function, use a deep reinforcement learning model to train the query logic of the education system. Take the dynamic state space as the input of the reinforcement learning model, and based on the policy optimization mechanism of the reinforcement learning model, through continuous iterative updates, generate a dynamic optimization query strategy for the education system data; S6. Apply the dynamic optimization query strategy of the education system data generated by the deep reinforcement learning model to the query process of the education system, dynamically adjust the query logic, generate a query plan, and perform dynamic optimization query on the target data in the education system according to the query plan; S7. Output the dynamic optimization query results of the education system data, and at the same time record the performance indicators of the query results. Feed the recorded performance indicators back to the fuzzy cognitive map and the deep reinforcement learning model for updating the causal relationship weights of the fuzzy cognitive map and the query strategy of the reinforcement learning model.

2. A method for dynamically optimizing query of educational system data based on deep reinforcement learning according to claim 1, characterized in that The specific steps of S1 include the following: S11. Collect multi-source data in the education system and represent the collected multi-source data as the original dataset D raw : D raw = {D record , D course , D plan , D exam}; Among them, D record is the learning record data of students, including the time of students' learning activities, the progress of task completion, and test scores. D course is the course resource data, including course content, course difficulty, and learning objectives of the course. D plan is the teaching plan data, including teaching time arrangement, teaching modules, and teacher resource allocation. D exam is the exam result data, including exam scores, knowledge point coverage, and error analysis data; S12. For the original dataset D raw Remove redundant items and noise data from the data to generate a denoised dataset. Interpolate and complete the missing data in the denoised dataset, and use the linear interpolation method or the mean filling method to generate a completed dataset. Normalize all data features in the completed dataset, and the normalized dataset D norm ; S13. Based on the normalized dataset D norm , extract key features and construct a feature vectorized dataset D vector : D vector = {v1, v2, v3, …, v m}; Each feature vector is represented as: v i = [f1, f2, f3, …, f n , i ∈ {1, 2..., m}; Among them, v i is the feature vector of a single data sample, and f1, f2, …, f n are respectively the extracted learning progress characteristics, course difficulty characteristics, teacher resource allocation intensity characteristics, and exam coverage characteristics.

3. A method for dynamically optimizing query of educational system data based on deep reinforcement learning according to claim 1, characterized in that, The specific steps of S2 include the following: S21. Vectorize the dataset D based on features vector , and determine the key factor nodes that affect the query logic in the education system. The key factor nodes include: Student learning efficiency node N efficiency Describes the efficiency of students in completing learning tasks, extracted based on learning progress characteristics; Course resource difficulty node N difficulty Describe the complexity of the course resource, extracted based on the course difficulty characteristics; Student learning interest node N interest Describes the degree of interest of students in the course, comprehensively extracted based on the learning progress characteristics and the course difficulty characteristics; Teacher query requirement node N demand Describe the target requirements of teachers for data query, extracted based on the characteristics of teacher resource allocation intensity; The node set is represented as: N = {N efficiency , N difficulty , N interest , N demand}; S22. Conduct a causal relationship analysis on each node in the key factor node N, and based on the actual requirements of the education system query logic, determine the causal relationship direction and association weight between nodes: Define slave node N i To Node N j The causal relationship E ij , if node N i For node N j If there is a positive promotion effect, the association weight w ij >0; if there is an inhibitory effect, then w ij <0; if there is no direct relationship, then w ij =0; Construct the causal relationship between nodes and their associated weights as an adjacency matrix \(W\), and the matrix element \(w\) ij represents the weight relationship from node \(N\) i to node \(N\) j : where n is the total number of nodes, and w ij represents the causal relationship weight from node N i to node N j ; S23. Based on the node set N and the adjacency matrix W, construct a fuzzy cognitive map FCM: FCM = (N, E, W); The state value A of each node in the fuzzy cognitive map i is initialized to the normalized value of the corresponding feature in the feature-vectorized dataset D vector .

4. A method for dynamically optimizing query of educational system data based on deep reinforcement learning according to claim 1, characterized in that The specific steps of S3 include the following: S31. Vectorize the feature dataset D based on the fuzzy cognitive map FCM vector Input the normalized feature values in it as the initial state values of the nodes, and the initial state vector A of the nodes (0) is expressed as: Among them, is the initial state value of node N i , and n is the total number of nodes in the fuzzy cognitive map; S32. Update the state value of each node N based on the adjacency matrix W of the fuzzy cognitive map i to calculate the dynamic node state value of the fuzzy cognitive map: Among them, represents the state value of node N after the (t + 1)-th round of reasoning i ; represents the state value of node N after the t-th round of reasoning j , and f(·) is the activation function. S33. Based on the node state values calculated by fuzzy inference, construct a dynamic state space of the education system query logic: S dynamic = {S query , S data , S user}; Among them, S query is the current query requirement status, provided by the status value of the teacher's query requirement node. S data is the educational system data characteristic status, comprehensively provided by the status values of the student learning efficiency node, the course resource difficulty node, and the student learning interest node. S user is the user behavior pattern status, and the set A of the status values of all nodes in the fuzzy cognitive map (t+1) provides an overall dynamic characteristic representation.

5. A method for dynamically optimizing query of educational system data based on deep reinforcement learning according to claim 1, characterized in that The specific steps of S4 include the following: S41. Based on the actual requirements of the education system, determine the query objectives, and the query objectives include three core indicators: the query efficiency of the education system data, the relevance of query results, and user satisfaction: Query efficiency E efficiency , which is used to measure the optimization degree of the time required for data query; Relevance of Query Results E relevance , which is used to measure the degree of match between the returned data and the query target; User satisfaction E satisfaction , which is used to measure the subjective satisfaction degree of users with the query results; S42. Construct a reward function R based on the query objectives. The reward function comprehensively considers the core indicators and is used to measure the optimization effect of the query logic: R = α1·E efficiency + α2·E relevance + α3·E satisfaction - β·T; Among them, α1, α2, α3 are weight factors used to balance the importance of different indicators, T represents the time required for the query, and β is a time penalty coefficient used to avoid the negative impact of high time overhead on the query efficiency.

6. The data dynamic optimization query method for an education system based on deep reinforcement learning according to claim 5, wherein The query efficiency E efficiency Calculation: Calculation of Query Result Relevance E relevance : where n1 is the number of data entries returned by the query, and R k is the relevance score of the k-th data to the query target, and S k is the weight factor of the k-th data; User satisfaction E satisfaction Calculation: where m1 is the total number of user feedbacks, and U j is the satisfaction score of the user for the j-th query result, with a value range of [0, 1].

7. A method for dynamically optimizing query of educational system data based on deep reinforcement learning according to claim 5, characterized in that The specific steps of S5 include the following: S51. Use the constructed dynamic state space S dynamic as the input state space S of the reinforcement learning model t ; S52. Based on the state space S t , the reinforcement learning model outputs a set of actions A t , the actions A t represent the dynamic adjustment of the query logic in the education system query logic, the optimization of the data retrieval path, and the generation of the query plan; S53. Calculate the current action A according to the reward function R, combined with the feedback values of the actual query efficiency, query result relevance, and user satisfaction t The corresponding immediate reward R t , the immediate reward R t is used to evaluate the optimization effect of the current action A t : R t = α1·E efficiency + α2·E relevance + α3·E satisfaction - β·T; S54. Use a deep reinforcement learning algorithm to optimize the strategy of the model and utilize policy updates: where Q(S t , A t ) is the value function of the current state space S t and action A t , η is the learning rate, γ is the discount factor, which is used to balance the importance of immediate reward and future reward, S t+1 is the state at the next moment, A t+1 is the action at the next moment; S55. The reinforcement learning model generates a dynamically optimized query policy for educational system data based on the optimized policy, uses the optimized query policy to guide the execution of the actual query logic, and at the same time iteratively updates the policy of the reinforcement learning model by continuously collecting the dynamic data of the educational system query environment.

Citation Information

Patent Citations

  • Database query optimization method based on deep reinforcement learning

    CN117349319A

  • Self-adaptive multi-granularity cloud resource arrangement method based on hierarchical reinforcement learning

    CN119225986A