Intelligent evaluation model dynamic optimization method based on reinforcement learning
Through the dynamic optimization method of intelligent evaluation model based on reinforcement learning, students' current interactive data and historical interaction sequences are used for personalized adjustments, solving the problem of lag in traditional evaluation models, and achieving rapid response and accurate evaluation of students' learning status.
Patent Information
- Application Number
- CN202510752294.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Traditional intelligent evaluation models are difficult to adapt to the dynamic changes in students' knowledge mastery during learning, resulting in lagging evaluation results and unable to provide continuous and accurate personalized evaluation and feedback.
Using a method based on reinforcement learning, by obtaining students' current interactive data and historical interaction sequences, fine-tuning is used to use the basic intelligent evaluation model and its general parameters to calculate the benchmark prediction error and recent performance characteristics, integrate it into the current reinforcement learning state, and use the RL Agent policy network to select adjustment actions and update personalized adjustment parameters.
It realizes rapid response and targeted adjustment to students' individual learning status, ensures the stability and generalization ability of the basic model, and improves the timeliness and accuracy of the evaluation results.
Smart Images

Figure CN120278397A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent evaluation, and more specifically, to a method for dynamically optimizing an intelligent evaluation model based on reinforcement learning. Background Art
[0002] In the process of educational intelligence, as a core tool for learner ability evaluation and personalized guidance, the intelligent evaluation model can effectively evaluate an individual's knowledge level, skill mastery, or ability status. However, an individual's learning process and ability are dynamically changing, and their performance will continue to evolve with factors such as time, learning experience, and forgetting. Traditional intelligent evaluation models usually adopt static parameter configurations and are difficult to adapt to the dynamic evolution of students' knowledge mastery during the learning process. Especially when facing a learning group with significant cognitive ability differences, the fixed model parameters often lead to evaluation results lagging behind the actual learning progress, and the prediction accuracy and evaluation accuracy of the model may decline over time, unable to provide continuous and accurate personalized evaluation or feedback.
[0003] Currently, the mainstream intelligent evaluation model parameter optimization methods mostly rely on the method of periodic batch updates. By regularly collecting students' learning data, the model parameters are retrained and adjusted using offline analysis. However, in an educational scenario, the continuous interaction data generated by students contains the changing rules of learning behavior patterns. This offline optimization mechanism cannot respond to individual learning state changes in a timely manner, and there are problems such as a long update cycle and lagging parameter optimization, making it difficult to ensure the timeliness and accuracy of evaluation results.
[0004] Therefore, an optimized method and system for dynamically optimizing an intelligent evaluation model based on reinforcement learning are expected. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed. The embodiments of this application provide a method for dynamically optimizing an intelligent evaluation model based on reinforcement learning. It fine-tunes the basic intelligent evaluation model and its general parameters according to the historical interaction data of a student object, uses the updated model parameters to predict the current interaction data of the student object, and then calculates the benchmark prediction error between the benchmark prediction result and the real interaction, and integrates the recent performance statistical features and current personalized adjustment parameters of the student object as the current reinforcement learning (RL) state. The RL Agent policy network is used to select adjustment actions, so as to calculate and update the personalized adjustment parameters of the student object based on the adjustment actions. On the basis of the core parameters of the basic evaluation model, this method realizes personalized adjustment for each student's real-time learning state, ensuring both the stability and generalization ability of the basic model and rapid response and targeted adjustment to the changes in the individual state of students.
[0006] According to one aspect of the present application, there is provided a method for dynamically optimizing an intelligent evaluation model based on reinforcement learning, which includes: Obtain the current interaction data of the first student object and the historical interaction sequence of the first student object; Use the basic intelligent evaluation model and its general parameters to predict the current interaction data based on the historical interaction sequence of the first student object to obtain a benchmark prediction result; Calculate the benchmark prediction error between the benchmark prediction result and the real interaction, and integrate the benchmark prediction error, recent performance statistical features, and current personalized adjustment parameters to obtain the current RL state; Input the current RL state into the RL Agent policy network to obtain the selected adjustment action; Calculate new personalized adjustment parameters based on the selected adjustment action.
[0007] Compared with the prior art, the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning provided by the present application fine-tunes the basic intelligent evaluation model and its general parameters according to the historical interaction data of the student object, predicts the current interaction data of the student object with the updated model parameters, and then calculates the benchmark prediction error between the benchmark prediction result and the real interaction, and integrates the recent performance statistical features and current personalized adjustment parameters of the student object as the current reinforcement learning (RL) state, and uses the RL Agent policy network to select adjustment actions, so as to calculate and update the personalized adjustment parameters of the student object based on the adjustment actions. Based on the core parameters of the basic evaluation model, this method makes personalized adjustments for the real-time learning status of each student individual, which not only ensures the stability and generalization ability of the basic model, but also realizes the rapid response and targeted adjustment to the changes in the individual state of the student. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0009] Figure 1 FIG. is a flowchart of a method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of the present application.
[0010] Figure 2 FIG. is a schematic diagram of data flow of a method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of the present application.
[0011] Figure 3 It is a flowchart of sub-step S2 of the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of the present application.
[0012] Figure 4 It is a flowchart of sub-step S5 of the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of the present application.
[0013] Figure 5 It is a flowchart of sub-step S52 of the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of the present application.
[0014] Figure 6 It is a flowchart of sub-step S522 of the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of the present application.
[0015] Figure 7 It is a flowchart of sub-step S5223 of the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of the present application. Detailed implementation manners
[0016] As shown in the present application and the claims, unless the context clearly indicates an exceptional situation, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "including" and "comprising" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0017] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0018] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. On the contrary, as needed, various steps can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several steps can be removed from these processes.
[0019] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0020] It should be noted that in this application, all actions of obtaining data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the data is located and obtaining the authorization given by the corresponding device owner.
[0021] In view of the technical problems described in the above background art, this application proposes a method for dynamically optimizing an intelligent evaluation model based on reinforcement learning. It fine-tunes the basic intelligent evaluation model and its general parameters according to the historical interaction data of the student object, predicts the current interaction data of the student object with the updated model parameters, and then calculates the baseline prediction error between the baseline prediction result and the real interaction, and integrates the recent performance statistical features and the current personalized adjustment parameters of the student object as the current reinforcement learning (RL) state, and uses the RL Agent policy network to select adjustment actions, so as to calculate and update the personalized adjustment parameters of the student object based on the adjustment actions. Based on the core parameters of the basic evaluation model, this method makes personalized adjustments according to the real-time learning state of each student individual, which not only ensures the stability and generalization ability of the basic model, but also realizes the rapid response and targeted adjustment to the changes in the individual state of the student.
[0022] Figure 1 It is a flowchart of the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of this application. Figure 2 It is a schematic diagram of data flow of the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of this application. As Figure 1 and Figure 2 shown, the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning includes the steps: S1, obtaining the current interaction data of the first student object and the historical interaction sequence of the first student object; S2, using the basic intelligent evaluation model and its general parameters, predicting the current interaction data based on the historical interaction sequence of the first student object to obtain a baseline prediction result; S3, calculating the baseline prediction error between the baseline prediction result and the real interaction, and integrating the baseline prediction error, recent performance statistical features and the current personalized adjustment parameters to obtain the current RL state; S4, inputting the current RL state into the RL Agent policy network to obtain the selected adjustment action; S5, calculating new personalized adjustment parameters based on the selected adjustment action.
[0023] In the above dynamic optimization method of the intelligent evaluation model based on reinforcement learning, in step S1, the current interaction data of the first student object and the historical interaction sequence of the first student object are obtained. It should be understood that since the learning behaviors of individual students have temporal correlation, and the evolution of cognitive states is deeply affected by historical learning trajectories. Therefore, in order to capture the long-term dependencies in the dynamic learning patterns of students, this application constructs a continuous learning state representation by collecting the current interaction data of students in real time (such as answer records, interaction durations, knowledge point residence times) and historical interaction sequences (including behavioral characteristics in multiple past learning stages). Specifically, the interaction data stream generated by students on the intelligent evaluation platform can be obtained in real time through a data interface, and a sliding window mechanism is used to extract structured data including dimensions such as timestamps, knowledge point labels, behavioral types, and answer results. At the same time, a distributed storage engine is combined to retrieve the historical interaction sequence of this student. In this way, a multi-granularity data foundation covering short-term behavioral characteristics and long-term learning rules can be established, providing temporal context support for the subsequent dynamic optimization of the evaluation model.
[0024] Specifically, first, a real-time data pipeline capable of continuously collecting and processing students' interaction behaviors needs to be constructed. This data pipeline is usually composed of a front-end application interface, a back-end service interface, and a distributed data storage system. When students perform operations such as answering questions, watching teaching videos, and participating in interactive exercises on the intelligent evaluation platform, these behaviors are immediately recorded and uploaded to the background server through a unified data collection interface. The collected content not only includes students' answer results, answer times, knowledge point labels, but also covers fine-grained behavioral characteristics such as page residence duration, click frequency, and jump path. These information constitute the core content of the current interaction data, reflecting the learning state and cognitive performance of students at a specific moment.
[0025] At the same time, in order to understand the learning patterns of students more comprehensively, it is also necessary to effectively integrate their historical interaction sequences. The historical interaction sequence refers to the interaction behavior records accumulated by students in multiple past learning stages, including but not limited to previous answer situations, stage assessment results, the change trend of knowledge mastery curves, and the evolution law of learning paths. These data are usually stored in a database in a structured or semi-structured form, may be distributed at multiple time nodes, and cover learning performances under different knowledge points, different question types, and different difficulty levels. By retrieving these historical data, a complete map of students' long-term learning behaviors can be established, thus providing a basis for their personalized modeling.
[0026] To ensure the effective integration of current interaction data with historical interaction sequences, a sliding window mechanism is often adopted within the system to extract time series features at multiple scales. Specifically, the sliding window can intercept a continuous segment of learning behavior in the time dimension, such as all interaction records within the past 24 hours, or relevant learning activities related to a specific knowledge point in the recent week. This mechanism helps capture the changing trends of students' short-term behavior patterns and also reflects the stability of their long-term learning habits. In addition, the window size can be dynamically adjusted according to actual needs to adapt to different types of learning tasks and assessment scenarios.
[0027] During the process of data collection and integration, it is also necessary to semantically annotate students' behavior data in combination with the knowledge point label system. For example, in a math problem, behaviors such as whether a student answers correctly, the time taken to answer the question, and whether they repeatedly attempt to modify the answer can all be labeled with corresponding knowledge point labels (such as "quadratic function", "trigonometric identity transformation", etc.), thus forming a structured representation of learning behavior. This annotation method enables the system to more accurately identify students' mastery of specific knowledge points in subsequent analysis and accordingly judge the changing direction of their cognitive state.
[0028] In addition to the knowledge point dimension, the organization of students' interaction data also needs to consider information from multiple dimensions such as behavior type, timestamp, device type, and network environment. For example, when the same question is answered at different times, the students' abilities reflected may vary; learning behaviors on mobile devices and desktop devices may also exhibit different attention distribution characteristics. Through the comprehensive analysis of these auxiliary dimensions, the expressive power and predictive value of the data can be further enhanced, providing richer input signals for subsequent model optimization.
[0029] To ensure the integrity and timeliness of data, the system usually adopts a distributed data storage engine to manage students' historical interaction sequences. Such storage systems have high concurrent read and write capabilities and flexible data scalability, and can support the rapid retrieval and update of massive user behavior logs. In actual deployment, an indexing mechanism can be established to classify and organize the learning records of each student, enabling rapid location of relevant historical data when needed. In addition, a caching strategy can be introduced to pre-load frequently accessed learning behavior segments into memory, thereby accelerating data retrieval speed and reducing system response latency.
[0030] In the process of specific implementation, data quality control is also an important link that cannot be ignored. Since students may make mistakes in operation, experience abnormal interruptions, or submit repeatedly during the use of the platform, corresponding data cleaning rules need to be set to eliminate invalid or incorrect interaction records. For example, for the behavior of answering questions repeatedly in a short period of time, the system can automatically identify and filter out the part that does not conform to the normal learning logic; for the residual data of users who have not logged in for a long time, it can be regularly cleaned or archived according to the preset strategy. These measures help to improve the accuracy and consistency of data and avoid model prediction deviation caused by noise interference.
[0031] It is worth noting that the collection and integration of students' interaction data are not only technical operations but also involve issues of data privacy and security protection. Therefore, at the beginning of system design, relevant data compliance requirements need to be followed, such as GDPR or information security standards in other countries and regions. Specific measures include desensitizing sensitive fields, restricting unauthorized access rights, and adopting encrypted transmission protocols to ensure that students' learning behavior data is not leaked or misused during the transfer process.
[0032] In the above dynamic optimization method of the intelligent evaluation model based on reinforcement learning, in step S2, the basic intelligent evaluation model and its general parameters are used to predict the current interaction data based on the historical interaction sequence of the first student object to obtain a benchmark prediction result. Specifically, since directly using a personalized model may have problems such as cold start or insufficient generalization. Therefore, in the initial stage, this application first adopts a well-trained basic intelligent evaluation model and its general parameters, and on this basis, fine-tune its general parameters according to the historical interaction sequence of this student object to obtain fine-tuned model parameters adapted to the current learning state of this student object, and then use the updated model parameters to predict the current interaction data to obtain a more accurate benchmark prediction result. In this way, both the stability and generalization ability of the basic model are utilized, and the adaptability of the model to the current student state is improved through personalized fine-tuning. Among them, Figure 3 It is a flowchart of sub-step S2 of the dynamic optimization method of the intelligent evaluation model based on reinforcement learning according to an embodiment of the present application. As Figure 3 shown, step S2 includes steps: S21, successively input each interaction data in the historical interaction sequence of the first student object into the basic intelligent evaluation model with general parameters to obtain updated output weights and bias definitions as updated general parameters; S22, input the current interaction data into the basic intelligent evaluation model with updated general parameters to obtain the benchmark prediction result.
[0033] Specifically, in step S21, each interaction data in the historical interaction sequence of the first student object is sequentially input into a basic intelligent evaluation model with general parameters to obtain updated output weights and bias definitions as updated general parameters. Specifically, since the learning trajectories of individual students have temporal continuity and the evolution of knowledge states depends on the implicit patterns in historical interaction behaviors. Therefore, in order to achieve initial personalized adaptation while maintaining the generalization ability of the model, this application fine-tunes the general parameters of the basic model online based on the principle of dynamic context awareness, using the historical interaction sequence of the first student object. In an embodiment of this application, the basic intelligent evaluation model is based on the LSTM (Long Short-Term Memory Network) architecture, loading pre-trained general parameters at the initial stage to form a general knowledge reasoning ability across student groups; when processing the target student, the historical interaction sequence of this student (such as the answer results sorted by time, knowledge point interaction labels, and residence duration, etc.) is sequentially input into the basic intelligent evaluation model with general parameters at each time step. Through the temporal memory characteristics of the LSTM network, in the forward propagation process of each time step, the internal representation of the network is dynamically adjusted through the hidden state transfer mechanism (the cell state update of the LSTM network), enabling the model to gradually learn the unique learning pattern and knowledge mastery trend of this student, and using the backpropagation algorithm to update the weights and bias parameters of the output layer of the model (setting the learning rate of other layers to 0), forming a context-aware fine-tuning model for this student. In this way, the basic intelligent evaluation model captures the long-term dependency relationships in individual historical behaviors while retaining the global knowledge structure, thus establishing a dynamically evolving reasoning benchmark for the prediction of current interaction data.
[0034] Specifically, in step S22, the current interaction data is input into the basic intelligent evaluation model with updated general parameters to obtain the benchmark prediction result. That is, the updated output layer weights and biases are locked as the temporary reasoning parameters of the basic intelligent evaluation model, and the current interaction data (such as the latest knowledge point interaction event) is input into the model. At this time, the forward calculation process of the basic intelligent evaluation model with updated general parameters will utilize two types of knowledge simultaneously: the cross-group regular knowledge in the basic general parameters (implicit in the hidden layer weights), and the individual short-term behavior characteristics contained in the output layer parameters after fine-tuning with historical interaction data, finally outputting a benchmark prediction result (such as the answer correct rate prediction value) that reflects both the group common law and integrates the individual historical behavior pattern.
[0035] In the above dynamic optimization method of the intelligent evaluation model based on reinforcement learning, in step S3, calculate the baseline prediction error between the baseline prediction result and the real interaction, and integrate the baseline prediction error, recent performance statistical features, and current personalized adjustment parameters to obtain the current RL state. It should be understood that this application takes into account that there may be a difference between the baseline prediction result obtained by the basic intelligent evaluation model and the real interaction result of the student, and this difference reflects the prediction accuracy of the basic intelligent evaluation model for the current performance of this student. Therefore, in order to enable the reinforcement learning agent to comprehensively perceive the current situation and make reasonable decisions, this application takes the baseline prediction error between the baseline prediction result and the real interaction as the core, quantifies the immediate deviation of the current parameter configuration, combines the recent performance statistical features of the student (such as the change trend of the answering correct rate in the past N interactions, high-frequency error knowledge points, knowledge point coverage entropy value, etc.) and the current personalized adjustment parameters (such as the result of the previous RL Agent decision, which is defaulted to be empty if it is the first time), and after feature splicing and dimension alignment, jointly constructs the current reinforcement learning (RL) state, so that the reinforcement learning agent can simultaneously perceive the immediate error, medium-term trend, and long-term policy impact, and provide comprehensive state information for the subsequent reinforcement learning strategy.
[0036] In the above method for dynamically optimizing an intelligent evaluation model based on reinforcement learning, in step S4, the current RL state is input into the RL Agent policy network to obtain the selected adjustment action. It should be understood that in order to achieve the autonomous evolution of the dynamic optimization strategy, based on the principle of deep reinforcement learning, this application realizes multi-modal adjustment decisions through the action space design of the policy network. Specifically, the RL Agent policy network adopts a two-branch structure: the value function branch is used to evaluate the state value, and the policy branch is used to output the probability distribution of four types of actions (directly output new parameters, output parameter adjustment amounts, select predefined adjustment strategies, maintain the status quo). The network is trained by the proximal policy optimization (PPO) algorithm, and the design of the reward function comprehensively considers the short-term error reduction rate, the achievement degree of long-term learning goals, and the parameter adjustment amplitude penalty term. In the inference stage, the RL Agent policy network calculates the probability distribution of each action according to the current state vector, and selects the optimal action through Gumbel-Softmax sampling. In this way, the RL Agent can intelligently weigh the pros and cons of different adjustment strategies and select actions that can quickly reduce the prediction error while maintaining the model stability and sustainable optimization. For example, when the prediction error is large and the recent performance of the student fluctuates significantly, the policy network may tend to directly output new personalized adjustment parameters to quickly adjust the model state; while when the model prediction is relatively stable and the student's knowledge mastery trend is relatively stable, it may choose to fine-tune the existing parameters or maintain the status quo to avoid the risk of model instability caused by excessive adjustment. Finally, the selected adjustment action will be used as the decision output of the reinforcement learning agent to guide the further personalized adjustment of the basic intelligent evaluation model, forming a closed-loop intelligent evaluation and optimization process.
[0037] In the above method for dynamically optimizing an intelligent evaluation model based on reinforcement learning, in step S5, based on the selected adjustment action, new personalized adjustment parameters are calculated. Among them, Figure 4 is a flowchart of sub-step S5 of the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of the present application. As Figure 4 shown, step S5 includes steps: S51, in response to the selected adjustment action being to output the adjustment amount of the personalized adjustment parameter; S52, fusing the output adjustment amount of the personalized adjustment parameter with the current personalized adjustment parameter to obtain the new personalized adjustment parameter.
[0038] Specifically, in step S51, in response to the selected adjustment action, an amount of adjustment of the personalized adjustment parameter is output. That is, according to the adjustment action selected by the RL Agent policy network, a new personalized adjustment parameter is calculated and determined. If it is selected to directly output the new personalized adjustment parameter, the new personalized adjustment parameter is directly applied; if it is selected to output the amount of adjustment of the personalized adjustment parameter, the parameter is adjusted based on the original parameter to obtain the new personalized adjustment parameter; if it is selected to use a predefined adjustment strategy, the corresponding predefined strategy is called through the strategy library to update the parameter (such as increasing the learning rate, reducing the adjustment step size, etc.); if it is selected not to make an adjustment, the current personalized adjustment parameter remains unchanged. The new personalized adjustment parameter will be used for the prediction of the next interaction data, and this cycle of iteration is continuously performed to dynamically adjust the evaluation model according to the student's learning progress and performance, so as to achieve more effective evaluation and optimization of the student's learning state.
[0039] Specifically, in step S52, the output amount of adjustment of the personalized adjustment parameter is fused with the current personalized adjustment parameter to obtain the new personalized adjustment parameter. In particular, considering that when the selected adjustment action is to output the amount of adjustment of the personalized adjustment parameter, the high-dimensional characteristic of the parameter adjustment space easily leads to a conflict between the adjustment action and the existing configuration, that is, direct numerical superposition may lead to parameter space conflict or semantic distortion. Therefore, in order to capture the implicit association between the output amount of adjustment of the personalized adjustment parameter and the current personalized adjustment parameter and maintain the stability of the adjustment process, the present application further introduces a deep learning algorithm to perform deep interaction fusion on the output amount of adjustment of the personalized adjustment parameter and the current personalized adjustment parameter to generate a new personalized adjustment parameter, so that the generated personalized adjustment parameter not only responds to the current RL state but also conforms to the internal parameter distribution law of the model. Among them, Figure 5 is a flowchart of sub-step S52 of the method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to an embodiment of the present application. As Figure 5 shown, step S52 includes the steps of: S521, performing structured encoding on the output amount of adjustment of the personalized adjustment parameter and the current personalized adjustment parameter to obtain a structured encoding vector of the adjustment amount and a structured encoding vector of the current adjustment parameter; S522, performing multi-scale progressive interaction on the structured encoding vector of the adjustment amount and the structured encoding vector of the current adjustment parameter to obtain a high-dimensional representation of the updated personalized parameter; S523, performing feature decoding on the high-dimensional representation of the updated personalized parameter to obtain the new personalized adjustment parameter.
[0040] More specifically, in step S521, the output personalized adjustment parameter adjustment amount and the current personalized adjustment parameter are structured encoded to obtain a parameter adjustment amount structured encoding vector and a current adjustment parameter structured encoding vector. It should be understood that, considering that the model parameters have high dimensions and complex topological relationships, the present application is based on the graph representation learning principle, and the parameter system is mapped into a resolvable topological space through structured encoding. Specifically, first, the current personalized adjustment parameter (such as the model weight matrix θ_t) and the adjustment amount Δθ output by reinforcement learning are modeled as graph structure data respectively: each graph structure node corresponds to the weight or bias item of a network layer in the model, the node feature contains meta-information such as parameter value and gradient history, and the edge weight is determined by the correlation measure between nodes, such as the gradient correlation in back propagation. Subsequently, the two parameter graphs are hierarchically aggregated using a graph convolutional network (GCN) to generate a parameter adjustment amount structured encoding vector and a current adjustment parameter structured encoding vector. In this way, not only the numerical characteristics of the parameters are retained, but also the functional coupling relationship between the parameters is deeply mined, providing a semantically aligned vector space for subsequent deep interactive fusion.
[0041] More specifically, in step S522, the parameter adjustment amount structured coding vector and the current adjustment parameter structured coding vector are subjected to multi-scale progressive interaction to obtain an updated personalized parameter high-dimensional representation. It should be understood that since parameter adjustment needs to take into account the global distribution law of model parameters and the coordinated changes of local sensitive areas, single-scale feature interaction is prone to information loss or over-smoothing. Therefore, in order to balance the overall consistency and fine-grained adaptability of parameter updates, the present application further processes the parameter adjustment amount structured coding vector and the current adjustment parameter structured coding vector through a hierarchical feature interaction mechanism to achieve multi-granularity knowledge fusion and form a composite feature expression that takes into account global stability and local sensitivity, namely, the updated personalized parameter high-dimensional representation. Among them, Figure 6 Flow chart of sub-step S522 of the dynamic optimization method of the intelligent evaluation model based on reinforcement learning according to an embodiment of the present application. Figure 6As shown, the step S522 includes the steps of: S5221, performing low-level feature interaction on the parameter adjustment amount structured coding vector and the current adjustment parameter structured coding vector to obtain a parameter adjustment amount-current adjustment parameter low-level joint perception feature vector; S5222, performing multi-level high-order feature interaction on the parameter adjustment amount structured coding vector and the current adjustment parameter structured coding vector to obtain a parameter adjustment amount-current adjustment parameter middle-level joint perception feature vector and a parameter adjustment amount-current adjustment parameter deep-level joint perception feature vector; S5223, performing progressive interactive fusion on the parameter adjustment amount-current adjustment parameter low-level joint perception feature vector, the parameter adjustment amount-current adjustment parameter middle-level joint perception feature vector and the parameter adjustment amount-current adjustment parameter deep-level joint perception feature vector to obtain the updated personalized parameter high-dimensional representation.
[0042] In a specific example of the present application, step S5221 is expressed by the formula:
[0043] in, Indicates adding by position point. represents the multi-layer perceptron model, Represents the parameter adjustment amount - the low-level joint perception feature vector of the current adjustment parameter.
[0044] That is, by directly performing low-level interaction on the structured encoding vector of the parameter adjustment amount and the structured encoding vector of the current adjustment parameter, high-frequency detail information and precise correspondence are retained, and the loss of details that may occur in the layer-by-layer abstraction process of the deep network is avoided. In this way, the association between the parameter adjustment amount and the current adjustment parameter in fine dimensions such as pixel level, timestamp level or word level is captured. The obtained low-level joint perception feature vector of the parameter adjustment amount-current adjustment parameter contains the fine correlation information in the original feature that has not been lost by abstraction, providing a more detailed basic feature representation for subsequent parameter optimization.
[0045] In a specific example of the present application, the step S5222 includes: first, performing multi-level implicit feature extraction on the parameter adjustment amount structured coding vector and the current adjustment parameter structured coding vector to obtain a parameter adjustment amount middle-layer implicit feature coding vector, a current adjustment parameter middle-layer implicit feature coding vector, a parameter adjustment amount deep-layer implicit feature coding vector, and a current adjustment parameter deep-layer implicit feature coding vector, which is expressed as:
[0046]
[0047]
[0048]
[0049] Among them, represents the structured encoding vector of the parameter adjustment amount, represents the structured encoding vector of the current adjustment parameter, and respectively represent the weight matrix and bias term of the middle-layer hidden feature encoding vector extraction network, represents the ReLU activation function, represents the middle-layer hidden feature encoding vector of the parameter adjustment amount, represents the middle-layer hidden feature encoding vector of the current adjustment parameter, and respectively represent the weight matrix and bias term of the deep-layer hidden feature encoding vector extraction network, represents the Sigmoid activation function, and respectively represent the deep-layer hidden feature encoding vector of the parameter adjustment amount and the deep-layer hidden feature encoding vector of the current adjustment parameter.
[0050] That is, by utilizing the non-linear transformation and hierarchical feature learning ability of the deep neural network, the structured encoding vector of the parameter adjustment amount and the structured encoding vector of the current adjustment parameter are independently processed respectively to extract hidden features with different abstraction granularities, providing multi-dimensional information support for subsequent cross-feature interaction, so that the model can understand the association between parameter adjustment and the current state from different levels. The middle-layer hidden feature encoding vector of the parameter adjustment amount, the middle-layer hidden feature encoding vector of the current adjustment parameter, the deep-layer hidden feature encoding vector of the parameter adjustment amount, and the deep-layer hidden feature encoding vector of the current adjustment parameter generated thereby respectively represent the internal relationship between the parameter adjustment amount and the current adjustment parameter at the structured information level and the global semantic level, not only retaining the stable structural features in the parameter adjustment process, but also refining the core semantic information related to the dynamic changes of the learning state, providing a hierarchical information basis for subsequent middle-level and deep-level feature interaction.
[0051] Then, middle-level feature interaction is performed on the middle-layer hidden feature encoding vector of the parameter adjustment amount and the middle-layer hidden feature encoding vector of the current adjustment parameter to obtain the middle-level joint perception feature vector of the parameter adjustment amount - current adjustment parameter, which is expressed by the formula:
[0052]
[0053] Among them, is the attention fusion network, represents pointwise multiplication by position, Denote the parameter adjustment amount - the hierarchical joint perception feature vector in the current adjustment parameter, as the feature scale value of, denote the transpose of the vector, is the normalization exponential function.
[0054] That is, by performing middle-level interaction on the middle-level implicit feature encoding vector in the parameter adjustment amount and the middle-level implicit feature encoding vector in the current adjustment parameter, on the basis of filtering redundant information and retaining stable structural patterns, modeling the dependency relationship between the two at the abstract structure level to capture cross-source relationships more complex than the low-level joint perception feature vector of the parameter adjustment amount - the current adjustment parameter, providing structural information support of medium complexity for subsequent parameter optimization, enabling the model to understand the adaptability of parameter adjustment and the current learning state from the perspective of structural association. Based on this, the generated parameter adjustment amount - the hierarchical joint perception feature vector in the current adjustment parameter effectively represents the interaction relationship between the parameter adjustment amount and the current adjustment parameter at the structural level, eliminates the noise interference in the original features, and encodes structural patterns with invariance, so that the model can perform dynamic adjustment based on the structural dependency relationship of medium complexity when updating personalized parameters, improving the response accuracy of the evaluation model to changes in the individual learning state at the structural level.
[0055] Finally, perform deep-level feature interaction on the deep-level implicit feature encoding vector of the parameter adjustment amount and the deep-level implicit feature encoding vector of the current adjustment parameter to obtain the parameter adjustment amount - the deep-level joint perception feature vector in the current adjustment parameter, which is expressed by the formula:
[0056]
[0057] Wherein, and are different low-rank projection matrices, denote the GELU activation function, denote the LayerNorm normalization function, is the weight matrix of the deep-level feature fusion network, denote the parameter adjustment amount - the current adjustment parameter linear projection gating interaction feature, denote the parameter adjustment amount - the deep-level joint perception feature vector in the current adjustment parameter.
[0058] That is, by performing deep interaction on the deep implicit feature encoding vector of the parameter adjustment amount and the deep implicit feature encoding vector of the current adjustment parameter, the collaborative association between the two at the levels of core semantics, global learning state context, or optimization objective is captured. Leveraging the characteristics of high information density and dimension reduction of deep features, semantic-level reasoning and alignment are achieved, providing abstract information support regarding the evolution trend of the learning state and the global significance of parameter adjustment for personalized parameter update, enabling the model to understand the adaptation relationship between parameter adjustment and the overall learning progress of the student from a global semantic perspective. Based on this, the generated deep-level joint perception feature vector of the parameter adjustment amount - current adjustment parameter extracts the core concept of learning state changes in the form of a low-dimensional and high-density representation, providing a semantic-level decision basis for subsequent cross-level feature fusion.
[0059] Figure 7 Flowchart of sub-step S5223 of the dynamic optimization method of the intelligent evaluation model based on reinforcement learning according to an embodiment of the present application. As Figure 7 shown, the step S5223 includes steps: S52231, performing complementary interaction fusion based on a gating mechanism on the low-level joint perception feature vector of the parameter adjustment amount - current adjustment parameter and the middle-level joint perception feature vector of the parameter adjustment amount - current adjustment parameter to obtain a low-middle-level joint perception feature vector of the parameter adjustment amount - current adjustment parameter; S52232, performing cross-level interaction based on a cross-attention mechanism on the low-middle-level joint perception feature vector of the parameter adjustment amount - current adjustment parameter and the deep-level joint perception feature vector of the parameter adjustment amount - current adjustment parameter to obtain the high-dimensional representation of the updated personalized parameter.
[0060] In a specific example of the present application, the step S52231 is represented by the formula:
[0061]
[0062] where represents feature concatenation, represents the weight matrix of the progressive complementary perception network, represents and the gating interaction weight vector between represents the low-middle-level joint perception feature vector of the parameter adjustment amount - current adjustment parameter.
[0063] That is, the gating mechanism is used to dynamically allocate weights to the low-level joint perception feature vector of the parameter adjustment amount-current adjustment parameter and the medium-level joint perception feature vector of the parameter adjustment amount-current adjustment parameter, so as to achieve complementary fusion of the two at the level of detail information and structured pattern. The generated low-level joint perception feature vector of the parameter adjustment amount-current adjustment parameter can adaptively integrate the detail accuracy of the low-level features and the structural stability of the medium-level features, and filter out redundant information and strengthen key associations through the gating mechanism, so that the model can take into account both subtle changes in learning behavior and structural evolution trends when updating personalized parameters, laying a feature foundation with both details and structure for deep-level semantic fusion.
[0064] In particular, in a preferred example of the present application, the step S52232 includes: first, obtaining a query embedding matrix, a key embedding matrix and a value embedding matrix, and performing multi-objective weight collaborative allocation on the query embedding matrix, the key embedding matrix and the value embedding matrix to obtain an optimized query embedding matrix, an optimized key embedding matrix and an optimized value embedding matrix. Here, considering that when the low-level joint perception feature vector of the parameter adjustment amount-current adjustment parameter and the middle-level joint perception feature vector of the parameter adjustment amount-current adjustment parameter are subjected to complementary interactive fusion based on a gating mechanism, the gated interaction weight vector substantially generated by the low-level joint perception feature vector of the parameter adjustment amount-current adjustment parameter and the middle-level joint perception feature vector of the parameter adjustment amount-current adjustment parameter are then fused based on the phase synchronization complementarity of the gated interaction weight vector, and further constrained with the deep-level joint perception coding vector of the edge thermal-geometric deviation via the query embedding matrix, the key embedding matrix and the value embedding matrix to realize interaction. Due to the instability of the constraint deconstruction of the structured measurement space, the phenomenon of limited interaction efficiency is caused. Based on this, the present application further suppresses the non-integrable accumulation of constraint associations in the structured measure space by calculating the field decomposition integral representations of the query embedding matrix, the key embedding matrix and the value embedding matrix, and by using the respective decomposition field integral expressions of the structured measure space, so as to avoid the geometric phase accumulation effect that inhibits the formation of mutual feedback of the local field structure.
[0065] Specifically, first, for the pre-trained matrix , and , taking the combination of the three as the coupled order parameter topological domain, the decomposed field path integral representation of each matrix is calculated:
[0066]
[0067]
[0068] in, , and respectively represent the query embedding matrix, the key embedding matrix, and the value embedding matrix. , and respectively represent the optimized query embedding matrix, the optimized key embedding matrix, and the optimized value embedding matrix.
[0069] Then, define the field state balance alignment loss function:
[0070] Among them, represents calculating the nuclear norm, that is, the sum of the eigenvalues of the matrix. represents the modulation hyperparameter. represents the value of the field state balance alignment loss function.
[0071] In this way, the deconstruction stability regulation can be constrained through the field state balance mapping relationship of the sum of the eigenvalues of the coupled order parameter topological domain of the embedding matrix, so as to realize the interaction preservation of the query embedding matrix, the key embedding matrix, and the value embedding matrix, and generate the optimized query embedding matrix, the optimized key embedding matrix, and the optimized value embedding matrix accordingly.
[0072] Then, use the optimized query embedding matrix to perform embedding encoding on the low-level joint perception feature vector in the parameter adjustment amount - current adjustment parameter to obtain a query vector, and use the optimized key embedding matrix and the optimized value embedding matrix to perform embedding encoding on the deep-level joint perception feature vector in the parameter adjustment amount - current adjustment parameter to obtain a key vector and a value vector, which is expressed by the formula:
[0073]
[0074]
[0075] Among them, , and respectively represent the query vector, the key vector, and the value vector.
[0076] That is, by optimizing the query embedding matrix, the key embedding matrix, and the value embedding matrix to perform spatial mapping on feature vectors at different levels, the low-level jointly perceived feature vectors in the parameter adjustment amount - current adjustment parameters are converted into query vectors that focus on details and structural information. At the same time, the deep-level jointly perceived feature vectors in the parameter adjustment amount - current adjustment parameters are respectively mapped into key vectors and value vectors carrying global semantic associations, so as to construct a feature representation space suitable for cross-level interaction for the subsequent cross-attention mechanism, enabling the model to retrieve key details in the low- and middle-level features targeted at deep-level semantic information, and achieving semantic alignment and weight allocation of multi-scale features.
[0077] Finally, perform cross-level interaction based on the transformer architecture on the query vector, the key vector, and the value vector to obtain the updated high-dimensional representation of the personalized parameter, which is expressed by the formula:
[0078] Where, represents the updated high-dimensional representation of the personalized parameter.
[0079] That is, perform progressive cross-level interaction on the query vector, the key vector, and the value vector to dynamically integrate the complementary information between features of different abstraction granularities, thereby generating an updated high-dimensional representation of the personalized parameter that fuses multi-scale interaction perception, providing a comprehensive feature basis with both detail accuracy, structural stability, and semantic consistency for personalized parameter update, and ensuring that the subsequent parameter adjustment results conform to the subtle evolution of the learning behavior pattern and the overall goal of cognitive development.
[0080] More specifically, in a specific example of the present application, the step S523 includes: inputting the updated high-dimensional representation of personalized parameters into a feature decoder based on a multi-layer perceptron to obtain the new personalized adjustment parameters. Specifically, to ensure that the output parameters comply with the model architecture constraints and maintain numerical stability, the present application further implements an inverse mapping from the feature space to the parameter space through a constrained decoding network. In an embodiment of the present application, the feature decoder based on a multi-layer perceptron (MLP) adopts a staged decoding strategy: the first-layer network projects the updated high-dimensional representation of personalized parameters into an intermediate space with the same dimension as the original parameters, and applies L2 regularization constraints to prevent numerical overflow; the second layer introduces a parameter group normalization mechanism to perform group normalization according to the parameter type (such as weight, bias) or functional module (such as attention layer, fully connected layer); the final output layer uses the Tanh activation function to constrain the numerical range within the interval [-1, 1], and then restores it to the actual parameter range through a learnable scaling coefficient. In addition, a parameter relationship constraint loss function is embedded in the decoding process to maintain the internal structure of the parameter system by minimizing the topological difference between the old and new parameter graphs (such as the KL divergence of the node similarity matrix). In this way, the new personalized adjustment parameters can effectively integrate the optimization direction of the reinforcement learning strategy and strictly follow the inherent distribution law of the model parameters, ensuring that the evaluation model always maintains stability and functional integrity during the dynamic adjustment process.
[0081] In summary, the dynamic optimization method of the intelligent evaluation model based on reinforcement learning according to the embodiments of the present application is elucidated. It fine-tunes the basic intelligent evaluation model and its general parameters according to the historical interaction data of the student object, predicts the current interaction data of the student object with the updated model parameters, and then calculates the baseline prediction error between the baseline prediction result and the real interaction, and integrates the recent performance statistical features and the current personalized adjustment parameters of the student object as the current reinforcement learning (RL) state, and uses the RL Agent policy network to select adjustment actions, so as to calculate and update the personalized adjustment parameters of the student object based on the adjustment actions. Based on the core parameters of the basic evaluation model, this method realizes personalized adjustment for the real-time learning state of each student individual, ensuring both the stability and generalization ability of the basic model and the rapid response and targeted adjustment to the changes in the student individual state.
[0082] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, advantages, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, advantages, effects, etc. are essential for each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and easy understanding, rather than limitations, and the above details do not limit the present invention to necessarily adopt the above specific details to implement.
[0083] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0084] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
[0085] In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.
[0086] Finally, it should be noted that the above description has been given for purposes of illustration and description. In addition, the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A dynamic optimization method for an intelligent evaluation model based on reinforcement learning, characterized in that Including: Obtain the current interaction data of the first student object and the historical interaction sequence of the first student object; Use the basic intelligent evaluation model and its general parameters to predict the current interaction data based on the historical interaction sequence of the first student object to obtain a benchmark prediction result; Calculate the benchmark prediction error between the benchmark prediction result and the real interaction, and integrate the benchmark prediction error, recent performance statistical features, and current personalized adjustment parameters to obtain the current RL state; Input the current RL state into the RL Agent policy network to obtain the selected adjustment action; Calculate new personalized adjustment parameters based on the selected adjustment action.
2. The dynamic optimization method of the intelligent evaluation model based on reinforcement learning according to claim 1, wherein The selected adjustment action includes directly outputting new personalized adjustment parameters, outputting the adjustment amount of personalized adjustment parameters, selecting a predefined adjustment strategy, and making no adjustment.
3. The dynamic optimization method of the intelligent evaluation model based on reinforcement learning according to claim 1, wherein Calculating new personalized adjustment parameters based on the selected adjustment action includes: In response to the selected adjustment action being to output the adjustment amount of personalized adjustment parameters; Fuse the output adjustment amount of personalized adjustment parameters with the current personalized adjustment parameters to obtain the new personalized adjustment parameters.
4. The method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to claim 1, wherein Using the basic intelligent evaluation model and its general parameters to predict the current interaction data based on the historical interaction sequence of the first student object to obtain a benchmark prediction result, including: Input each interaction data in the historical interaction sequence of the first student object into the basic intelligent evaluation model with general parameters in sequence to obtain the updated output weight and bias definition as the updated general parameters; Input the current interaction data into the basic intelligent evaluation model with the updated general parameters to obtain the benchmark prediction result.
5. The method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to claim 3, wherein Fusing the output adjustment amount of personalized adjustment parameters with the current personalized adjustment parameters to obtain the new personalized adjustment parameters, including: Perform structured encoding on the output adjustment amount of personalized adjustment parameters and the current personalized adjustment parameters to obtain a structured encoding vector of the parameter adjustment amount and a structured encoding vector of the current adjustment parameters; Perform multi-scale progressive interaction on the structured encoding vector of the parameter adjustment amount and the structured encoding vector of the current adjustment parameters to obtain a high-dimensional representation of the updated personalized parameters; Perform feature decoding on the high-dimensional representation of the updated personalized parameters to obtain the new personalized adjustment parameters.
6. The method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to claim 5, wherein Performing multi-scale progressive interaction on the structured encoding vector of the parameter adjustment amount and the structured encoding vector of the current adjustment parameters to obtain a high-dimensional representation of the updated personalized parameters, including: Perform low-level feature interaction on the structured encoding vector of the parameter adjustment amount and the structured encoding vector of the current adjustment parameters to obtain a low-level joint perception feature vector of the parameter adjustment amount - current adjustment parameters; Perform multi-level high-order feature interaction on the structured encoding vector of the parameter adjustment amount and the structured encoding vector of the current adjustment parameters to obtain a middle-level joint perception feature vector of the parameter adjustment amount - current adjustment parameters and a deep-level joint perception feature vector of the parameter adjustment amount - current adjustment parameters; Perform progressive interactive fusion on the parameter adjustment amount - current adjustment parameter low - level joint perception feature vector, the parameter adjustment amount - current adjustment parameter middle - level joint perception feature vector, and the parameter adjustment amount - current adjustment parameter deep - level joint perception feature vector to obtain the updated personalized parameter high - dimensional representation.
7. The method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to claim 6, wherein Perform multi - level high - order feature interaction on the parameter adjustment amount structured encoding vector and the current adjustment parameter structured encoding vector to obtain the parameter adjustment amount - current adjustment parameter middle - level joint perception feature vector and the parameter adjustment amount - current adjustment parameter deep - level joint perception feature vector, including: Perform multi - level implicit feature extraction on the parameter adjustment amount structured encoding vector and the current adjustment parameter structured encoding vector to obtain the parameter adjustment amount middle - level implicit feature encoding vector, the current adjustment parameter middle - level implicit feature encoding vector, the parameter adjustment amount deep - level implicit feature encoding vector, and the current adjustment parameter deep - level implicit feature encoding vector; Perform middle - level feature interaction on the parameter adjustment amount middle - level implicit feature encoding vector and the current adjustment parameter middle - level implicit feature encoding vector to obtain the parameter adjustment amount - current adjustment parameter middle - level joint perception feature vector; Perform deep - level feature interaction on the parameter adjustment amount deep - level implicit feature encoding vector and the current adjustment parameter deep - level implicit feature encoding vector to obtain the parameter adjustment amount - current adjustment parameter deep - level joint perception feature vector.
8. The dynamic optimization method for the intelligent evaluation model based on reinforcement learning according to claim 7, characterized in that, Perform progressive interactive fusion on the parameter adjustment amount - current adjustment parameter low - level joint perception feature vector, the parameter adjustment amount - current adjustment parameter middle - level joint perception feature vector, and the parameter adjustment amount - current adjustment parameter deep - level joint perception feature vector to obtain the updated personalized parameter high - dimensional representation, including: Perform complementary interactive fusion based on the gating mechanism on the parameter adjustment amount - current adjustment parameter low - level joint perception feature vector and the parameter adjustment amount - current adjustment parameter middle - level joint perception feature vector to obtain the parameter adjustment amount - current adjustment parameter middle - low - level joint perception feature vector; Perform cross - level interaction based on the cross - attention mechanism on the parameter adjustment amount - current adjustment parameter middle - low - level joint perception feature vector and the parameter adjustment amount - current adjustment parameter deep - level joint perception feature vector to obtain the updated personalized parameter high - dimensional representation.
9. The method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to claim 8, wherein, Perform cross - level interaction based on the cross - attention mechanism on the parameter adjustment amount - current adjustment parameter middle - low - level joint perception feature vector and the parameter adjustment amount - current adjustment parameter deep - level joint perception feature vector to obtain the updated personalized parameter high - dimensional representation, including: Obtain the query embedding matrix, the key embedding matrix, and the value embedding matrix, and perform multi - objective weight collaborative allocation on the query embedding matrix, the key embedding matrix, and the value embedding matrix to obtain the optimized query embedding matrix, the optimized key embedding matrix, and the optimized value embedding matrix; Use the optimized query embedding matrix to perform embedding encoding on the low-level joint perception feature vectors in the parameter adjustment amount - current adjustment parameter to obtain query vectors, and use the optimized key embedding matrix and the optimized value embedding matrix respectively to perform embedding encoding on the deep-level joint perception feature vectors in the parameter adjustment amount - current adjustment parameter to obtain key vectors and value vectors; Perform cross-level interaction based on the transformer architecture on the query vectors, the key vectors and the value vectors to obtain the updated high-dimensional representation of the personalized parameters.
10. The method for dynamically optimizing an intelligent evaluation model based on reinforcement learning according to claim 9, wherein, Perform feature decoding on the updated high-dimensional representation of the personalized parameters to obtain the new personalized adjustment parameters, including: Input the updated high-dimensional representation of the personalized parameters into a feature decoder based on a multi-layer perceptron to obtain the new personalized adjustment parameters.
Citation Information
Patent Citations
Accurate teaching management method and system based on adaptive learning analysis
CN118396804A
Self-adaptive high-temperature damage model construction method and system
CN119783537A
Education effect evaluation method and system based on artificial intelligence
CN119831186A
Personalized dynamic weight teaching evaluation model and parameter adjustment method thereof
CN119991374A
Student psychological evaluation system based on AI reinforcement learning optimization
CN120048517A
Cited By
Intelligent detection system for range hood
CN120252042A
Student psychological health dynamic assessment method and system based on deep learning
CN120753654A