Learning path generation method and system in virtual power plant simulation training, terminal and medium

By collecting trainee feature data to generate competency profiles, updating and optimizing training paths in real time, and combining these with path optimization models to recommend training scenarios, the problem of mismatch between training content and trainees' abilities in virtual power plant simulation training has been solved. This has enabled the generation of personalized and adaptive learning paths, improving training efficiency and resource utilization efficiency.

CN121858637APending Publication Date: 2026-04-14STATE GRID HENAN INTEGRATED ENERGY SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing virtual power plant simulation training systems cannot dynamically adjust to individual differences in trainees' knowledge base, skill level, and decision-making style, resulting in a mismatch between training content and trainees' actual abilities, low learning efficiency, one-sided evaluation results and delayed feedback, and low utilization efficiency of training resources.

Method used

Initial ability profiles are generated by collecting initial feature data of trainees, and operation process data is collected in real time to dynamically update the ability profiles. Training scenarios are intelligently recommended using a path optimization model to generate personalized learning paths. Combined with dual recommendation technologies of content matching and collaborative filtering, training resource allocation is optimized.

Benefits of technology

It enables dynamic adjustment of learning paths based on trainees' real-time ability status, improving the relevance and efficiency of training, achieving comprehensive evaluation and real-time feedback of multi-dimensional data, optimizing the allocation of training resources, and enhancing the effectiveness of trainees' ability improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858637A_ABST
    Figure CN121858637A_ABST
Patent Text Reader

Abstract

The invention relates to the field of power system simulation training, and particularly provides a learning path generation method and system in virtual power plant simulation practical training, a terminal and a medium, and the method comprises the steps: collecting initial feature data of a student, generating an initial ability portrait of the student, and carrying out the matching and recommendation of a first training scene from a preset scene library; in the process of executing the recommended training scene by the student, collecting operation process data in real time; dynamically updating the initial ability portrait of the student based on the operation process data, and generating a real-time ability portrait of the student; inputting the real-time ability portrait of the student, the completed training scene sequence and the operation process data into a path optimization model, and outputting a next recommended training scene by the path optimization model; and generating and outputting a learning path report. According to the method, personalized and self-adaptive intelligent generation of the virtual power plant simulation training learning path is realized, and the training pertinence, efficiency and resource allocation effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system simulation training, specifically to a method, system, terminal, and medium for generating learning paths in virtual power plant simulation training. Background Technology

[0002] Virtual Power Plants (VPPs), as core carriers for aggregating distributed energy resources and participating in grid regulation and electricity market transactions, are facing increasing operational complexity and professional requirements. Cultivating professionals with VPP operation capabilities has become an urgent industry need. Currently, VPP simulation training systems are being used in training, but existing systems mostly adopt uniform and fixed training processes and scenario sequences, failing to dynamically adjust according to individual differences in trainees' knowledge base, skill levels, and decision-making styles. This leads to a mismatch between training content and trainees' actual abilities, resulting in low learning efficiency and difficulty in achieving personalized instruction. Traditional training assessments rely heavily on final theoretical exams or operational results, lacking real-time collection and analysis of multi-dimensional data such as operational behavior, decision-making logic, and response efficiency during training. Assessment results are one-sided and feedback is delayed, failing to provide timely and accurate improvement guidance for trainees. Although there are resource libraries containing various training scenarios, scenario recommendations are usually based on simple tags or manual selection, failing to deeply and intelligently match with trainees' dynamically changing ability profiles. Training content is not strongly correlated with the trainees' most pressing skill areas, resulting in low utilization efficiency of training resources. The training process is usually a one-off, linear process. The system cannot dynamically adjust the subsequent training content, difficulty, and strategies based on the trainees' real-time performance. Once the learning path is set, it is difficult to change, and there is a lack of a system that can self-optimize based on the trainees' learning outcomes. Summary of the Invention

[0003] To address the aforementioned issues, this invention provides a method, system, terminal, and medium for generating learning paths in virtual power plant simulation training. This enables personalized, adaptive, and intelligent generation of learning paths for virtual power plant simulation training, thereby improving the relevance, efficiency, and resource allocation effectiveness of the training.

[0004] In a first aspect, the technical solution of the present invention provides a method for generating learning paths in virtual power plant simulation training, comprising the following steps: S1. Collect initial characteristic data of trainees, including identity information, historical training records, pre-test scores and decision-making style questionnaire results, and generate an initial ability profile of trainees based on the initial characteristic data. S2, based on the student's initial ability profile, matches and recommends the first training scenario from a pre-set scenario library; S3 collects operation process data in real time while the trainee performs the recommended training scenario; S4, based on the operation process data, dynamically updates the initial ability profile of the trainee to generate a real-time ability profile of the trainee; S5 inputs the student's real-time ability profile, the completed training scenario sequence, and operation process data into the path optimization model, and the path optimization model outputs the next recommended training scenario. S6. Repeat steps S3 to S5 until the preset training termination conditions are met, and generate and output a learning path report.

[0005] Secondly, the technical solution of the present invention provides a learning path generation system for virtual power plant simulation training, comprising: The initial competency profile generation module is used to collect the trainees' initial characteristic data, including identity information, historical training records, pre-test scores and decision-making style questionnaire results, and generate the trainees' initial competency profile based on the initial characteristic data. The first training scenario recommendation module is used to match and recommend the first training scenario from a pre-set scenario library based on the student's initial ability profile. The operation process data acquisition module is used to collect operation process data in real time during the training scenario performed by the trainee; The real-time competency profile generation module is used to dynamically update the trainee's initial competency profile based on the operation process data, and generate the trainee's real-time competency profile. The recommended training scenario output module is used to input the student's real-time ability profile, the completed training scenario sequence and operation process data into the path optimization model, and the path optimization model outputs the next recommended training scenario. The learning path report output module is used to repeatedly execute the operation process data acquisition module, real-time capability profile generation module, and recommended training scenario output module until the preset training termination conditions are met, and then generate and output a learning path report.

[0006] Thirdly, the technical solution of the present invention provides a terminal, comprising: The memory is used to store the learning path generation program in the virtual power plant simulation training; The processor is used to implement the steps of the learning path generation method in the virtual power plant simulation training when executing the learning path generation program in the virtual power plant simulation training.

[0007] Fourthly, the present invention provides a computer-readable storage medium storing a learning path generation program for virtual power plant simulation training. When the learning path generation program for virtual power plant simulation training is executed by a processor, it implements the steps of the learning path generation method for virtual power plant simulation training as described above.

[0008] As can be seen from the above technical solutions, this application has the following advantages: (1) By constructing a multi-dimensional profile of trainees’ abilities and dynamically updating it based on operational data during the training process, the system can accurately perceive the real-time ability status and weaknesses of each trainee. Based on this, it can intelligently recommend the most suitable training scenario, generate and continuously optimize a personalized learning path for each trainee, and greatly improve the training’s relevance and learning efficiency. (2) It not only focuses on training results, but also pays more attention to the collection and analysis of multi-dimensional data in the operation process, realizing a comprehensive evaluation of trainees’ knowledge, skills and strategic abilities, achieving dynamic and real-time assessment, and providing timely feedback on trainees’ performance, providing an accurate basis for dynamic adjustment of the path; (3) By integrating content matching and collaborative filtering for dual recommendation, and combining a path optimization model based on reinforcement learning for global decision-making, the system can accurately select the next training scenario that can best promote the improvement of students' abilities from a massive scenario library, so that training resources can be optimally allocated. Attached Figure Description

[0009] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of a learning path generation method in a virtual power plant simulation training provided by an embodiment of the present invention.

[0011] Figure 2 This is a schematic block diagram of a learning path generation system in a virtual power plant simulation training exercise, provided as an embodiment of the present invention.

[0012] Figure 3 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present invention. Detailed Implementation

[0013] To make the purpose, features, and advantages of this application more apparent and understandable, specific embodiments and accompanying drawings will be used to clearly and completely describe the technical solution protected by this application. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this application and in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0015] Figure 1 This is a schematic flowchart illustrating a learning path generation method in a virtual power plant simulation training exercise, as provided in an embodiment of the present invention. Figure 1 The executing entity can be a learning path generation system for virtual power plant simulation training. The learning path generation method for virtual power plant simulation training provided in this embodiment of the invention is executed by a computer device; correspondingly, the learning path generation system for virtual power plant simulation training runs on the computer device. Depending on different needs, the order of the steps in this flowchart can be changed, and some steps can be omitted.

[0016] like Figure 1 As shown, the method includes the following steps.

[0017] S1. Collect initial characteristic data of trainees, including identity information, historical training records, pre-test scores and decision-making style questionnaire results, and generate an initial ability profile of trainees based on the initial characteristic data.

[0018] S2, based on the trainee's initial ability profile, matches and recommends the first training scenario from a pre-set scenario library.

[0019] S3 collects operation data in real time as the trainee performs the recommended training scenario.

[0020] S4 dynamically updates the trainee's initial competency profile based on operational process data, generating a real-time competency profile for the trainee.

[0021] S5 inputs the trainee's real-time ability profile, the completed training scenario sequence, and operation process data into the path optimization model, which then outputs the next recommended training scenario.

[0022] S6. Repeat steps S3 to S5 until the preset training termination conditions are met, and generate and output a learning path report.

[0023] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process in this embodiment, another method for generating learning paths in virtual power plant simulation training is provided, which includes the following steps.

[0024] S101: Collect initial characteristic data of trainees and generate initial ability profiles of trainees.

[0025] Initial feature data includes identity information, historical training records, pre-test scores, and decision-making style questionnaire results.

[0026] Identity Information: Basic identity information of students is obtained through student registration or import from the management backend. This information is mainly used for student identification and basic grouping, and provides background dimensions for subsequent profile features, which may include: job role, years of work experience, and affiliated unit / department.

[0027] Historical Training Records: If a trainee has previously used this system or related training platforms, the system will extract relevant records from their historical learning archives as an important reference for assessing their existing skill level. The extracted content includes: a list of completed training scenarios, key performance indicators (KPIs) for historical operations, and past competency assessment results. The list of completed training scenarios records all scenarios completed by the trainee in the past, along with the completion time and final mastery score. Key performance indicators (KPIs) for historical operations refer to statistical values ​​extracted for each completed scenario, such as average operational accuracy score, average task completion time, and decision compliance rate. Past competency assessment results can be the competency profile score at the end of the most recent training session.

[0028] Pre-test Results: Before the training begins, trainees are required to complete a targeted online pre-test. This pre-test combines theoretical questions with basic operational simulations to quickly assess their understanding of the core knowledge of virtual power plant operation. The pre-test content includes basic virtual power plant concepts, market trading rules, basic equipment principles, interface function recognition, and basic operation process sequencing. After the pre-test, the system automatically grades the test and decomposes the score into a score vector across multiple knowledge dimensions according to a pre-defined knowledge point classification system.

[0029] Decision-Making Style Questionnaire Results Collection: To capture trainees' potential tendencies in strategy formulation and risk decision-making, the system sends out a short psychological questionnaire on decision-making styles before the training begins. This questionnaire is designed based on a mature theoretical framework of decision-making styles, guiding trainees to demonstrate their preferences in virtual power plant operation-related decisions (such as pricing strategy selection and risk response methods) through situational judgment questions or scale-based questions. After the questionnaire is submitted, the system outputs quantified strategy preference indices through a scoring model, such as "Risk Tolerance Index," "Strategic Innovation Tendency Index," and "Rule Compliance Index."

[0030] After collection, the initial feature data will undergo the following unified processing for subsequent steps: data cleaning and formatting, feature vectorization, and temporary storage and association. Feature vectorization refers to converting non-numerical data into numerical feature vectors through encoding techniques; numerical data is then standardized or normalized to eliminate the influence of units. The processed feature data is temporarily stored in the student's temporary feature pool and strongly associated with the student's unique identifier (User ID), providing a structured input data source for the next step of "generating the student's initial competency profile."

[0031] The process of generating an initial competency profile of trainees based on the initial feature data includes the following steps S101.1 to S101.4.

[0032] S101.1, the identity information and historical training records in the initial feature data are vectorized and encoded to generate a basic feature vector; the pretest scores are classified according to knowledge points to form an initial score vector for the knowledge dimension; the decision style questionnaire results are quantified into a strategy preference vector.

[0033] Identity information and historical training records are vectorized and encoded, for example, job type and years of work experience are mapped to numerical feature vectors to obtain basic feature vectors. .

[0034] The pre-test scores were categorized by the subjects assessed, forming an initial score vector for each knowledge dimension. ,in Representing the The score for each knowledge point.

[0035] The results of the decision-making style questionnaire were quantified into a strategy preference vector. ,in Represents a risk appetite index. This index represents a conservative strategy inclination. This represents an index indicating a propensity for innovative strategies.

[0036] S101.2 Establish a three-dimensional evaluation system encompassing knowledge, skills, and strategy dimensions; use the analytic hierarchy process (AHP) to determine the initial static weights for each dimension.

[0037] Establish from the knowledge dimension Skills dimension and strategy dimensions The three-dimensional assessment system comprises: Skills Dimension. In the initial stage, the operation proficiency index is initialized based on the historical training records.

[0038] Using the analytic hierarchy process (AHP), domain experts were invited to conduct pairwise comparisons of the importance of the three dimensions, constructing a judgment matrix. After passing a consistency check, the initial static weights of each dimension were calculated. .

[0039] S101.3 Based on the initial score vector and strategy preference vector of the knowledge dimension, as well as the historical operation records, fuzzy evaluation matrices for the knowledge, skill and strategy dimensions are constructed respectively; the weights of each dimension and the corresponding fuzzy evaluation matrices are fuzzy synthesized to obtain the comprehensive fuzzy evaluation results of each dimension; the comprehensive fuzzy evaluation results are defuzzified to obtain the quantitative scores of each dimension.

[0040] Regarding the knowledge dimension Previous test score vector As input, four evaluation levels are set: "Excellent", "Good", "Pass", and "Weak". A membership function is constructed, and the membership function is calculated. The membership degrees of each level form a fuzzy evaluation matrix R_k for the knowledge dimension.

[0041] Regarding the skill dimension In the absence of initial operational data, a fuzzy evaluation matrix R_s is generated based on the historical operational accuracy and completion rate indicators in historical training records, with reference to the knowledge dimension method.

[0042] For the strategy dimension , policy preference vector The model is matched with a pre-defined library of typical strategy patterns. By calculating the Euclidean distance or cosine similarity, the membership degree of the model on each typical pattern is obtained, forming a fuzzy evaluation matrix R_a for the strategy dimension.

[0043] The weight vectors are respectively Perform fuzzy synthesis operations (such as using the M(•, +) operator) with the corresponding fuzzy evaluation matrix (R_k, R_s, R_a) to obtain the comprehensive fuzzy evaluation results for each dimension. The weighted average method was used to evaluate the comprehensive fuzzy evaluation results. Perform deblurring to convert it into a specific score. This serves as a quantitative score for each dimension.

[0044] S101.4 concatenates the basic feature vector with the quantitative scores of each dimension to form a vectorized representation of the student's initial ability profile.

[0045] Quantify the score Quantifying identity features, also known as basic feature vectors The features are concatenated to form a multi-dimensional feature vector, which serves as the mathematical representation of the trainee's initial ability profile. .

[0046] S102, output the first training scene.

[0047] The pre-built scenario library adopts a structured data storage architecture, including a scenario metadata table, a scenario content table, and a capability tag table.

[0048] Scene metadata table: Stores basic information for each training scene, including scene unique identifier, scene name, creation time, last update time, version number, scene type, etc.

[0049] Scene Content Table: Stores the specific training content of the scene, including simulation model parameters, training task description, evaluation criteria, operation procedure instructions, etc.

[0050] Capability Tag Table: This table labels each scenario with multi-dimensional capability requirement tags, including: Knowledge dimension tags: Related theoretical knowledge points required for operating a virtual power plant, such as electricity market trading rules, equipment working principles, and power grid operation principles; Skills dimension tags: Identify the required operational skills, such as equipment control, strategy formulation, and troubleshooting; Strategy dimension tags: These indicate the requirements of decision-making thinking, such as risk aversion, profit optimization, and collaborative scheduling. Difficulty levels: Divided into three levels: beginner, intermediate, and advanced, corresponding to different levels of complexity; Training duration: The estimated average time required to complete the scenario.

[0051] The scenario library is constructed using the following process: a) Based on the actual business needs of virtual power plant operation, training objectives are determined; typical work scenarios are identified, and key skill points are determined; complex business processes are broken down into independently trainable sub-scenarios. b) 3D visualization scenarios are developed based on simulation engines such as Unity3D / Unreal Engine; power system simulation models (such as Matlab / Simulink, RT-LAB, etc.) are integrated; and an interactive operation interface is developed to support trainees in equipment control, strategy formulation, and other operations. c) Domain experts annotate each scenario in multiple dimensions, including knowledge dimension, skill dimension, strategy dimension, difficulty coefficient, prerequisite scenarios, and related scenarios. d) Based on scenario tags, a standardized capability requirement vector is generated for each scenario, quantifying the tag requirements of the three dimensions of knowledge, skills, and strategies into numerical requirement scores, and establishing a mapping table between scenario tags and capability dimension scores.

[0052] This step, based on the trainee's initial ability profile, matches and recommends the first training scenario from a pre-set scenario library, specifically including the following steps S102.1 to S102.4.

[0053] S102.1 Extract the ability score vector from the trainee's initial ability profile. The elements of this vector are the quantitative scores of each dimension. For each candidate scenario in the scenario library, obtain its preset ability requirement vector. The elements of this vector are the preset minimum requirement scores of each dimension. Calculate the cosine similarity between the ability score vector and the ability requirement vector to obtain the content matching score between the trainee and the candidate scenario.

[0054] From the initial ability profile of the trainees Extract the scores from each dimension to construct the current ability score vector. Simultaneously, the system reads the preset ability requirement vector for each candidate scenario from the scenario library. This vector contains three elements: the minimum required scores for knowledge, skills, and strategies in that scenario. The cosine similarity is calculated between the student's ability score vector and the preset ability requirement vector of the candidate scenario to obtain the content matching score between the student and that candidate scenario.

[0055] S102.2, retrieve the set of historical students that are similar to the current student's initial ability profile, count the frequency of each candidate scene in the set being selected as the first training scene, divide the selection frequency of the candidate scene by the maximum selection frequency, and obtain the collaborative filtering recommendation score of the candidate scene.

[0056] The N historical trainees with the highest similarity to the current trainee in terms of initial ability profile are retrieved from the historical training database and formed into a set of similar trainees. The similarity between trainees is calculated based on the cosine similarity of their initial ability profiles.

[0057] We statistically analyze the scenes selected by all trainees in the similar trainee set during their first training session, and calculate the frequency of selection for each candidate scene, denoted as . , For the j-th candidate scene, calculate the collaborative filtering recommendation score using the following formula. :

[0058] in This represents all candidate scenes in the scene library.

[0059] S102.3, the content matching score and the collaborative filtering recommendation score are weighted and summed to obtain the final recommendation score for each candidate scenario.

[0060] S102.4, the candidate scenes with content matching scores not less than the preset first matching threshold are formed into a preliminary screening set, and the candidate scene with the highest final recommendation score is selected from the preliminary screening set as the first training scene for the student.

[0061] A first matching threshold is set. Initial screening is performed based on the content matching score, that is, candidate scenarios with a content matching score greater than or equal to the matching threshold constitute the initial screening set. Then, collaborative screening is performed based on the final recommendation score, that is, the candidate scenario with the highest final recommendation score is selected from the initial screening set as the student's first training scenario.

[0062] S103 collects operation process data in real time.

[0063] During the training scenarios recommended to trainees, operational process data is collected in real time, including operation sequence data, response time data, decision result data, and market return indicator data.

[0064] The interface operation listener captures all interactive operations of the trainees on the simulation interface, including: mouse clicks, drags, scroll wheel operations, keyboard input and shortcut key usage, and touch screen gesture operations.

[0065] Operation events are recorded in chronological order by an operation sequence recorder to form a structured operation sequence.

[0066] High-frequency time-series data is collected through a time-series data acquisition device, including: device status change time points, control command sending time, and system response delay time.

[0067] (1) Manipulating sequence data Identify the types of operation events, including: equipment control operations (such as starting and stopping equipment, adjusting output), strategy formulation operations (such as setting prices, adjusting parameters), monitoring and viewing operations (such as viewing curves, viewing reports), and fault handling operations (such as confirming alarms, executing contingency plans).

[0068] Record operation-related parameter values, such as: target settings for device control, specific values ​​of strategy parameters, and operation timestamps.

[0069] (2) Response time data Decision response time: The time interval from when the problem is presented to when the student takes their first action.

[0070] Operation completion time: The time from the start of the operation to the completion of all operation steps.

[0071] Critical operation delay: Statistics on response delays to critical operation nodes.

[0072] Thinking time distribution: Analyze the thinking time patterns of trainees before different types of operations.

[0073] (3) Decision outcome data Decision correctness assessment: Based on pre-set standard answers or expert rules, assess the correctness of decisions.

[0074] Decision quality scoring: Quantitatively assess the quality of decisions based on economic benefit indicators (such as revenue and cost), safe operation indicators (such as voltage over-limit and frequency deviation), and compliance indicators.

[0075] Decision consistency analysis: Analyzes the degree of consistency in the decisions made by trainees in similar situations.

[0076] (4) Market return indicator data Based on a virtual electricity market simulation model, the real-time revenue of trainees' operations is calculated.

[0077] Risk metrics for calculation operations include: return volatility, maximum drawdown, and Sharpe ratio.

[0078] S104, real-time updates of trainee competency profiles.

[0079] This step dynamically updates the trainee's initial competency profile based on the operation process data, generating a real-time competency profile for the trainee. Specifically, it includes the following steps S104.1 to S104.6.

[0080] S104.1, Analyze and extract a set of quantifiable key performance indicators from the operational process data. The indicator set includes time-based, accuracy-based, result-based, and strategy-based indicators.

[0081] The collected operational process data (including operation sequences, response times, decision results, and market return indicators) are analyzed and features are extracted to obtain a set of quantifiable performance indicators. .

[0082] The current set of indicators includes, but is not limited to: Time-related metrics: average response time, critical decision delay; Accuracy-related indicators: compliance rate of operation sequence, deviation of set value; Outcome-related metrics: Market return achievement rate, risk control score; Strategy-related indicators: Strategy consistency index, adoption rate of innovative strategies.

[0083] S104.2, Construct the mapping matrix Matrix elements Indicators The contribution weights to the ability dimension d, where the ability dimension includes the knowledge dimension. Skills dimension With strategy dimension .

[0084] matrix Each row in the table corresponds to a performance metric. Each column corresponds to a capability dimension, including the knowledge dimension. Skills dimension With strategy dimension .

[0085] Matrix elements Indicators The contribution weights to capability dimension d are obtained by training with expert experience or historical data, and satisfy the following: .

[0086] S104.3, set the key performance indicators Normalization yields Calculate the performance increment vector for each capability dimension using the following formula. :

[0087] Specifically, the performance metrics are normalized to the [0,1] interval to obtain a standardized vector. .

[0088] according to Calculate the performance increment vector for each capability dimension ,in, Represents matrix multiplication. These represent the performance increments in the knowledge, skills, and strategy dimensions, respectively.

[0089] S104.4, based on performance increment vector For the initial static weights Make dynamic adjustments to obtain the current dynamic weights. :

[0090] in, These are preset adjustment coefficients used to control the adjustment range; j=k,s,a, representing the knowledge dimension, skill dimension, and strategy dimension, respectively.

[0091] S104.5, based on the aforementioned performance increment vector With current dynamic weights Update scores for each dimension:

[0092] in, The learning rate parameter controls the step size for score updates. The score of the j-th dimension of the initial capability profile.

[0093] S104.6 concatenates the updated score with the basic feature vector to generate a real-time ability profile of the trainee.

[0094] The real-time ability profile of trainees is represented as follows .

[0095] Through the above process, the trainee's initial competency profile has been dynamically updated based on their operational data, forming a real-time competency profile that reflects their current true level. This profile serves as real-time input for subsequent recommendations and path optimization.

[0096] S105, output the next recommended training scenario based on the path optimization model.

[0097] In this embodiment, a path optimization model based on reinforcement learning is employed. A reinforcement learning environment is pre-constructed, including the definition of the state space, the definition of the action space, and the design of the reward function.

[0098] (1) The state space definition includes: Student ability status: a three-dimensional ability score vector; Training history status: encoded representation of trained scene sequences, sub-sequences of training performance for each scene, and training duration statistics; Learning progress status: number of training scenarios completed, current training round, and cumulative training time; Scenario library status: number of remaining trainable scenarios, and the matching degree between each scenario and the student's current ability.

[0099] (2) Definition of action space The motion space includes the following optional actions: Scene recommendation action: Select a scene from the candidate scenes for recommendation; Difficulty Adjustment Action: Adjust the difficulty of the current training scenario (increase / decrease / maintain); Assistive intervention actions: Insert supplementary learning materials or prompts; Recommended rest action: It is recommended that trainees pause training and rest before continuing.

[0100] (3) Reward function design It adopts a multi-objective design, including rewards for enhancing capabilities and rewards for supporting objectives.

[0101] Ability enhancement rewards are based on the magnitude of the improvement in ability scores, and are expressed as:

[0102] in, , , These are the weighting coefficients for each dimension.

[0103] The auxiliary objective reward refers to the training efficiency reward, which is based on the improvement in ability per unit of time, and is expressed as:

[0104] in Rewards for training efficiency; The change in ability refers to the extent to which a trainee's overall ability score improves during a training period. Training time refers to the actual time spent completing this training session.

[0105] Set constraints, including: difficulty suitability constraint, which is to penalize when the difficulty of the scenario does not match the student's ability; scenario diversity constraint, which is to penalize when the training scenario type is too monotonous; and learning fatigue constraint, which is to penalize when the training time is too long.

[0106] During training, complete training data from at least 1,000 trainees is collected. Historical data is cleaned, normalized, and serialized. Experts are invited to conduct demonstration training to generate high-quality training trajectories. The training set, validation set, and test set are divided in a 7:2:1 ratio.

[0107] The neural network architecture is designed to include a state encoder, a policy network, and a value function. The state encoder consists of a 3-layer fully connected network with ReLU activation; the policy network outputs the action probability distribution using Softmax activation; and the value network evaluates the state value for calculating the advantage function.

[0108] The hyperparameters were set as follows: learning rate of 0.001, Adam optimizer, discount factor of 0.99, batch size of 64, and replay buffer size of 10000. Proximal Policy Optimization (PPO) algorithm was used for training.

[0109] The real-time ability profile of the trainees, the sequence of completed training scenarios, and the operation process data are input into the path optimization model, which then outputs the next recommended training scenario.

[0110] S105.1, the current state vector is formed by concatenating the ability scores of each dimension in the real-time ability profile of the trainee, the encoded features of the completed training scene sequence, and the scene performance statistical features extracted from the operation process data.

[0111] Obtain the real-time ability profile of the trainees Extract real-time capability vectors from them .

[0112] Record the completed training scene sequences as an ordered list. ,in q is a unique identifier for completed scenarios, and q is the number of completed scenarios.

[0113] For each completed scenario Extract key performance indicator vectors from the corresponding operational process data. And calculate its overall performance score in this scenario. The It can be calculated by a preset weighting function. The weighted sum of the various indicators is obtained.

[0114] The above information is integrated into a state vector. Its composition is as follows:

[0115] in, For scene sequences Encoding functions, such as one-hot coding, embedding coding, or scene-feature-based aggregation statistics; To differentiate the performance of the sequence Statistical feature extraction functions, such as mean, variance, and recent trend.

[0116] S105.2, the current state vector is input into the path optimization model trained by reinforcement learning; the model processes the state vector through its internal state encoder and policy network, and outputs a probability distribution of multiple optional actions including "recommend the next training scenario"; the current optimal action is selected according to the probability distribution.

[0117] The path optimization model is a neural network model trained based on a reinforcement learning framework. Its network structure includes a state encoder, a policy network, and a value network.

[0118] The state vector The input path optimization model's state encoder outputs a high-dimensional state representation. ,Will Input the policy network π, and output the probability distribution of the available actions in the current state. . The set of optional actions is defined as A={a1,a2,…,aT}, where aT represents “recommend the next training scenario”, and other actions may include “adjust the scenario difficulty”, “insert auxiliary learning materials”, etc.

[0119] S105.3 If the selected optimal action is "recommend the next training scene", then through the scene generation sub-network in the model, a recommendation weight vector corresponding to all scenes in the scene library is output.

[0120] According to probability distribution If the current optimal action is "Recommend the next training scenario", then the following steps are executed; otherwise, the corresponding processing flow is executed, such as calling the difficulty adjustment module or the learning material library.

[0121] When the action is "recommend the next training scenario", the state representation will be... The input is fed into the scene generation subnetwork, which outputs a recommendation weight vector corresponding to the candidate scenes in the scene library.

[0122] S105.4 Calculate the real-time matching degree between the student's real-time ability vector and the preset ability requirement vector for each scenario in the scenario library; based on the preset second matching degree threshold, select scenarios that meet the matching degree to form a subset of candidate recommended scenarios.

[0123] Similar to step S102.1, cosine similarity is used to calculate the real-time matching degree between the student's real-time ability vector and the preset ability requirement vector of each scene in the scene library. Scenes with a real-time matching degree not less than the second matching degree threshold are screened out as candidate recommended scenes, forming a subset of candidate recommended scenes.

[0124] S105.5 In the subset of candidate recommendation scenarios, for each scenario, its corresponding weight in the recommendation weight vector is weighted and fused with its real-time matching degree to obtain the final selection score of the scenario; the scenario with the highest final selection score is selected as the next recommendation training scenario.

[0125] The final selection score is calculated as follows:

[0126] in, For the scene in the recommendation weight vector The weight, The students and scenarios calculated in step S105.4 Real-time matching accuracy These are the model weight coefficients.

[0127] S106, Output the learning path report.

[0128] Repeat steps S103 to S105 until the preset training termination conditions are met, and generate and output a learning path report.

[0129] S106.1 Collect and integrate various types of data generated throughout the training process, including trainee information, initial competency profiles, executed training scenario sequences, operation process data corresponding to each scenario, decision logs of the path optimization model, and evolution records of trainees' real-time competency profiles, to form a structured original dataset of learning paths.

[0130] S106.2 generates visual analysis content based on the original dataset of the learning path, including the changing trends of learners' ability scores in the three dimensions of knowledge, skills, and strategies as the training process changes, and displays the learners' comprehensive mastery scores in different types of training scenarios in matrix form.

[0131] From the evolution sequence of competency profiles, score changes in the three dimensions of knowledge, skills, and strategies are extracted to form a time series. , where T is the total training duration or the number of key nodes.

[0132] Generate a capability growth curve using a line chart or area chart. The horizontal axis represents time or training scenario sequence, and the vertical axis represents capability score. The three curves correspond to the three capability dimensions, and key improvement points and regression points are marked. Additional annotations can be added to the graph to explain the specific training scenarios or events associated with significant score changes.

[0133] For each trained scenario, a comprehensive mastery score is calculated based on key performance indicators from the operational process data. The calculation formula is a weighted sum and normalized sum of all indicators to the [0,1] interval. A two-dimensional matrix is ​​constructed with the training scenario as the rows and the ability dimension or operational task type as the columns. The matrix elements represent the mastery scores for the corresponding scenario in that dimension or task. A mastery heatmap is generated using a gradient color scheme, with the color level continuously changing from low mastery (cool colors) to high mastery (warm colors), visually displaying the distribution of trainees' performance in different types of scenarios.

[0134] The foregoing has described in detail an embodiment of a learning path generation method in virtual power plant simulation training. Based on the learning path generation method in virtual power plant simulation training described in the above embodiment, this invention also provides a learning path generation system in virtual power plant simulation training corresponding to the method.

[0135] Figure 2 This is a schematic block diagram of a learning path generation system for virtual power plant simulation training provided by an embodiment of the present invention. In this embodiment, the learning path generation system 200 for virtual power plant simulation training can be divided into multiple functional modules according to its functions. A module, as referred to in this invention, is a series of computer program segments that can be executed by at least one processor and perform a fixed function, and is stored in memory.

[0136] The initial competency profile generation module 210 is used to collect the trainees' initial characteristic data, including identity information, historical training records, pre-test scores and decision-making style questionnaire results, and generate the trainees' initial competency profile based on the initial characteristic data.

[0137] The first training scenario recommendation module 220 is used to match and recommend the first training scenario from a preset scenario library based on the student's initial ability profile.

[0138] The operation process data acquisition module 230 is used to collect operation process data in real time during the training scenario recommended by the trainee.

[0139] The real-time competency profile generation module 240 is used to dynamically update the trainee's initial competency profile based on the operation process data, and generate the trainee's real-time competency profile.

[0140] The recommended training scenario output module 250 is used to input the trainee's real-time ability profile, the completed training scenario sequence and operation process data into the path optimization model, and the path optimization model outputs the next recommended training scenario.

[0141] The learning path report output module 260 is used to repeatedly execute the operation process data acquisition module 230, the real-time capability profile generation module 240, and the recommended training scenario output module 250 until the preset training termination conditions are met, and to generate and output the learning path report.

[0142] The learning path generation system in the virtual power plant simulation training of this embodiment is used to implement the aforementioned learning path generation method in the virtual power plant simulation training. Therefore, the specific implementation of this system can be found in the embodiment section of the virtual power plant simulation training learning path generation method above. So, its specific implementation can be referred to the description of the corresponding embodiments, and will not be described in detail here.

[0143] Furthermore, since the learning path generation system in the virtual power plant simulation training of this embodiment is used to implement the aforementioned learning path generation method in the virtual power plant simulation training, its function corresponds to the function of the above method, and will not be repeated here.

[0144] Figure 3 This is a schematic diagram of a terminal 300 provided in an embodiment of the present invention, including: a processor 310, a memory 320, and a communication unit 330. The processor 310 is used to implement the process steps of the above-described embodiment of the learning path generation method in virtual power plant simulation training when implementing the learning path generation program in virtual power plant simulation training stored in the memory 320.

[0145] This invention also provides a computer storage medium, which may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. The computer storage medium stores a learning path generation program for virtual power plant simulation training. When the learning path generation program is executed by a processor, it implements the process steps of the above-described embodiment of the learning path generation method for virtual power plant simulation training.

[0146] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating learning paths in virtual power plant simulation training, characterized in that, Includes the following steps: S1. Collect initial characteristic data of trainees, including identity information, historical training records, pre-test scores and decision-making style questionnaire results, and generate an initial ability profile of trainees based on the initial characteristic data. S2, based on the student's initial ability profile, matches and recommends the first training scenario from a pre-set scenario library; S3 collects operation process data in real time while the trainee performs the recommended training scenario; S4, based on the operation process data, dynamically updates the initial ability profile of the trainee to generate a real-time ability profile of the trainee; S5 inputs the student's real-time ability profile, the completed training scenario sequence, and operation process data into the path optimization model, and the path optimization model outputs the next recommended training scenario. S6. Repeat steps S3 to S5 until the preset training termination conditions are met, and generate and output a learning path report.

2. The learning path generation method in virtual power plant simulation training according to claim 1, characterized in that, Based on the initial feature data, an initial competency profile of the trainee is generated, specifically including: The identity information and historical training records in the initial feature data are vectorized and encoded to generate a basic feature vector; the pre-test scores are classified according to knowledge points to form an initial score vector for the knowledge dimension; and the results of the decision-making style questionnaire are quantified into a strategy preference vector. Establish a three-dimensional evaluation system encompassing knowledge, skills, and strategy dimensions; employ the analytic hierarchy process (AHP) to determine the initial static weights for each dimension; Based on the initial score vector and strategy preference vector of the knowledge dimension, as well as historical operation records, fuzzy evaluation matrices for the knowledge, skill and strategy dimensions are constructed respectively. The weights of each dimension and the corresponding fuzzy evaluation matrices are fuzzy synthesized to obtain the comprehensive fuzzy evaluation results of each dimension. The comprehensive fuzzy evaluation results are then defuzzified to obtain the quantitative scores of each dimension. The basic feature vector is concatenated with the quantitative scores of each dimension to form a vectorized representation of the student's initial ability profile.

3. The learning path generation method in virtual power plant simulation training according to claim 2, characterized in that, Based on the trainee's initial ability profile, the first training scenario is matched and recommended from a pre-set scenario library, specifically including: Extract the ability score vector from the initial ability profile of the learner. The elements of this vector are the quantitative scores of each dimension. For each candidate scenario in the scenario library, obtain its preset ability requirement vector. The elements of this vector are the preset minimum requirement scores of each dimension. Calculate the cosine similarity between the ability score vector and the ability requirement vector to obtain the content matching score between the learner and the candidate scenario. Retrieve a set of historical students that are similar to the current student's initial ability profile, count the frequency of each candidate scene in the set being selected as the first training scene, divide the selection frequency of the candidate scene by the maximum selection frequency, and obtain the collaborative filtering recommendation score of the candidate scene. The content matching score and the collaborative filtering recommendation score are weighted and summed to obtain the final recommendation score for each candidate scenario; Candidate scenarios with content matching scores not less than a preset first matching threshold are formed into an initial screening set. The candidate scenario with the highest final recommendation score is selected from the initial screening set as the student's first training scenario.

4. The learning path generation method in virtual power plant simulation training according to claim 3, characterized in that, Based on operational process data, the initial competency profile of trainees is dynamically updated to generate a real-time competency profile, which specifically includes: Analyze and extract a set of quantifiable key performance indicators from the operational process data. The indicator set includes time-based, accuracy-based, result-based, and strategy-based indicators; Construct a mapping matrix Matrix elements Indicators The contribution weight to the ability dimension d, where the ability dimension includes the knowledge dimension. Skills dimension With strategy dimension ; Key Performance Indicators Normalization yields Calculate the performance increment vector for each capability dimension using the following formula. : Based on performance increment vector For the initial static weights Make dynamic adjustments to obtain the current dynamic weights. : in, These are preset adjustment coefficients; j = k, s, a, representing the knowledge dimension, skill dimension, and strategy dimension, respectively. Based on the performance increment vector With current dynamic weights Update scores for each dimension: in, The learning rate parameter, The score of the j-th dimension of the initial capability profile; The updated score is concatenated with the basic feature vector to generate a real-time ability profile of the trainee.

5. The learning path generation method in virtual power plant simulation training according to claim 4, characterized in that, The path optimization model is a path optimization model based on reinforcement learning training.

6. The learning path generation method in virtual power plant simulation training according to claim 5, characterized in that, The real-time ability profile of the trainee, the completed training scenario sequence, and the operation process data are input into the path optimization model, which then outputs the next recommended training scenario, specifically including: The current state vector is formed by concatenating the ability scores of each dimension in the real-time ability profile of the trainee, the encoded features of the completed training scene sequence, and the scene performance statistical features extracted from the operation process data. The current state vector is input into a path optimization model trained based on reinforcement learning; the model processes the state vector through its internal state encoder and policy network, and outputs a probability distribution of multiple optional actions, including "recommend the next training scenario"; the optimal action is selected based on the probability distribution. If the selected optimal action is "recommend the next training scene", then the scene generation sub-network in the model will output a recommendation weight vector corresponding to all scenes in the scene library. Calculate the real-time matching degree between the student's real-time ability vector and the preset ability requirement vector for each scenario in the scenario library; based on the preset second matching degree threshold, select scenarios that meet the matching degree to form a subset of candidate recommended scenarios; In the candidate recommendation scenario subset, for each scenario, its corresponding weight in the recommendation weight vector is weighted and fused with its real-time matching degree to obtain the final selection score for that scenario; the scenario with the highest final selection score is selected as the next recommendation training scenario.

7. The learning path generation method in virtual power plant simulation training according to claim 6, characterized in that, Generate and output a learning path report, specifically including: Collect and integrate various types of data generated throughout the training process, including trainee information, initial competency profiles, executed training scenario sequences, operation process data corresponding to each scenario, decision logs of the path optimization model, and evolution records of trainees' real-time competency profiles, to form a structured original dataset of learning paths. Based on the original dataset of the learning path, visual analysis content is generated, including the changing trends of learners' ability scores in the three dimensions of knowledge, skills, and strategies as the training process progresses. The comprehensive mastery scores of learners in different types of training scenarios are displayed in matrix form.

8. A learning path generation system for virtual power plant simulation training, characterized in that, include: The initial competency profile generation module is used to collect the trainees' initial characteristic data, including identity information, historical training records, pre-test scores and decision-making style questionnaire results, and generate the trainees' initial competency profile based on the initial characteristic data. The first training scenario recommendation module is used to match and recommend the first training scenario from a pre-set scenario library based on the student's initial ability profile. The operation process data acquisition module is used to collect operation process data in real time during the training scenario performed by the trainee; The real-time competency profile generation module is used to dynamically update the trainee's initial competency profile based on the operation process data, and generate the trainee's real-time competency profile. The recommended training scenario output module is used to input the trainee's real-time ability profile, the completed training scenario sequence and operation process data into the path optimization model, and the path optimization model outputs the next recommended training scenario. The learning path report output module is used to repeatedly execute the operation process data acquisition module, real-time capability profile generation module, and recommended training scenario output module until the preset training termination conditions are met, and then generate and output a learning path report.

9. A terminal, characterized in that, include: The memory is used to store the learning path generation program in the virtual power plant simulation training; A processor is configured to implement the steps of the learning path generation method in the virtual power plant simulation training as described in any one of claims 1 to 7 when executing the learning path generation program in the virtual power plant simulation training.

10. A computer-readable storage medium, characterized in that, The readable storage medium stores a learning path generation program for virtual power plant simulation training. When the learning path generation program for virtual power plant simulation training is executed by the processor, it implements the steps of the learning path generation method for virtual power plant simulation training as described in any one of claims 1 to 7.