Practical training teaching effect evaluation method based on big data analysis
By constructing a standardized case investigation knowledge graph and graph theory alignment algorithm, the problem of insufficient multi-dimensional quantification in existing practical training assessment methods is solved, enabling refined assessment and process feedback of student behavior, thereby improving teaching quality and guidance.
Patent Information
- Application Number
- CN202511714415.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2025-12-23
AI Technical Summary
Existing practical training assessment methods are unable to provide a refined, multi-dimensional quantitative assessment of trainees' behavioral patterns during practical operations, particularly the compliance of investigation procedures, the logical rationality of evidence chain construction, and the dynamic accuracy of crime type identification. This results in delayed teaching feedback and unclear directions for improvement.
A standardized case investigation knowledge graph is constructed as a benchmark for expert behavior models. The full and high-dimensional behavioral event sequences of trainees in the virtual training environment are collected in real time. The trainees' behavioral trajectory graph is deeply compared with the expert knowledge graph through graph theory alignment algorithm to generate multi-dimensional evaluation indicators, including compliance of investigation process, completeness of evidence chain construction, convergence of crime type identification, and efficiency index of investigation behavior.
It enables objective, in-depth, and process-oriented evaluation of the effectiveness of practical training, provides comprehensive and in-depth quantitative analysis and visual feedback, enhances the guidance and effectiveness of teaching, and establishes a unified evaluation benchmark to support the iterative optimization of teaching methods.
Smart Images

Figure CN121189650A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of data processing methods for prediction purposes, and specifically relates to a practical training teaching effect evaluation method based on big data analysis. BACKGROUND
[0002] With the deep integration of information technology and education, practical training teaching occupies a core position in the cultivation of professionals in public security, justice and network security. Traditional practical training effect evaluation mainly relies on final examination results or teacher subjective evaluation, and its evaluation system is based on static and discrete output results, which is difficult to reflect the dynamic behavior characteristics and cognitive process of students in complex practical scenes.
[0003] Especially in the investigation of network-related crimes, students need to complete the complete operation chain from clue discovery, evidence fixation to legal application. The ability not only reflects whether the final conclusion is correct, but also whether the investigation process meets the legal procedures, the evidence chain is logically rigorous, and the technical identification of crime types is accurate.
[0004] However, the existing evaluation methods lack the ability to capture and quantitatively analyze the above key behavior dimensions in detail, resulting in delayed teaching feedback, unclear improvement direction, and inability to support personalized teaching intervention and curriculum system iteration optimization.
[0005] Big data-driven education evaluation methods have gradually become a research hotspot. This direction aims to collect learners' whole-process interaction data in a digital teaching environment, build a multi-dimensional behavior portrait, and achieve dynamic and objective description of learning effectiveness. In the practical training teaching scene, such methods can theoretically cover operation trajectory, task response timing, tool invocation logic, and collaboration interaction mode, providing technical possibilities for breaking through the single nature of traditional evaluation.
[0006] Although existing technologies have tried to introduce log analysis or simple statistical indicators for teaching monitoring, there are still defects:
[0007] First, data collection is limited to surface behaviors such as system login and correct or incorrect answers, and cannot deeply analyze high-level ability elements such as process compliance, evidence relevance, and type discrimination accuracy in network-related crime operations;
[0008] Second, the evaluation model mainly uses linear weighting or rule threshold judgment, lacks the ability to mine non-linear relationships between multi-dimensional behavior characteristics, and is difficult to reveal the internal structure of skill mastery;
[0009] Third, the analysis results do not form a closed-loop mapping with specific teaching goals, and cannot effectively distinguish whether operation errors are caused by knowledge gaps, process omissions or strategy biases.
[0010] The above problems are particularly prominent in network crime investigation practical training which emphasizes procedural justice and rigorous evidence, and seriously restrict the accurate diagnosis and continuous improvement of teaching quality, and a new evaluation method which can deeply integrate practical training business logic and big data intelligent analysis is urgently needed. SUMMARY
[0011] The purpose of the present application is to provide a practical training teaching effect evaluation method based on big data analysis, aiming to solve the technical problem that the existing practical training teaching evaluation only focuses on the final result performance, and cannot finely and multi-dimensionally quantitatively evaluate the behavior pattern of students in the operation process, especially the compliance of the investigation process, the logical rationality of the evidence chain construction and the dynamic accuracy of the crime type identification.
[0012] In order to achieve the above purpose, the practical training teaching effect evaluation method based on big data analysis provided by the present application is characterized in that a standardized case investigation knowledge graph is constructed as an expert behavior model benchmark, full-dimensional behavior event sequence data of students in a virtual practical training environment is collected in real time, the behavior sequence of the students is mapped into a personal behavior trajectory graph, the graph theory alignment algorithm is used to deeply compare and analyze the behavior trajectory graph of the students and the expert knowledge graph, and finally a group of accurate quantitative evaluation indexes are generated from four dimensions of process compliance, logical reasoning, cognitive accuracy and operation efficiency, so as to realize the objective, in-depth and process evaluation of the teaching effect.
[0013] According to one aspect of the present application, a practical training teaching effect evaluation method based on big data analysis is provided, which comprises the following steps:
[0014] Multi-dimensional behavior event sequence data of students in a network crime virtual practical training environment, i.e. a virtual practical training case, is obtained, the multi-dimensional behavior event sequence data is structured log which records each atomic operation of the students in the practical training process, and each log data contains a timestamp, an operation subject identifier, a session identifier, an operation type code, an operation object and a system state snapshot before and after operation execution;
[0015] A standardized investigation knowledge graph corresponding to the virtual practical training case is constructed, the standardized investigation knowledge graph is a directed acyclic graph which is based on the key path of case investigation and defines four node types including evidence nodes, procedure nodes, hypothesis nodes and conclusion nodes, and three directed edge types including pointing relationship, pre-constraint relationship and support relationship, and all nodes and edges are assigned with preset weight values;
[0016] The multi-dimensional behavior event sequence data of the trainee is processed and mapped into a trainee behavior trajectory graph, which is a time-series directed graph, the nodes of which correspond to specific operation behaviors of the trainee, and the edges represent the sequence and time interval of the operation behaviors;
[0017] The trainee behavior trajectory graph is aligned and semantically associated with the standardized investigation knowledge graph based on a time-series graph kernel function, the mapping path of the trainee behavior trajectory graph on the standardized investigation knowledge graph is calculated, and four deviation types, i.e., positive matching, redundant operation, key omission and sequence error, between the mapping path and the preset optimal investigation path in the knowledge graph are identified;
[0018] Based on the results of the alignment and association analysis, the multi-dimensional evaluation indicators of the trainee's practical training process are quantitatively calculated, including investigation process compliance index, evidence chain construction completeness index, crime type identification convergence index and investigation behavior efficiency index;
[0019] According to the quantitative calculation results, a structured comprehensive practical training teaching effect evaluation report is generated, which contains the numerical values of the multi-dimensional evaluation indicators, the alignment visualization results of the trainee behavior trajectory graph and the standardized investigation knowledge graph, and specific improvement suggestions for the identified deviation types.
[0020] As an embodiment of the present application, the acquisition of multi-dimensional behavior event sequence data of the trainee in the virtual practical training environment specifically includes:
[0021] A behavior data collection agent is deployed in the bottom operating system kernel layer of the virtual practical training platform; the behavior data collection agent captures all the interactive behaviors of the trainee by hooking system calls and monitoring user interface events;
[0022] The interactive behaviors include command line input, file system operation, sending and receiving of network communication data packets, clicking, dragging, hovering of graphical user interface elements, and content change of text input boxes;
[0023] The behavior data collection agent structures and encapsulates the captured raw behavior data according to a predefined log format, which includes fields: global unique timestamp, trainee anonymous identifier, current practical training session unique identifier, predefined digital code of operation behavior, unique resource identifier of operation target, set of system key state variables before operation, and set of system key state variables after operation.
[0024] As an embodiment of the present application, the construction of the standardized investigation knowledge graph specifically includes:
[0025] The expert in the field of network security and criminal investigation defines all key elements required for case investigation and formalizes them into graph nodes through a special knowledge graph construction software interface for each online crime training case;
[0026] The evidence node represents various types of digital evidence in the case, and the attributes include the acquisition method, storage path and hash value of the evidence;
[0027] The program node represents the legal or technical procedures that must be followed during the investigation, and the attributes include the preconditions and execution standards;
[0028] The hypothesis node represents intermediate reasoning based on existing evidence, and the attributes include the evidence set supporting the hypothesis; the conclusion node represents the final qualification of the case;
[0029] The expert also defines the directed edges connecting the nodes and sets the weight, which is represented by a 0-1 floating point number, reflecting the importance of the node or path in the entire investigation process.
[0030] As an embodiment of the present application, the similarity alignment and semantic correlation analysis of the student behavior trajectory graph and the standardized investigation knowledge graph are realized by the following steps:
[0031] First, the node semantic vectorization processing is performed on the student behavior trajectory graph and the standardized investigation knowledge graph, and a pre-trained language model is used to encode each node description text and its associated operation parameters into a fixed-dimensional real vector;
[0032] Secondly, a time-decay weighted graph kernel function is used to calculate the overall structure feature vector of the student behavior trajectory graph and the standardized investigation knowledge graph; the graph kernel function assigns a lower weight to the association between continuous behaviors with a long time interval and a higher weight to behaviors with a short time interval during the calculation process;
[0033] Thirdly, the cosine similarity between the structure feature vector of the student behavior trajectory graph and the structure feature vector of all preset investigation paths in the standardized investigation knowledge graph is calculated in the vector space;
[0034] Finally, the preset investigation path with the highest cosine similarity is selected as the optimal alignment target, and a dynamic programming-based sequence alignment algorithm is used to accurately match each operation node in the student behavior trajectory graph to the corresponding node in the optimal alignment target path, thereby identifying four specific deviations: matching, redundancy, omission and misordering.
[0035] As an embodiment of the present application, the quantitative calculation of the multi-dimensional evaluation index specifically includes:
[0036] The calculation method of the investigation process compliance index is that: all operations in the student behavior trajectory graph that match the procedure nodes in the standardized investigation knowledge graph are counted, and whether the execution sequence meets the defined pre-constraint relationship between the procedure nodes is checked, and the value of the index is equal to the sum of the weights of the matching procedure nodes that meet the pre-constraint relationship divided by the sum of the weights of all procedure nodes in the knowledge graph;
[0037] The calculation method of the evidence chain construction completeness index is that: in the alignment result, the longest continuous and logically correct evidence path constructed by the student is identified, and the value of the index is equal to the sum of the weights of the evidence nodes and the pointing relationship edges contained in the longest continuous and logically correct evidence path, divided by the total weight of the optimal evidence path in the standardized investigation knowledge graph;
[0038] The calculation method of the crime type identification convergence index is that: at the beginning of the training, a uniform probability distribution for all candidate crime types is set for the student; during the training process, whenever the student obtains a strongly related evidence node pointing to a specific crime type, the Bayesian update rule is used to adjust the probability judgment of each crime type; the index is calculated by calculating the Kullback-Leibler divergence between the final probability distribution of the student at the end of the training and the one-hot encoding distribution corresponding to the real crime type of the case, and the smaller the divergence value is, the higher the index value is;
[0039] The calculation method of the investigation behavior efficiency index is that: the value of the index is equal to the total number of nodes contained in the optimal investigation path in the standardized investigation knowledge graph, divided by the total number of nodes contained in the student behavior trajectory graph, and the ratio reflects the simplicity and efficiency of the student's operation.
[0040] As an embodiment of the present application, the generating a structured comprehensive training teaching effect evaluation report specifically includes:
[0041] The first page of the report displays the scores of the student in the four dimensions of investigation process compliance, evidence chain construction completeness, crime type identification convergence and investigation behavior efficiency in the form of a radar chart;
[0042] The main part of the report presents in a visual way, the behavior trajectory graph of the student is drawn in one color or line type, and is superimposed on the standardized investigation knowledge graph drawn in another color or line type; in the superimposed graph, the redundant operation nodes in the student behavior path, the missed key nodes and the path segments with incorrect execution sequence are clearly marked;
[0043] At the end of the report, a diagnostic text description is attached, which lists the student behavior deviations identified by the system, and provides specific and actionable improvement suggestions for each deviation, explains why the operation is redundant or wrong, and the potential consequences of missing a certain step, according to the content of the standardized investigation knowledge graph.
[0044] Compared with the prior art, the present application has the beneficial effects that:
[0045] 1. The multi-dimension and refinement of the evaluation dimension are realized, the four core indexes of investigation process compliance, evidence chain construction completeness, crime type recognition convergence and investigation behavior efficiency are constructed, the traditional single result score is decomposed into full and deep quantitative analysis of the student investigation thinking process, and the evaluation granularity is deepened from macro whether correct to micro how to think and operate.
[0046] 2. The objectivity and data-driven nature of the evaluation process are ensured, the present application is completely based on the objective behavior data flow generated by the student in the virtual environment, and the deviation caused by the subjective score in the traditional evaluation is excluded through the algorithm-driven alignment and comparison with the expert knowledge graph, so that the evaluation result has high reproducibility, consistency and public credibility.
[0047] 3. Process diagnosis and feedback are provided, the present application not only gives a final evaluation score, but also aligns the visualization graph and specific deviation analysis, directly reveals each logical breakpoint, knowledge blind area or operation error of the student in the investigation process, and generates targeted improvement suggestions, changes the evaluation from one-time terminal evaluation to a continuous improvement formative teaching link, greatly improves the guidance and effectiveness of the practical teaching.
[0048] 4. A standardized evaluation benchmark is established, the standardized investigation knowledge graph is introduced as an expert model, a unified measurement scale is provided for the practical training effect of different students and different batches, horizontal and vertical comparison of teaching effect in large scale and across time and space becomes a reality, and solid data support is provided for iterative optimization of teaching methods. BRIEF DESCRIPTION OF DRAWINGS
[0049] Figure 1 It is the overall technical scheme architecture schematic diagram of the practical teaching effect evaluation method based on big data analysis proposed by the present application;
[0050] Figure 2 It is the core principle framework schematic diagram of the alignment analysis of the student behavior trajectory graph and the standardized investigation knowledge graph based on the time sequence graph kernel function in the present application;
[0051] Figure 3 It is the logical flow framework diagram of the multi-dimensional behavior event sequence data collection and structured processing in the present application;
[0052] Figure 4 is the logical flow framework of the standardized investigation knowledge graph construction and its node-edge semantic formal definition in the application;
[0053] Figure 5 is the multi-level interaction relationship and data flow schematic diagram of the multi-dimensional evaluation index quantification and structured report generation of the student practical training process in the application; DETAILED DESCRIPTION
[0054] Please refer to Figures 1 to 5 The application provides a practical training teaching effect evaluation method based on big data analysis, which is characterized in that a standardized investigation knowledge graph is constructed as an expert behavior model benchmark, full-dimensional behavior event sequences of students in a network crime virtual training environment are collected in real time, the sequences are mapped into student behavior trajectory graphs, and then a graph theory alignment algorithm is used to realize deep comparison and deviation analysis with the expert knowledge graph, and finally accurate quantitative evaluation indexes are generated from four dimensions of investigation process compliance, evidence chain construction rationality, crime type identification accuracy and operation efficiency, so that process, multi-dimensional and objective evaluation of the practical training teaching effect is realized.
[0055] The method comprises the following steps:
[0056] S1, obtaining multi-dimensional behavior event sequence data of students in a network crime virtual training environment, i.e. a virtual training case;
[0057] S2, constructing a standardized investigation knowledge graph corresponding to the virtual training case;
[0058] S3, processing and mapping the multi-dimensional behavior event sequence data of the students into a student behavior trajectory graph;
[0059] S4, performing similarity alignment and semantic association analysis of the student behavior trajectory graph and the standardized investigation knowledge graph based on a time series graph kernel function;
[0060] S5, based on the results of the alignment and association analysis, quantitatively calculating multi-dimensional evaluation indexes of the student's practical training process;
[0061] S6, generating a structured comprehensive practical training teaching effect evaluation report according to the quantitative calculation results.
[0062] In step S1, multi-dimensional behavior event sequence data of students in a network crime virtual training environment is obtained. The multi-dimensional behavior event sequence data is structured log, and each log records an atomic operation performed by the student in the training process.
[0063] Each log contains seven core fields: a globally unique timestamp, a student anonymous identifier, a current hands-on session unique identifier, a predefined numerical encoding of the operation behavior, a unique resource identifier of the operation target, a set of system key state variables before operation, and a set of system key state variables after operation.
[0064] To ensure the integrity and non-tamperability of data collection, a behavior data collection agent is deployed in the underlying operating system kernel layer of the virtual hands-on platform.
[0065] The agent captures all the interactive behaviors of students through hooking system call interfaces and monitoring graphical user interface events.
[0066] The captured interactive behaviors include command line input instructions, file system read-write operations, network communication data packet transmission and reception behaviors, graphical user interface element clicks, drags, hover actions, and text input box content change events.
[0067] All raw behavior data is immediately structured and encapsulated according to the predefined log format after being captured, and written to a distributed log storage system.
[0068] The timestamp of the log has a precision greater than milliseconds to ensure the accuracy of subsequent time series analysis.
[0069] The predefined numerical encoding of the operation behavior adopts a unified coding specification, such as encoding 001 representing accessing a specified IP address, encoding 002 representing extracting a hard disk image, and encoding 003 representing parsing DNS logs, etc. All encodings are preloaded into the behavior mapping dictionary by the system before the hands-on begins.
[0070] The unique resource identifier of the operation target is used to uniquely identify the object being operated, such as a file path hash value, a network endpoint quadruple, or a database table name. The set of system key state variables includes but is not limited to current memory usage, CPU load, network connection number, extracted evidence list, current hypothesis set, etc. These variables are snapshot saved before and after each operation to restore the operation context.
[0071] In step S2, a standardized investigation knowledge graph corresponding to the virtual hands-on case is constructed. The knowledge graph is a directed acyclic graph, and its construction process is led by experts in the fields of network security and criminal investigation.
[0072] For each online crime hands-on case, the expert defines all the key elements required for case cracking through a dedicated knowledge graph construction software interface, and formalizes them as graph nodes.
[0073] The nodes are divided into four types: evidence nodes, procedure nodes, hypothesis nodes, and conclusion nodes.
[0074] An evidence node represents a type of digital evidence in the case, whose attributes include the acquisition method, storage path, hash value, time range, and associated original data source.
[0075] A procedure node represents a legal or technical procedure that must be followed during the investigation process, such as the need to apply for an electronic data subpoena before accessing a cloud server. Its attributes include preconditions, execution standards, and compliance verification rules.
[0076] A hypothesis node represents an intermediate inference based on existing evidence, such as the possibility that the attacker used a jump machine. Its attributes include the set of evidence supporting the hypothesis, the initial confidence value, and the falsifiable conditions.
[0077] A conclusion node represents the final characterization of the case, such as constituting the crime of illegally obtaining computer information system data. Its attributes include legal citation, sentencing recommendation basis, and reasons for excluding other charges.
[0078] The expert defines three types of directed edges connecting the nodes: a pointing relationship, indicating that evidence directly supports a hypothesis or conclusion; a pre-constraint relationship, indicating that a procedure node must be executed before another node; and a support relationship, indicating that multiple pieces of evidence collectively support a hypothesis.
[0079] All nodes and edges are assigned a floating-point weight between 0 and 1, which is assigned by the expert based on their importance in the entire investigation logic chain. For example, the weight of a key evidence node is 0.9, and the weight of a supplementary procedure node is 0.3. The constructed knowledge graph is stored in a graph database format, and a unique graph version identifier is generated for each training case to ensure consistency and traceability of the evaluation benchmark.
[0080] In step S3, the multidimensional behavior event sequence data of the student is processed and mapped into a student behavior trajectory graph. The trajectory graph is a time-series directed graph, and its construction process first cleans and denoises the original logs collected in S1, removing invalid operations and system-generated logs.
[0081] Subsequently, each valid log is mapped to a node in the graph, and the type of the node is automatically determined based on the operation behavior code, such as code 002 for evidence extraction nodes and code 005 for procedure execution nodes. The attributes of the node include operation time, operation object identifier, and difference summary of pre-operation and post-operation state snapshots.
[0082] Then, according to the time stamp order of the log, a directed edge is established between adjacent operation nodes, and the attribute of the edge contains the time interval and the operation continuity flag. The final student behavior trajectory graph completely retains the operation sequence, time distribution and state transition path of the student during the training process, and its topology reflects the student's investigation thinking process and behavior habits.
[0083] In step S4, the student behavior trajectory graph and the standardized investigation knowledge graph are aligned and analyzed based on the time sequence graph kernel function.
[0084] The analysis process first performs semantic vectorization processing on all nodes in the two graphs.
[0085] A pre-trained language model is used to input the description text of each node and its associated operation parameters after splicing, and output a fixed-dimensional real vector as the semantic embedding of the node.
[0086] Secondly, a time-decay-weighted graph kernel function is used to calculate the overall structural feature vector of the student behavior trajectory graph and the standardized investigation knowledge graph. The graph kernel function is defined as follows:
[0087] ;
[0088] where, is the student behavior trajectory graph, is the standardized investigation knowledge graph, and are the node sets thereof, and are the time stamps of the nodes, is the time decay constant, is the semantic embedding vector of the node , and denotes the inner product of vectors.
[0089] The kernel function gives a lower weight to behaviors with a longer time interval when calculating the structural similarity, reflecting the characteristic that recent operations have a greater impact on current decisions in human cognition.
[0090] Thirdly, the cosine similarity between the structural feature vector of the student behavior trajectory graph and the structural feature vectors of all preset investigation paths in the standardized investigation knowledge graph is calculated in the vector space.
[0091] Finally, the preset investigation path with the highest cosine similarity is selected as the optimal alignment target, and a sequence alignment algorithm based on dynamic programming is used to accurately match each operation node in the student behavior trajectory graph to the corresponding node in the optimal alignment target path.
[0092] The alignment algorithm defines a matching cost matrix, where the matching cost is jointly determined by the semantic distance of nodes and the time offset. By backtracking the minimum cost path, four types of deviation are identified: positive matching: the student's operation is successfully aligned with the knowledge graph node, redundant operation: the student performs an operation that does not exist in the knowledge graph, key omission: a key node exists in the knowledge graph but is not executed by the student, and sequence error: the student executes the correct node but violates the precedence constraint relationship.
[0093] In step S5, based on the results of the alignment and association analysis, the practical training process of the student is quantitatively calculated based on multi-dimensional evaluation indicators. The multi-dimensional evaluation indicators include four core indexes.
[0094] The calculation method of the investigation process compliance index is to traverse all operation nodes matched with the program nodes in the standardized investigation knowledge graph in the alignment result, and check whether the execution sequence satisfies the precedence constraint relationship defined between the program nodes.
[0095] If the precedence constraint of a program node A requires node B to be executed first, and A appears before B in the student's behavior trajectory, the alignment is considered invalid.
[0096] The value of this index is equal to the sum of the weights of all valid matching program nodes that satisfy the precedence constraint relationship, divided by the total weight of all program nodes in the knowledge graph. Its mathematical expression is:
[0097] ;
[0098] where, is the set of valid matching program nodes, is the set of all program nodes in the knowledge graph, is the weight of node .
[0099] The calculation method of the evidence chain construction completeness index is to identify the longest continuous and logically correct evidence path constructed by the student in the alignment result.
[0100] This path must satisfy: the starting point is the initial evidence, the ending point is the evidence supporting a hypothesis or conclusion, all edges on the path are pointing or supporting relationships, and there is no logical break.
[0101] The value of this index is equal to the sum of the weights of all evidence nodes and connecting edges contained in the longest continuous and logically correct evidence path, divided by the total weight of the optimal evidence path in the standardized investigation knowledge graph, i.e. the shortest complete path preset by experts. Its calculation formula is:
[0102] ;
[0103] where, and are the node and edge sets in the longest evidence path of the student, respectively, is the total weight of the optimal evidence path in the knowledge graph.
[0104] The calculation method of the crime type recognition convergence index is: at the beginning of the training, a uniform probability distribution is set for the student for all candidate crime types, assuming that there are possible charges, then the initial probability vector is During the training process, whenever the student obtains a strongly related evidence node pointing to a specific crime type, the Bayesian update rule is used to adjust its probability judgment for each crime type. The update formula is:
[0105] ;
[0106] wherein, is the th crime type, is the th obtained evidence, is the likelihood value of the evidence under the charge , which is determined by the pre-set evidence-charge association strength in the knowledge graph. At the end of the training, the final probability distribution is obtained. The index is measured by calculating the Kullback-Leibler divergence between and the one-hot encoding distribution corresponding to the true crime type of the case , the smaller the divergence value, the higher the index value. The final index is defined as:
[0107]
[0108] wherein, is the Kullback-Leibler divergence, and the denominator is used for normalization to ensure that the index value falls within the range of 0-1.
[0109] The calculation method of the investigation behavior efficiency index is: the value of the index is equal to the total number of nodes contained in the optimal investigation path in the standardized investigation knowledge graph , divided by the total number of nodes contained in the student's behavior trajectory graph . This ratio reflects the simplicity and efficiency of the student's operation, avoiding unnecessary exploratory operations. Its expression is:
[0110] ;
[0111] In step S6, according to the quantitative calculation results, a structured comprehensive training teaching effect evaluation report is generated.
[0112] The first page of the report displays the student's scores in the above four dimensions in the form of a radar chart, with each axis corresponding to the index value. The main part of the report uses visualization techniques to plot the student's behavior trajectory graph in blue solid lines and superimposes it on the standardized investigation knowledge graph plotted in red dashed lines.
[0113] In the superimposed graph, redundant operation nodes are marked with yellow triangles, missing key nodes are marked with red hollow circles, and path segments with incorrect execution order are highlighted with flashing animations or thick red lines.
[0114] At the end of the report, there is a diagnostic text description that lists the student's behavior deviations identified by the system.
[0115] For example: The student directly assumes the attack source IP without obtaining network traffic logs, violating the principle of evidence first, and is advised to perform the capture and analyze PCAP file operation first; The student repeatedly attempts to crack the same encrypted file three times, which is a redundant operation, and is advised to switch to other evidence clues after the first failure; The student missed the key procedure node of verifying timestamp consistency, which may cause the evidence chain to break, and is advised to perform time alignment verification immediately after extracting multiple source evidence.
[0116] Each suggestion references the corresponding node definition and weight in the knowledge graph, ensuring the professionalism and operability of the feedback.
[0117] The method described in this embodiment achieves a full-process, multi-dimensional, data-driven precise evaluation of the effect of network-related crime training through the above six steps.
[0118] The entire process is based entirely on objective behavior data, relies on a standardized knowledge graph as an evaluation benchmark, and through graph alignment and deviation analysis techniques, abstract investigation capabilities are converted into quantifiable numerical indicators, supplemented by visualization and diagnostic feedback, improving the scientificity, guidance and teaching value of the training evaluation.
[0119] The execution of the method relies on a complete software system architecture, which includes a behavior data collection module, a knowledge graph management module, a trajectory graph construction module, a graph alignment analysis engine, an evaluation index calculation unit, and a report generator.
[0120] The behavior data collection module runs in the form of a kernel-level agent to ensure comprehensive data collection and low latency.
[0121] The knowledge graph management module provides a graphical editing interface for experts to build and maintain the graph, and supports version control and permission management.
[0122] The trajectory graph construction module consumes log streams in real time and dynamically generates student behavior trajectory graphs.
[0123] The graph alignment analysis engine integrates a pre-trained language model and a graph kernel function calculation library, and is responsible for performing complex semantic alignment tasks.
[0124] The evaluation index calculation unit calls a preset formula to perform numerical calculation according to the alignment result.
[0125] The report generator integrates all analysis results and outputs a structured report meeting the teaching needs. The modules are loosely coupled through message queues and API interfaces to ensure the scalability and stability of the system.
[0126] In summary, the embodiment discloses a practical training teaching effect evaluation method based on big data analysis, which solves the technical defects of single dimension, strong subjectivity and lack of process feedback of traditional evaluation methods by constructing an expert knowledge graph, collecting full behavior data, performing graph alignment analysis, quantifying multi-dimensional indexes and generating a diagnosis report, and provides a scientific, objective and efficient evaluation tool for network security practical training teaching.
[0127] It should be noted that, in this paper, relational terms such as first and second are used only to distinguish one entity or operation from another, and do not necessarily require or imply that these entities or operations exist in any such actual relationship or order. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed or inherent to such a process, method, article or device.
[0128] Although embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for evaluating the effectiveness of practical training based on big data analysis, characterized in that, include: Acquire multi-dimensional behavioral event sequence data of trainees in a virtual training environment for cybercrime, i.e., virtual training cases; A standardized investigative knowledge graph corresponding to the virtual training case is constructed. The standardized investigative knowledge graph is a directed acyclic graph. Based on the key path of case investigation, it defines four types of nodes, including evidence nodes, procedure nodes, hypothesis nodes and conclusion nodes, as well as three types of directed edges, including pointing relationship, precondition relationship and support relationship. Preset weight values are assigned to all nodes and edges. The multi-dimensional behavioral event sequence data of the trainees is processed and mapped into a trainee behavior trajectory graph. The trainee behavior trajectory graph is a temporal directed graph, in which the nodes correspond to the specific operational behaviors of the trainees, and the edges represent the order and time interval of the operational behaviors. The similarity alignment and semantic association analysis based on the temporal graph kernel function are performed on the trainee behavior trajectory map and the standardized reconnaissance knowledge graph to calculate the mapping path of the trainee behavior trajectory map on the standardized reconnaissance knowledge graph, and four types of deviations between the mapping path and the preset optimal reconnaissance path in the knowledge graph are identified: positive matching, redundant operation, key omission and sequence error. Based on the results of the alignment and correlation analysis, the training process of trainees is quantitatively calculated using multi-dimensional evaluation indicators, including the compliance index of the investigation process, the completeness index of the evidence chain construction, the convergence index of crime type identification, and the efficiency index of investigation behavior. Based on the quantitative calculation results, a structured comprehensive practical training teaching effectiveness evaluation report is generated. The report includes the values of evaluation indicators for each dimension, the alignment visualization results of student behavior trajectory maps and standardized investigation knowledge graphs, and specific improvement suggestions for the identified deviation types.
2. The method for evaluating the effectiveness of practical training based on big data analysis according to claim 1, characterized in that, The multi-dimensional behavioral event sequence data is a structured log, which records every atomic operation performed by the trainee during the training process. Each log entry includes a timestamp, operation subject identifier, session identifier, operation type code, operation object, and a snapshot of the system state before and after the operation.
3. The method for evaluating the effectiveness of practical training based on big data analysis according to claim 1, characterized in that, Obtain multi-dimensional behavioral event sequence data of trainees in a virtual training environment for cybercrime, including: Deploy a behavioral data acquisition agent at the underlying operating system kernel layer of the virtual training platform; The behavior data acquisition agent captures all student interaction behaviors by hooking into system calls and monitoring user interface events; The interactive behaviors include command line input, file system operations, sending and receiving network communication data packets, clicking, dragging, and hovering graphical user interface elements, and changing the content of text input boxes; The behavior data acquisition agent encapsulates the captured raw behavior data in a structured manner according to a predefined log format. The log format includes the following fields: a globally unique timestamp, a student anonymous identifier, a unique identifier for the current training session, a predefined numerical code for the operation behavior, a unique resource identifier for the operation target, a set of key system state variables before the operation, and a set of key system state variables after the operation.
4. The method for evaluating the effectiveness of practical training based on big data analysis according to claim 1, characterized in that, Constructing a standardized investigative knowledge graph corresponding to the virtual training case, including: Experts in cybersecurity and criminal investigation will use a dedicated knowledge graph to build a software interface for each cybercrime training case, manually defining all the key elements required for case investigation and formalizing them as graph nodes. The evidence nodes represent various types of digital evidence in the case, and their attributes include the method of obtaining the evidence, the storage path, and the hash value. The program node represents the legal or technical procedures that must be followed during the investigation process, and its attributes include prerequisites and execution standards. The hypothesis node represents an intermediate inference based on existing evidence, and its attributes include a set of evidence supporting the hypothesis. The conclusion node represents the final determination of the case; Experts also defined directed edges connecting each node and assigned weights, which were represented by floating-point numbers and reflected the importance of the node or path throughout the reconnaissance process.
5. The method for evaluating the effectiveness of practical training based on big data analysis according to claim 1, characterized in that, The similarity alignment and semantic association analysis between the trainee behavior trajectory map and the standardized investigation knowledge graph include: The trainee behavior trajectory map and standardized investigation knowledge graph are processed by node semantic vectorization. A pre-trained language model is used to encode the description text of each node and its associated operation parameters into a fixed-dimensional real number vector. A time-decay weighted graph kernel function is used to calculate the overall structural feature vectors of the trainee behavior trajectory map and the standardized reconnaissance knowledge graph, respectively. During the calculation process, the graph kernel function assigns lower weights to the correlation between continuous behaviors with long time intervals and higher weights to behaviors with short time intervals. Calculate the cosine similarity between the structural feature vector of the trainee's behavior trajectory graph and the structural feature vector of all preset reconnaissance paths in the standardized reconnaissance knowledge graph in the vector space; The preset detection path with the highest cosine similarity is selected as the optimal alignment target. A sequence alignment algorithm based on dynamic programming is used to accurately match each operation node in the student behavior trajectory map to the corresponding node in the optimal alignment target path, thereby identifying four specific deviations: matching, redundancy, omission, and misorder.
6. The method for evaluating the effectiveness of practical training based on big data analysis according to claim 5, characterized in that, The trainee behavior trajectory map and standardized investigative knowledge graph are processed by node semantic vectorization, including: The description text of each node is concatenated with its associated operation parameters and then input into a pre-trained language model, which outputs a real number vector as the semantic embedding of that node.
7. The method for evaluating the effectiveness of practical training based on big data analysis according to claim 5, characterized in that, The overall structural feature vector is calculated using a time-decay-weighted graph kernel function, including: According to the formula: ; Calculate the kernel value of the graph, where To create a behavioral trajectory map of trainees. To standardize the investigative knowledge graph, and Each represents its set of nodes. and For the timestamp of the node, The time decay constant, For nodes semantic embedding vector, This represents the vector dot product.
8. The method for evaluating the effectiveness of practical training based on big data analysis according to claim 1, characterized in that, The quantitative calculation of multi-dimensional evaluation indicators includes: The calculation method for the compliance index of the investigation process is as follows: count all operations in the staff behavior trajectory graph that match the program nodes in the standardized investigation knowledge graph, and check whether their execution order satisfies the pre-constraint relationship defined between the program nodes. The value of the index is equal to the sum of the weights of the matching program nodes that satisfy the pre-constraint relationship divided by the sum of the weights of all program nodes in the knowledge graph. The method for calculating the completeness index of the evidence chain construction is as follows: In the alignment results, the longest continuous and logically correct evidence path constructed by the trainee is identified. The value of this index is equal to the sum of the weights of the evidence nodes and the pointing relation edges contained in the longest continuous and logically correct evidence path, divided by the total weight of the optimal evidence path in the standardized investigation knowledge graph. The convergence index for crime type identification is calculated as follows: At the beginning of the training, a uniform probability distribution for all candidate crime types is set for the trainees; during the training, whenever a trainee obtains a strongly correlated evidence node pointing to a specific crime type, a Bayesian update rule is used to adjust their probability judgment for each crime type; the index is measured by calculating the Körbek-Leibler divergence between the trainee's final probability distribution at the end of the training and the one-hot coding distribution corresponding to the actual crime type of the case. The smaller the divergence value, the higher the index value. The reconnaissance behavior efficiency index is calculated as follows: the value of the index is equal to the total number of nodes contained in the optimal reconnaissance path in the standardized reconnaissance knowledge graph, divided by the total number of nodes contained in the trainee behavior trajectory graph.
9. The method for evaluating the effectiveness of practical training based on big data analysis according to claim 8, characterized in that, The calculation of the crime type identification convergence index includes: Set the initial probability vector ,in Number of alternative crime types; Whenever a student obtains a node indicating a charge, according to the formula: ; Update the probability distribution, where For evidence in the charge The likelihood value below; Calculate the final probability distribution Unique hot coding of the real crime The Körbeck-Leibler divergence between them, and expressed by the formula: ; The normalized exponent value is obtained.
10. The method for evaluating the effectiveness of practical training based on big data analysis according to claim 1, characterized in that, Generate a structured, comprehensive evaluation report on the effectiveness of practical training, including: The report's homepage displays trainees' scores in four dimensions—compliance of investigation procedures, completeness of evidence chain construction, convergence of crime type identification, and efficiency of investigation actions—in the form of a radar chart. The main body of the report draws the trainees' behavioral trajectory map in the first color or line type, and overlays it on the standardized investigation knowledge graph drawn in the second color or line type. Redundant operation nodes, missing key nodes, and path segments with incorrect execution order are marked in the overlay map. The report concludes with a diagnostic text description, listing each behavioral deviation identified by the system and providing specific, actionable improvement suggestions for each deviation based on the content of the standardized investigative knowledge graph.
Citation Information
Patent Citations
Teaching quality evaluation and analysis method based on knowledge graph
CN118195415A
Public security AI-driven data generation type investigation teaching training system
CN120356378A
Practical training teaching and evaluation management system
CN120410809A
Similarity determination device and abnormality detection device
JP2020107016A
System and method for providing personalized explainable response by generating multimedia prompt using contextual information
US20250272323A1
Cited By
Big data full-process practical training and assessment method and system based on multi-resource collaboration
CN121120333A
College culture and vocational education integrated intelligent teaching system
CN121544439A