MR police practical training method and system based on four-layer behavior model
By combining a four-layer behavior model and a multi-window temporal matching algorithm, the problems of incomplete behavior modeling, feedback delay, and insufficient scene fusion in existing police combat training systems are solved, and high-fidelity combat simulation training and quantitative evaluation are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG POLICE COLLEGE (GUANGDONG PROVINCIAL PUBLIC SECURITY JUDICIAL MANAGEMENT CADRE COLLEGE)
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-31
AI Technical Summary
Existing police combat training systems lack a complete logical closed loop in behavioral modeling, have time delays in real-time feedback mechanisms, vague evaluation criteria, and insufficient scenario integration, making it difficult to achieve high-fidelity combat simulation training.
A four-layer behavioral model-based MR police combat training method is adopted, including structured modeling of the field control layer, control layer, tactical layer and closed-loop layer. It combines multi-window temporal matching algorithm for real-time behavioral analysis and provides instant feedback and evaluation through mixed reality devices to achieve dual-field fusion of virtual and real-world scenarios.
It achieves standardized and model-based expression of law enforcement behavior, supports real-time intelligent guidance and quantitative evaluation, enhances the immersiveness of training and the effect of practical simulation, and improves the quantifiability and verifiability of training results.
Smart Images

Figure REF-OBJ-1775568039550-000074 
Figure REF-OBJ-1775568039550-000075 
Figure REF-OBJ-1775568039550-000076
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of police training and mixed reality technology, and in particular to a method and system for MR police combat training based on a four-layer behavioral model. Background Technology
[0002] Practical police training is a core component in enhancing law enforcement officers' on-site response capabilities. Currently, practical police training primarily employs the traditional lecture-demonstration-practice model, which involves classroom lectures, instructor demonstrations, and repetitive, mechanical drills by trainees to impart skills. This model heavily relies on on-site instructor guidance, and the construction of training scenarios depends mainly on venue setup, role-playing, and instructor verbal descriptions, making it difficult for trainees to gain an immersive experience close to real-world scenarios. Furthermore, traditional training methods struggle to effectively simulate complex and ever-changing real-world scenarios, resulting in limited scenario fidelity and difficulty in effectively eliciting psychological stress responses from trainees, leading to a significant gap between training effectiveness and actual combat needs.
[0003] In recent years, with the development of virtual simulation technology, some training systems have begun to incorporate virtual reality or augmented reality technologies to improve the training experience. For example, Chinese invention patent CN116030684A discloses an interactive training system based on virtual reality technology. This system uses simulated dummies in conjunction with sensing devices and AR glasses for wartime first aid training. It utilizes an interactive service platform to simulate scenarios and collect and analyze data for multiple first aid projects, thereby improving the immersion and data-driven level of first aid training to some extent. However, this system is mainly geared towards first aid training scenarios, and its behavioral assessment logic is relatively simple, lacking the ability to model multi-level behaviors for complex law enforcement procedures. Another example is Chinese invention patent CN107111894B, which discloses an augmented or virtual reality simulator for professional and educational training. This system uses VR / AR headsets to create virtual training environments and supports the construction of multi-person collaborative training scenarios. However, its technical solution focuses on the presentation of virtual scenarios and basic interactions, without involving real-time matching and deviation detection mechanisms based on structured behavioral models. Furthermore, Chinese invention patent CN107479699A discloses a virtual reality interaction method, device, and system. This system uses a motion capture camera to collect user location information and sensors to determine user behavior, rendering user actions in a virtual scene to enhance the naturalness and vividness of the interaction. While this system provides valuable technical reference for behavioral data acquisition and rendering, its core focus is on the motion capture and virtual reconstruction of single behaviors, rather than the construction and real-time evaluation of logical chains of multiple behavioral nodes in law enforcement procedures.
[0004] However, the aforementioned existing technologies still have the following shortcomings when applied to police combat training: First, in terms of behavioral modeling, existing systems mostly break down law enforcement procedures into several isolated action units for separate training, lacking the technical means to integrate fragmented actions into a complete behavioral chain. This makes it impossible to establish a complete logical closed loop of situational awareness, decision-making, action execution, and effect evaluation, making it difficult for trainees to quickly recall and integrate learned skills when facing real police situations. Second, in terms of real-time feedback, the feedback mechanism of existing systems has a significant time delay. Instructors need to conduct manual analysis and provide verbal guidance after observing the trainees' actions, with feedback cycles typically measured in minutes or even hours. This makes it impossible to conduct real-time behavioral analysis and deviation detection while the trainees are performing actions, hindering immediate correction during training. Third, in terms of training evaluation, existing systems mainly rely on instructors' subjective observation and experience judgment, lacking a systematic, multi-dimensional evaluation index system. Evaluation standards are vague and inconsistent, resulting in low data utilization efficiency. Fourth, in terms of scene integration, existing VR systems place trainees entirely in a virtual environment, which is not sufficiently integrated with the real physical space. Trainees cannot obtain immersive feedback from the virtual scene while operating simulated police equipment, thus limiting the transfer of training experience to actual combat. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a MR (Mixed Reality) police combat training method based on a four-layer behavior model, comprising the following steps: Scene loading and behavior model initialization: A preset training scene is loaded into the mixed reality interactive environment, and the parameters of the four-layer behavior model corresponding to the training scene are initialized, including a control layer, a tactical layer, and a closed-loop layer. The control layer manages the behavior nodes for the initial control posture at the law enforcement scene; the tactical layer manages the behavior nodes for the use of force and coercive measures; the tactical layer manages the behavior nodes for police team coordination and tactical decision-making; and the closed-loop layer manages the behavior nodes for legal notification and evidence preservation. Real-time user behavior data collection: Spatial location data, body posture data, voice data, and interactive action data of trainees are collected in real time through a mixed reality interactive device. Multi-dimensional behavior feature vector construction and node matching: The behavior data is converted into a multi-dimensional behavior feature vector containing spatial features, action features, and voice features, and multi-window time... The sequence matching algorithm integrates feature vectors with preset behavior nodes for real-time matching; behavior deviation detection and classification compares the matching results with standard behavior paths to detect three types of deviations: node missing deviation, execution order deviation, and timeliness exceeding the limit deviation; intelligent generation and real-time presentation of prompts determine the prompt level based on the deviation type and current behavior state, and generate real-time prompts from a hierarchical prompt library, presenting them to trainees in real-time through the voice broadcast and visual overlay channels of the mixed reality device; behavior execution chain time sequence recording records the complete behavior execution chain with a unified timestamp, covering behavior data, status identifiers, deviation events, and prompt records; six-dimensional weighted evaluation calculation calculates evaluation scores and generates an overall evaluation result from six dimensions: behavior chain completeness, program expression standardization, tactical decision rationality, coordination effectiveness, timeliness control accuracy, and overall execution smoothness; evaluation result output and training feedback output an evaluation report containing sub-scores for each dimension and improvement suggestions.
[0006] This invention also provides a Mixed Reality (MR) police combat training system based on a four-layer behavior model, including an MR interaction module, a behavior recognition module, a logic judgment module, a prompt generation module, and an evaluation module. The MR interaction module is configured to load a virtual training scene through a mixed reality head-mounted display and collect behavioral data from trainees, while also serving as an output terminal for real-time prompt information. The behavior recognition module is configured to convert behavioral data into multi-dimensional behavioral feature vectors and perform real-time matching of behavioral nodes using a multi-window temporal matching algorithm. The logic judgment module is configured to store the parameters of the four-layer behavior model and standard behavioral paths, and perform behavior deviation detection and classification based on the matching results. The prompt generation module is configured to determine the prompt level based on the deviation detection results, retrieve matching prompt templates from a hierarchical prompt library, and generate real-time prompt information. The evaluation module is configured to record the behavior execution chain and calculate the evaluation score using a six-dimensional weighted evaluation algorithm.
[0007] Compared with existing technologies, this invention has the following advantages: First, by constructing a four-layer behavioral model consisting of a control layer, a tactical layer, and a closed-loop layer, this invention, for the first time, decomposes the complex police law enforcement process into a structured and quantifiable hierarchical system of behavioral nodes, achieving a standardized and model-based expression of law enforcement behavior and overcoming the technical deficiency of low behavioral modeling in existing technologies. Second, this invention employs a multi-window temporal matching algorithm to achieve real-time matching of behavioral nodes and detection of three types of deviations. During training, targeted prompts can be presented to trainees through mixed reality devices, achieving real-time intelligent guidance in the training process. Third, this invention uses a six-dimensional weighted evaluation algorithm to construct a quantitative evaluation system covering six dimensions: behavioral integrity, procedural standardization, tactical rationality, collaborative effectiveness, timeliness and accuracy, and execution fluency. The evaluation indicators for each dimension are calculable, verifiable, and comparable. Fourth, this invention, through a dual-domain fusion mode of virtual and real-world domains, simultaneously provides visual immersion and physical tactile feedback in a mixed reality environment, constructing a high-fidelity practical simulation environment. Attached Figure Description
[0008] Figure 1 This is a structural diagram of a four-layer behavioral model provided in an embodiment of the present invention.
[0009] Figure 2 This is a system architecture diagram provided in an embodiment of the present invention.
[0010] Figure 3 This is a flowchart of the method provided in an embodiment of the present invention.
[0011] Figure 4 This is a schematic diagram of the dual-field training mode provided in an embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0013] This invention provides a method for MR police combat training based on a four-layer behavioral model, such as... Figure 3 As shown, the method includes steps S1 to S8. The following detailed description of each step is provided in conjunction with specific embodiments.
[0014] Step S1: Scene Loading and Behavior Model Initialization. During the training startup phase, the system first performs scene loading and behavior model initialization. Specifically, the system selects the target training scene module corresponding to the current training task from a preset scene module library. In one embodiment of the invention, the preset scene module library covers four main categories of core training scenes: inspection module, arrest module, emergency response module, and rescue module. Each category of scene module is further subdivided into several specific police incident sub-scenes. Preferably, the inspection module includes sub-scenes such as street inspection and vehicle control; the arrest module includes sub-scenes such as indoor arrest and outdoor arrest; and the emergency response module includes sub-scenes such as handling mass incidents and handling armed standoffs.
[0015] After scene selection, the system configures scene difficulty and environmental parameters based on the training objective. Scene difficulty parameters include the number of virtual target objects (e.g., 1 to 5), threat level (low, medium, and high), and time limit (e.g., 120 to 600 seconds). Environmental parameters include lighting conditions (daytime, nighttime, indoors, etc.), spatial complexity (open space, narrow passage, multi-room structure, etc.), and noise interference level. The system generates scene instance data based on the target training scene module and the above parameters, and loads the scene instance data into the mixed reality interactive environment to complete the rendering and deployment of the virtual scene.
[0016] Simultaneously with scene loading, the parameters of the four-layer behavior model are initialized. For example... Figure 1 As shown, the four-layer behavioral model comprises a control layer, a control layer, a tactical layer, and a closed-loop layer, employing a hierarchical structure. The control layer, as the foundational layer, includes three behavioral nodes: a situational awareness node, a tactical positioning node, and a verbal control node. Specifically, the situational awareness node assesses trainees' ability to identify hazards and assess potential threats in a virtual environment; the tactical positioning node assesses trainees' spatial positioning choices, including elements such as cover utilization, angle control, and distance maintenance; and the verbal control node assesses the trainees' proficiency in delivering verbal instructions, covering key law enforcement language such as identification, command issuance, and warnings.
[0017] The control layer, acting as the execution layer, implements specific control measures based on the control layer. It includes force assessment nodes, enforcement measures nodes, and escalation nodes. Force assessment nodes require trainees to accurately determine the necessity and proportion of force use based on the level of threat on-site and select the appropriate level of force. Enforcement measures nodes cover the assessment of enforcement measures at three levels: unarmed control, use of police equipment, and weapons readiness. Escalation nodes assess the timeliness and accuracy of trainees' escalation procedures, such as initiating reinforcement requests, tactical transitions, or site lockdowns under specific conditions.
[0018] The tactical layer, acting as the coordination layer, provides tactical support during the execution of the control layer. It includes police team coordination nodes, spatial search nodes, and dynamic decision-making nodes. The police team coordination nodes assess the clarity of task division, information transmission efficiency, and action synchronization in collaborative training scenarios involving multiple trainees. The spatial search nodes evaluate the effectiveness of trainees' search strategies in virtual spatial areas, including zoned searches, cross-coverage, and blind spot investigations. The dynamic decision-making nodes evaluate the quality of trainees' solution selection and response speed in the face of emergencies.
[0019] The closed-loop layer, serving as a support layer, completes the law enforcement procedure loop after the tactical layer's execution. It includes three nodes: legal notification, evidence preservation, and reporting / handover. The legal notification node assesses the completeness and standardization of trainees' notification of rights, explanation of obligations, and procedural interpretations during the law enforcement process. The evidence preservation node assesses trainees' ability to identify, protect, and preserve simulated evidence in a virtual scenario. The reporting / handover node assesses the standardization of the post-training report, material organization, and handover procedures.
[0020] There are preset hierarchical triggering conditions and state transition rules between each layer. Preferably, when all behavior nodes in the control layer are matched or the cumulative time exceeds a preset threshold (e.g., 60 seconds), the system automatically activates the set of behavior nodes in the control layer. In one embodiment of the present invention, hierarchical triggering supports two modes: conditional triggering and time triggering. Conditional triggering requires the previous layer to complete at least a preset proportion (e.g., 70%) of behavior node matching, while time triggering forces progress after a preset time window is reached.
[0021] Step S2: Real-time collection of user behavior data. After the scene is loaded, the system collects the trainees' behavior data in the mixed reality interactive environment in real time through the mixed reality interactive device. In this embodiment of the invention, the behavior data includes four dimensions of collection content.
[0022] In terms of spatial location data, the system uses the spatial tracking sensors built into the mixed reality headset to collect the trainee's head coordinates (x, y, z) and body orientation angle (θ) at a frequency of 60Hz. Preferably, the system simultaneously collects the spatial coordinates and motion trajectory data of the hand controllers to support accurate capture of hand movements. The spatial location data is mapped to the world coordinate system of the virtual scene after coordinate system transformation, with the transformation error controlled within 0.02m.
[0023] In terms of limb posture data, the system collects joint angle data (represented as quaternions), motion velocity vectors, and acceleration vectors of the trainee through an inertial measurement unit (IMU) and external motion capture sensors. In one embodiment of the invention, the system collects posture data from 25 key skeletal points, covering the major joint nodes of the head, trunk, upper limbs, and lower limbs. The collected raw data is processed by Kalman filtering for noise reduction before outputting a stable posture feature data stream.
[0024] In terms of voice data, the system collects the trainee's voice signal through the microphone array integrated into the mixed reality headset, with a sampling rate of 16kHz and a quantization precision of 16bit. After endpoint detection (VAD) and noise suppression processing, the voice signal is transcribed into text using speech recognition, while simultaneously extracting voice emotion features and volume level parameters. In this embodiment, the text result from speech recognition is used for subsequent speech control node matching, and the emotion features are used to assist in assessing the trainee's psychological state.
[0025] In terms of interactive action data, the system records interaction events between trainees and virtual objects in the virtual scene, including event types and trigger timestamps for operations such as picking up virtual items, triggering virtual machine shutdown, and operating virtual control panels.
[0026] The behavioral data from the four dimensions mentioned above are timestamped and then encapsulated into standardized data packets with an alignment accuracy controlled within 16ms (i.e., one frame rendering cycle) to ensure the synchronization of multi-source data on the timeline. Each standardized data packet contains a timestamp T, a spatial location vector P, pose data A, speech features V, and scene state S. The data packet output frequency is synchronized with the rendering frame rate, preferably 60Hz.
[0027] Step S3: Multi-dimensional behavioral feature vector construction and node matching. Based on the collected behavioral data, the system performs multi-dimensional behavioral feature vector construction and node matching. This step is one of the core technical aspects of this invention, achieving accurate matching between the trainee's real-time behavior and standard behavioral nodes through a multi-window sequential matching algorithm (MW-SM).
[0028] First, the system constructs a multi-dimensional behavioral feature vector. Specifically, the system extracts three sets of sub-feature vectors from continuously collected behavioral data. Spatial feature vector It includes 6 dimensional components: the trainee's x-axis coordinate, z-axis coordinate, orientation angle θ, x-axis velocity, z-axis velocity, and distance from the target object. Among them, the distance to the target object The calculation formula is: , Where: x and z are the horizontal coordinates of the trainee in the virtual scene's world coordinate system, in meters; and These are the coordinates of the virtual target object in the current scene within the same coordinate system, in meters (m). The unit is m, and the value ranges from 0 to 50m.
[0029] Action feature vector It includes five dimensions: arm joint angle, torso tilt angle, hand movement speed, hand movement acceleration, and gesture type identifier. (Speech feature vector) It includes three dimensions: keyword matching degree, voice sentiment value, and response time. The system concatenates these three sets of sub-feature vectors into a fused feature vector. The total number of dimensions is 14. , Secondly, the system employs the MW-SM algorithm for behavior node matching. The core idea of the MW-SM algorithm is to establish a sliding time window to segment the behavior data, and calculate the comprehensive matching score between the fused feature vector and the standard behavior node feature template within each time window. In this embodiment of the invention, the width of the sliding time window is set to 500ms (i.e., 30 frames of data), and the step size is set to 166ms (i.e., 10 frames of data), thereby achieving overlapping segmentation processing of the behavior data.
[0030] Within each time window, the system calculates spatial similarity. Action similarity and speech similarity Spatial similarity is calculated using normalized Euclidean distance: , in: This represents the spatial feature vector of the trainees within the current time window. The spatial reference feature vector for standard behavior nodes; The L2 norm (Euclidean distance) is expressed in meters. The normalization factor is set to the maximum effective distance of the scene (preferably set to 20m), so that... The range of values is normalized to [0, 1].
[0031] Action similarity is calculated using cosine similarity: , in: This represents the action feature vector of the trainee within the current time window; The action reference feature vector for the standard behavior node; the numerator is the inner product of the two vectors, and the denominator is the product of the magnitudes of the two vectors. The value range is [-1, 1], and in practical applications, it is truncated and mapped to [0, 1].
[0032] Voice similarity is determined using the keyword Jaccard similarity coefficient: , in: This is a set of law enforcement keywords extracted from the text output by speech recognition. The set of keywords is a pre-defined reference set for standard behavior nodes; the numerator is the number of elements in the intersection of the two sets; the denominator is the number of elements in the union of the two sets. The value range is [0, 1].
[0033] The overall matching score Score_i is calculated using a weighted fusion strategy: , in: , and These are spatial similarity weights, action similarity weights, and speech similarity weights, respectively, and their sum is 1. In a preferred embodiment of the present invention, , , The technical basis for this weighting configuration is that, in practical police scenarios, the spatial location selection and the execution of physical actions have roughly equal weight in their impact on the handling effect, while voice commands, although a necessary procedural element, have a relatively low impact on the immediate handling result. Preferably, the system supports dynamically adjusting the weighting configuration according to different training scenarios; for example, in interrogation scenarios that emphasize verbal control, the weighting can be adjusted accordingly. Increased to 0.3.
[0034] When the overall matching score When the matching threshold θ_match is exceeded, the system determines that the current behavior node has matched successfully. Preferably, θ_match is set to 0.75. This threshold was obtained through experimental calibration. When θ_match is below 0.65, the false match rate increases significantly, and when θ_match is above 0.85, the missed match rate increases. After a successful match, the system generates a current behavior status identifier containing a hierarchy identifier and a node identifier, and records the matching timestamp.
[0035] Step S4: Behavioral Deviation Detection and Classification. Based on behavioral node matching, the system performs behavioral deviation detection and classification. This step compares the current behavioral state identifier with the preset standard behavioral path to detect three types of behavioral deviations.
[0036] The first category is node missing deviation detection. The system maintains a set of necessary nodes in the standard behavior path. When a necessary node is not matched within a preset time window (preferably 80% of the total time of the current level), the system determines that the node is missing and generates a missing deviation flag. For example, in the control layer, if a trainee does not perform the identity declaration operation corresponding to a speech control node within a specified time, the system detects a missing deviation of the speech control node.
[0037] The second category is execution order deviation detection. The system compares the execution order of the actual action nodes with the preset order of the standard action path. When the actual execution order is inconsistent with the preset order, it is determined to be a sequence deviation. In this embodiment of the invention, the system quantifies the degree of sequence deviation by comparing the edit distance between the actual action node sequence and the standard sequence; the larger the edit distance, the more severe the deviation.
[0038] The third category is timeliness deviation detection. The system sets a preset reasonable execution time range for each behavior node. When the actual execution time of a behavior node exceeds this range, it is judged as a timeliness deviation. Preferably, the reasonable execution time range is parameterized according to the difficulty level of the training scenario. For example, the reasonable duration range for situational awareness nodes in low-difficulty scenarios is 5s to 15s, while in high-difficulty scenarios, this range is adjusted to 3s to 10s.
[0039] Each deviation detection generates a deviation type identifier, a deviation quantification value, and a deviation severity level (low, medium, high, and severe). The deviation quantification value is calculated as follows: for missing deviations, the quantification value is set to 1.0; for sequence deviations, the quantification value is the normalized edit distance between the actual sequence and the standard sequence; for timeliness deviations, the quantification value is the ratio of the difference between the actual duration and the upper limit of the reasonable range to the span of the reasonable range.
[0040] Step S5: Intelligent Generation and Real-Time Presentation of Prompt Messages. When the system detects a behavioral deviation, the intelligent prompt message generation mechanism is triggered. The core objective of this step is to transform abstract deviation detection data into guidance information that trainees can directly understand and execute, thereby achieving real-time correction during the training process.
[0041] First, the system determines the deviation category based on the deviation type identifier. The deviation categories include four types: missing procedures, operational errors, timing delays, and deviations from specifications. The missing procedures category corresponds to node-missing deviations, indicating that the trainee missed a required procedural step. The operational errors category corresponds to situations where the matching score exceeds the threshold but falls below the excellent standard (e.g., below 0.85), meaning that although the trainee's operation was identified, the execution quality was substandard. The timing delays category corresponds to timeliness exceeding limits, indicating that although the trainee's operation was correct, the response was too late. The deviations from specifications category correspond to sequence deviations, indicating that the trainee's operation execution order does not conform to the standard procedure.
[0042] Secondly, the system determines the prompt level based on the deviation category and the current training stage. The prompt levels are divided into three levels: guiding prompts, corrective prompts, and warning prompts. Guiding prompts are suitable for low-level deviations, gently guiding trainees to pay attention to overlooked steps, such as suggesting they verify the target's identity. Corrective prompts are suitable for medium-level deviations, clearly pointing out the deviation and providing standard operating procedure suggestions, such as currently not implementing safe distance control and suggesting maintaining a distance of at least 3 meters. Warning prompts are suitable for high or severe deviations, reminding trainees to immediately correct critical errors, such as warning: "Executing mandatory control without conducting a force level assessment is a serious procedural violation."
[0043] During the prompt generation process, the system retrieves prompt templates from the hierarchical prompt library that match the current deviation category, prompt level, and behavior node identifier, and fills in dynamic parameters (such as the specific deviation item name, time difference, suggested operation content, etc.) to generate complete real-time prompt information. Preferably, the prompt library is organized hierarchically according to the four-layer behavior model, with each layer containing prompt entries corresponding to the behavior nodes at that layer. Each entry has preset template texts at three levels: guidance, correction, and warning. In one embodiment of the present invention, the number of templates in the prompt library is no less than 200, covering the prompt scenarios for all 12 behavior nodes of the four-layer behavior model under the combination of four deviation categories and three prompt levels.
[0044] The real-time prompts are output via two channels: voice broadcast and visual overlay display. Voice broadcast is achieved through the built-in speakers or bone conduction headphones of the mixed reality headset. The speech synthesis engine converts the prompt text into natural speech, and the broadcast duration is controlled between 3 and 8 seconds to avoid information overload. Visual overlay display is achieved by overlaying a semi-transparent text prompt box onto the trainee's field of vision. The prompt box is located in the peripheral area of the trainee's field of vision (preferably in the lower right area) to avoid obstructing the trainee's main field of vision of the virtual scene. The prompt box automatically fades after 5 seconds. In one embodiment of the invention, the system sets a prompt cooldown time (preferably 5 seconds) to prevent repeated triggering of the same type of prompt during the cooldown time, thus avoiding high-frequency prompts interfering with the training process. Furthermore, the system supports training administrators in presetting prompt output strategies, such as enabling full prompt output in the basic training mode, outputting only warning-level prompts in the advanced training mode, and disabling all real-time prompts and only recording in the background in the assessment mode.
[0045] Step S6: Time-series recording of the behavior execution chain. During training, the system continuously records the complete behavior execution chain. The behavior execution chain is a data record of the entire training process organized according to a time sequence. It is the core data foundation for evaluating training effectiveness and post-training analysis, and includes three components: time-series data records, hierarchical state records, and deviation event records.
[0046] The time-series data records the occurrence time (accurate to 1ms) and duration of each behavioral event using a unified timestamp. In one embodiment of the invention, the time-series data is stored in a structured time-series format, with each record containing four fields: timestamp, event type identifier, associated behavioral node identifier, and event parameters. During training, the system continuously writes time-series data at a sampling frequency of 60Hz, generating approximately 36,000 time-series records in a typical 10-minute training session.
[0047] The hierarchical state record saves the currently active hierarchical level identifier and the currently active node identifier at each time point, forming a timeline of hierarchical state transitions. Through the hierarchical state record, the system can reconstruct the switching trajectory of trainees between different levels of the four-layer behavioral model during training, and analyze the distribution of trainees' dwell time at different levels and the timeliness of level switching. Preferably, when the system detects a level switching event, it automatically generates a switching timestamp and a switching trigger condition record for subsequent analysis of whether the level switching is driven by condition triggering (the previous level's completion target is met) or time triggering (reaching a preset time window).
[0048] The deviation event log stores the deviation type, deviation quantification value, deviation severity, and associated prompts for each deviation detection, as well as whether the trainee corrected the deviation in subsequent operations. In one embodiment of the invention, the deviation event log further includes a deviation correction time field to quantify the trainee's response speed to the prompts.
[0049] At the end of training, the system performs data integrity verification on the behavior execution chain. This verification includes: timestamp continuity verification (ensuring no data loss exceeding 500ms), hierarchical state coverage verification (ensuring each level of the four-layer behavior model has state records, with coverage time not less than 50% of the theoretical minimum duration), and behavior event completeness verification (ensuring records cover the entire period from training start to training end). After successful verification, the behavior execution chain, used as input data for subsequent evaluation calculations, is locked in a read-only state to ensure the immutability of the evaluation data.
[0050] Step S7: Six-Dimensional Weighted Evaluation Calculation. After training, the system performs a six-dimensional weighted evaluation calculation based on the behavior execution chain. This step uses a six-dimensional weighted evaluation algorithm (6D-WE) to objectively and quantitatively evaluate the training effect from six dimensions.
[0051] The first dimension is the completeness of the behavioral chain (D1), which evaluates the completeness of the law enforcement process. Its calculation formula is: , in: The number of behavior nodes that were actually successfully matched during training is a dimensionless positive integer, ranging from 0 to 1. ; This represents the total number of action nodes in the standard action path, a positive integer determined by the selected training scenario (e.g., an interrogation scenario). Number 8, arrest scene (12). The value ranges from 0 to 100, and the unit is minutes.
[0052] The second dimension is the procedural expression standardization dimension (D2), which evaluates the degree of standardization in legal notification and procedural execution. Its calculation is based on the matching rate between the speech-recognized text and standard law enforcement terminology templates, as well as the conformity of the execution order of procedural nodes.
[0053] , in: The speech standard matching rate is defined as the ratio of the number of standard law enforcement terms correctly used in actual speech to the total number of standard terms, and its value ranges from [0, 1]. The program order compliance rate is defined as the ratio of the number of program nodes executed in the correct order to the total number of program nodes, and its value ranges from [0, 1]. and For sub-dimension weight coefficients, preferably , ,make The value range is from 0 to 100.
[0054] The third dimension is the tactical decision rationality dimension (D3), which evaluates the rationality of tactical choices and execution. This dimension is calculated based on a combination of the matching scores of dynamic decision nodes and the search coverage of spatial search nodes.
[0055] The fourth dimension is the Collaboration Effectiveness Dimension (D4), which evaluates the effectiveness of police team collaboration. This dimension is applicable to multi-person collaborative training scenarios and is calculated based on the temporal synchronization and spatial coordination of the behavioral chains of multiple trainees. In single-person training mode, D4 is set to a fixed value of 80 points.
[0056] The fifth dimension is the accuracy of time control (D5), which evaluates the accuracy of time control at each stage. Its calculation formula is: , in: The number of behavioral nodes that require timeliness evaluation, which is a positive integer; The actual execution time of the j-th action node is expressed in seconds. This is the reference execution time for the j-th behavior node, in seconds. The allowable time deviation tolerance for the j-th behavior node is expressed in seconds, and preferably ranges from 2 to 5 seconds. This is the maximum timeliness deviation value for a single node, preferably taken as 30s; The value range is from 0 to 100.
[0057] The sixth dimension is the overall execution fluency dimension (D6), which evaluates the overall coherence of law enforcement actions. This dimension is calculated based on the statistical characteristics of the transition time between adjacent nodes in the action chain. The smaller the fluctuation of the transition time and the closer the mean is to the reference value, the higher the fluency score.
[0058] The overall evaluation result is obtained by weighted summation of the scores of the six dimensions and their corresponding weight coefficients: , in: The weight coefficient for the k-th dimension is... ; In a preferred embodiment of the present invention, the weights are configured as follows: (Completeness of the behavioral chain) (Program expression standardization) (Rationality of tactical decision-making) (Effectiveness of collaboration) (Accuracy of timeliness control) (Overall execution smoothness). The value range is from 0 to 100.
[0059] The system supports dynamically adjusting weight configurations based on training type and objectives. For example, in law enforcement training emphasizing procedural standardization, weights can be dynamically adjusted. Increase to 0.30; in police team training that emphasizes tactical coordination, it can be... and They were increased to 0.25 and 0.20 respectively.
[0060] Step S8: Evaluation Result Output and Training Feedback. After training, the system outputs a complete evaluation report, including a visual representation of the behavior execution chain, sub-scores for six dimensions, overall evaluation score, and evaluation level. In one embodiment of the invention, the evaluation level is divided into four levels based on the overall evaluation score: Excellent (90 to 100 points), Good (75 to 89 points), Pass (60 to 74 points), and Unsatisfactory (0 to 59 points). The evaluation report also includes detailed analysis and targeted improvement suggestions for each dimension, providing data support and directional guidance for the trainees' subsequent training. Preferably, the evaluation report also includes a behavior chain sequence diagram, which uses the time axis as the horizontal axis and the four-layer behavior model hierarchy as the vertical axis, intuitively displaying the trainees' behavior node trigger trajectory, level switching time points, and deviation event distribution during the training process. Instructors can quickly locate the trainees' weaknesses through the behavior chain sequence diagram. In addition, the evaluation module supports longitudinal comparative analysis of historical training data, which compares the current training score with the trainee's previous historical scores to generate a capability growth curve, helping instructors and trainees to quantitatively track the continuous improvement of training effectiveness.
[0061] The dual-field training mode used in this embodiment of the invention is as follows: Figure 4 As shown, the training consists of two parts: a virtual training area and a physical training area. The virtual training area comprises a mixed reality interactive environment, presenting a 3D virtual training scene to trainees through a mixed reality head-mounted display, including a virtual architectural environment, virtual target characters (NPCs), virtual props, and real-time prompts. The physical training area consists of a physical training space, equipped with simulated police equipment (such as simulated handcuffs and batons), simulated weapons (such as laser-emitting simulated guns), and protective gear (such as training bulletproof vests and helmets) to provide trainees with realistic physical tactile feedback. The virtual and physical training areas are mapped using spatial registration technology. During spatial registration, a coordinate system is established using a fixed point in the physical training area as the origin, aligning the virtual scene coordinate system with the physical space coordinate system with an alignment accuracy controlled within 0.05m. By integrating the two fields, the physical actions performed by trainees in the real field (such as raising a gun, running, and fighting) can generate corresponding visual responses in the virtual field. At the same time, changes in the scene in the virtual field (such as NPC behavior reactions) can also be fed back to the trainees through visual and auditory channels, thereby achieving an immersive training experience that integrates the virtual and real.
[0062] like Figure 2As shown in the figure, this embodiment of the invention also provides an MR police combat training system based on a four-layer behavior model. This system includes five core functional modules: an MR interaction module, a behavior recognition module, a logic judgment module, a prompt generation module, and an evaluation module. The structure and working principle of each module are described in detail below.
[0063] The MR interaction module is the system's human-computer interaction front end, undertaking three core functions: scene presentation, data acquisition, and prompt output. It comprises three sub-units: a scene engine unit, a tracking and positioning unit, and a display output unit. The scene engine unit is responsible for the 3D rendering and real-time updating of the virtual training scene. Internally, it integrates a scene resource manager and rendering pipeline, supporting the loading and instantiation of various training scenes from the scene module library. The scene rendering frame rate is maintained at 60fps to ensure smooth visuals. The tracking and positioning unit integrates a spatial tracking sensor (preferably using an inside-out tracking scheme, covering a physical space of 10m × 10m), an inertial measurement unit (six-axis IMU, 200Hz sampling rate), and a microphone array (dual-channel noise-canceling microphone, 16kHz sampling rate) to achieve multimodal real-time acquisition of trainees' spatial position data, limb posture data, and voice data. The display output unit presents the virtual scene and real-time prompts to the trainees through the optical display system of the mixed reality head-mounted display. The display resolution is preferably 2K per eye, with a field of view greater than 52 degrees. In one embodiment of the present invention, the MR interaction module adopts a perspective mixed reality head-mounted display device (such as a head-mounted device based on optical waveguide technology). This device supports the optical overlay display of virtual content and real physical environment, so that trainees can see the overlaid virtual scene elements while observing the real training site, thereby realizing a dual-field fusion training experience.
[0064] The behavior recognition module is the system's perception processing center, responsible for transforming raw behavior data into structured behavior feature descriptions. It comprises three sub-units: a data acquisition unit, a feature extraction unit, and a matching calculation unit. The data acquisition unit receives the raw behavior data stream transmitted by the MR interaction module via a high-speed data bus, performs timestamp alignment (alignment error less than 16ms) and multi-source data synchronization processing, unifying sensor data with different sampling rates to a processing frequency of 60Hz. The feature extraction unit sequentially performs Kalman filtering noise reduction (for position data) and moving average filtering (for action data) on the synchronized behavior data, then constructs spatial feature vectors (6-dimensional), action feature vectors (5-dimensional), and speech feature vectors (3-dimensional), concatenating them into a 14-dimensional fused feature vector. The matching calculation unit implements the MW-SM algorithm described in step S3 of the aforementioned method embodiment, performing real-time matching calculations of behavior nodes within a sliding time window (window width 500ms, step size 166ms). Preferably, when performing matching, the matching calculation unit first pre-filters the candidate node set based on the current scene context, retaining only behavior nodes related to the current training scene and the current activation level for matching calculation, thereby reducing computational complexity and improving matching efficiency. The output of the matching calculation unit includes the matching node identifier, matching score, and confidence parameter.
[0065] The logic judgment module is the core of the system's decision-making and reasoning, responsible for behavior deviation detection and training state management. It includes three sub-units: a model storage unit, a rule engine unit, and a path management unit. The model storage unit stores the complete parameter set of the four-layer behavior model in the form of a structured database, including standard feature vectors for each of the 12 behavior nodes in each layer, layer-level trigger condition parameters (condition trigger threshold of 70% and time trigger window value), and a state transition rule matrix. The rule engine unit implements the behavior deviation detection logic described in step S4 of the aforementioned method embodiment, performing three types of deviation judgments on the matching results output by the behavior recognition module: missing value detection, sequence detection, and timeliness detection, and generating deviation type identifiers, deviation quantification values, and deviation severity levels. The path management unit maintains a standard behavior path database, storing standard behavior node sequences corresponding one-to-one with each training scenario, along with their execution order constraints and time constraint parameters, supporting training administrators in customizing and modifying standard behavior paths according to training needs.
[0066] The prompt generation module is the system's feedback generation engine, responsible for converting deviation detection results into understandable prompt information. It comprises three sub-units: a corpus unit, a generation engine unit, and an output control unit. The corpus unit stores a prompt template library organized hierarchically according to a four-layer behavioral model. In one embodiment of this invention, the prompt template library contains over 200 prompt templates, covering three levels of prompts: guidance, correction, and warning, for 12 behavioral nodes in the four-layer behavioral model. The generation engine unit, based on the deviation detection results output by the logic judgment module, first determines the deviation category and prompt level. Then, it retrieves a prompt template from the corpus that matches the current deviation category, prompt level, and behavioral node identifier, and fills the template with dynamic parameters (such as the specific deviation item name, suggested operation content, and remaining time) to generate complete real-time prompt information. The output control unit manages the timing of real-time prompt information output, the selection of the output channel (voice broadcast or visual overlay), and the control of the cooling time (a default 5-second cooling interval for similar prompts). This ensures that the output of prompt information does not excessively interfere with the trainees' normal training process. The generated prompt information is then transmitted to the display output unit of the MR interaction module via the system's internal communication interface for presentation.
[0067] The evaluation module is the system's performance quantification engine, responsible for the complete recording of training process data and the quantitative evaluation of training effects. It includes three sub-units: a recording unit, a calculation unit, and a reporting unit. The recording unit implements the behavior execution chain timing recording function described in step S6 of the aforementioned method embodiment. It continuously records behavior data, status identifiers, deviation events, and prompt trigger records throughout the training process and performs data integrity verification at the end of training. The calculation unit implements the 6D-WE algorithm described in step S7 of the aforementioned method embodiment. It extracts evaluation indicators for each dimension from the behavior execution chain and calculates the six-dimensional sub-scores and the overall evaluation result. During the calculation process, it obtains weight configuration parameters from the logic judgment module to support dynamic weight adjustment under different training types. The reporting unit formats the evaluation results into a structured evaluation report. The report includes a six-dimensional radar chart, a behavior chain timing diagram, a list of deviation events, and targeted improvement suggestions. It supports presentation to instructors and trainees via external display devices (such as tablets, projectors, or large screens).
[0068] The data flow between modules forms a complete closed-loop processing chain: the MR interaction module sends the collected behavioral data to the behavior recognition module; the behavior recognition module sends the matching results to the logic judgment module for deviation detection, and simultaneously sends them to the evaluation module for recording; the logic judgment module sends the deviation detection results to the prompt generation module to trigger prompt generation; the prompt generation module returns the generated real-time prompt information to the MR interaction module for output display; the evaluation module obtains the evaluation rule parameters from the logic judgment module and outputs the calculated evaluation results to an external display device. This data flow ensures a closed-loop processing chain from behavioral data collection, feature extraction, node matching, deviation detection, prompt generation to evaluation calculation, with each module achieving efficient collaboration through standardized data interfaces.
[0069] To verify the effectiveness and practicality of the technical solution of this invention, the inventors organized a systematic pre- and post-test controlled experiment at a police training base in Guangdong Province. Fifty in-service law enforcement officers were recruited as trainees and randomly divided into an experimental group (25 participants) and a control group (25 participants). There were no significant differences between the two groups in terms of age, years of service, and training background. The experimental group was trained using the system described in this invention. During the training process, the system continuously implemented four-layer behavioral model matching, real-time deviation detection and feedback, and six-dimensional evaluation functions. The control group received training using traditional instructor guidance methods, led by the same group of senior instructors. Both groups underwent the same training period of four weeks, with three training sessions per week, each lasting 60 minutes. The training scenarios covered three core police scenarios: interrogation, arrest, and emergency response.
[0070] After training, both groups of trainees underwent testing in the same standardized assessment scenario. The system automatically calculated six-dimensional evaluation scores. The results showed that the experimental group's scores in all six dimensions were significantly higher than the control group and the pre-training baseline. Specifically, the average score for the Behavioral Chain Integrity dimension (D1) increased by 14.1 points (from 71.2 before training to 85.3 after training), the average score for the Procedural Expression Normativity dimension (D2) increased by 14.8 points (from 68.5 to 83.3), the average score for the Tactical Decision-Making Rationality dimension (D3) increased by 11.7 points, and the average score for Timeliness Control Accuracy dimension (D5) increased by 10.9 points. The weighted total score across all six dimensions showed an overall improvement of 12.4 points for the experimental group. Paired-samples t-test analysis of the pre- and post-test data showed that the improvement in scores for each dimension was statistically significant (P < 0.01). Meanwhile, the independent samples t-test between the experimental group and the control group also showed that the difference was statistically significant (P < 0.05), proving that the technical solution of the present invention has significant advantages over traditional training methods in improving training effectiveness.
[0071] This invention forms a closed-loop collaborative technical system through the deep coupling of five core technical links: four-layer behavior model construction, multi-window time-series matching algorithm, real-time behavior deviation detection, intelligent generation of prompt information, and six-dimensional weighted evaluation algorithm. There are close data dependencies and functional synergistic effects among the links.
[0072] From a data flow perspective, the four-layer behavioral model provides a structured knowledge representation foundation for the entire system. The 12 defined behavioral nodes and their standard feature vectors serve as the reference template for the MW-SM matching algorithm. Standard behavioral paths and their temporal constraints provide the basis for the deviation detection module's judgment. The hierarchical correspondence between behavioral nodes and the prompt corpus forms the foundation for corpus retrieval in prompt information generation. The dimensional mapping between the six-dimensional evaluation index system and the four-layer behavioral model forms the computational framework for the 6D-WE evaluation algorithm. All four technical aspects utilize the four-layer behavioral model as a common knowledge base, ensuring semantic consistency throughout the entire system chain from behavior acquisition to effect evaluation.
[0073] From a functional synergy perspective, the real-time matching results of the MW-SM algorithm directly drive the judgment logic of the deviation detection module—only after the behavioral nodes are matched can the system determine whether there are missing, sequential, or timeliness deviations. The deviation detection results then trigger the workflow of the prompt generation module, with the prompt level and content depending on the quantitative output of the deviation type and severity. The real-time presentation of prompt information guides the trainees' subsequent behavior. This guided behavioral data returns to the MW-SM matching stage for a new round of processing, forming a real-time feedback loop of matching-detection-prompt-behavior correction-rematching. After training, the complete behavioral execution chain data (including matching records, deviation events, and prompt trigger records) serves as input to the 6D-WE algorithm, objectively and quantitatively evaluating the training effect from six complementary dimensions.
[0074] From a technical perspective, the dual-domain training mode, through spatial registration and fusion of virtual and real-world domains, simultaneously provides visual immersion and physical tactile feedback within a mixed reality environment. This overcomes the limitations of single VR systems that lack physical interaction realism, allowing trainees to receive immersive feedback from virtual scenarios while operating realistic simulation equipment, significantly improving the transfer of training experience to real-world practice. The hierarchical structure of the four-layer behavioral model closely matches the phased characteristics of real law enforcement procedures, enabling trainees to systematically master complete law enforcement skills from on-site control to procedural closure under the guidance of a structured framework. The coupling relationship between the aforementioned technical components ensures a closed-loop operation of the entire training process, from scenario construction, behavior collection, real-time feedback to effect evaluation. This allows the training system to provide immediate corrective guidance while trainees perform operations and output quantifiable, comparable, and traceable evaluation data after training, providing effective technical support for the standardization and scientification of police training.
[0075] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A method for MR police practical training based on a four-layer behavior model, characterized in that, Includes the following steps: S1: Scene loading and behavior model initialization: Load the preset training scene into the mixed reality interactive environment and initialize the four-layer behavior model corresponding to the training scene; S2: Real-time collection of user behavior data. The behavior data of trainees in the mixed reality interactive environment is collected in real time through mixed reality interactive devices. The behavior data includes spatial location data, body posture data, voice data and interactive action data. S3: Multidimensional behavior feature vector construction and node matching: The behavior data is converted into a multidimensional behavior feature vector, and a multi-window temporal matching algorithm is used to match the multidimensional behavior feature vector with the preset behavior nodes in the four-layer behavior model in real time to generate the current behavior status identifier. S4: Behavioral deviation detection and classification, compare the current behavioral state identifier with the preset standard behavioral path, detect node missing deviation, execution order deviation and time limit exceeding deviation, and generate deviation type identifier and deviation quantification value; S5: Intelligent generation and real-time presentation of prompt information. Based on the deviation type identifier and the current behavior status identifier, the prompt level is determined and a matching prompt template is retrieved from the hierarchical prompt library to generate real-time prompt information and present it to the trainees through the mixed reality interactive device. S6: Behavior execution chain timing record, recording the behavior data, the current behavior status identifier, the deviation type identifier and the real-time prompt information according to a unified timestamp, forming a complete behavior execution chain; S7: Six-dimensional weighted evaluation calculation: Based on the aforementioned behavior execution chain, a six-dimensional weighted evaluation algorithm is used to calculate evaluation scores from six dimensions: completeness of the behavior chain, standardization of program expression, rationality of tactical decision-making, effectiveness of coordination, accuracy of timeliness control, and overall smoothness of execution, and generate an overall evaluation result; S8: Evaluation result output and training feedback, outputting the behavior execution chain, the sub-scores of the six dimensions, and the overall evaluation result.
2. The method of claim 1, wherein, The four-layer behavioral model includes a control layer, a tactical layer, and a closed-loop layer. Each layer contains preset behavioral nodes and hierarchical triggering conditions. The control layer includes situational awareness nodes, tactical positioning nodes, and verbal control nodes. The tactical layer includes force assessment nodes, coercive measures nodes, and procedural escalation nodes. The tactical layer includes police team coordination nodes, spatial search nodes, and dynamic decision-making nodes. The closed-loop layer includes legal notification nodes, evidence preservation nodes, and reporting and handover nodes.
3. The method of claim 2, wherein, The four-layer behavioral model adopts a hierarchical structure. The control layer is the basic layer used to establish the initial control situation at the law enforcement site. The control layer is the execution layer used to implement specific control measures based on the control layer. The tactical layer is the coordination layer used to provide tactical support during the execution of the control layer. The closed-loop layer is the guarantee layer used to complete the closed loop of the law enforcement procedure after the execution of the tactical layer. There are preset hierarchical trigger conditions and state transition rules between each layer.
4. The method of claim 3, wherein, The multi-window temporal matching algorithm includes: establishing a sliding time window to segment the behavior data; calculating spatial similarity, action similarity, and speech similarity within each time window; calculating a comprehensive matching score using a weighted fusion strategy; and determining that the current behavior node is successfully matched when the comprehensive matching score exceeds a preset matching threshold, and recording the matching timestamp.
5. The method of claim 4, wherein, The behavior deviation detection and classification include: missing detection, where a behavior node in the standard behavior path is determined to be missing when it is not matched within a preset time window; sequence detection, where the execution order of the actual behavior nodes is inconsistent with the preset order of the standard behavior path, which is determined to be a sequence deviation; and timeliness detection, where the execution time of a behavior node exceeds a preset reasonable range, which is determined to be a timeliness deviation.
6. The method of claim 5, wherein, In the six-dimensional weighted evaluation algorithm, the weight coefficients of each dimension are dynamically adjusted according to the training type and training objective. The overall evaluation result is obtained by weighted summation of the scores of each dimension and the corresponding weight coefficients.
7. The method of claim 6, wherein, The intelligent generation of the prompt information includes: determining the deviation category based on the deviation type identifier, wherein the deviation category includes program missing, operation error, timing delay, and deviation from standard; determining the prompt level based on the deviation category and the current training stage, wherein the prompt level includes guidance prompts, correction prompts, and warning prompts; and the output method of the real-time prompt information includes voice broadcast and visual overlay display.
8. The method of claim 7, wherein, The method employs a dual-field training mode, comprising: a virtual field, consisting of the mixed reality interactive environment, used to present an overlay display of virtual training scenarios and real-time prompts; and a physical field, consisting of a physical training space, equipped with simulated police equipment and physical tactile feedback devices; the virtual field and the physical field are mapped to each other through spatial registration.
9. The method according to claim 8, characterized in that, The scene loading and behavior model initialization include: selecting a target training scene module from a preset scene module library, which includes a search module, an arrest module, an emergency response module, and a rescue module; configuring scene difficulty parameters and environmental parameters according to the training target; generating scene instance data according to the target training scene module and the scene difficulty parameters and loading it into the mixed reality interactive environment.
10. A four-layer behavioral model-based MR police combat training system, used to implement the method described in any one of claims 1-9, characterized in that, include: The MR interaction module is configured to load and present virtual training scenes, collect spatial position data, body posture data, voice data and interactive action data of trainees, and output real-time prompt information. The behavior recognition module is configured to receive behavior data collected by the MR interaction module, convert the behavior data into a multi-dimensional behavior feature vector, and perform real-time matching of behavior nodes using a multi-window temporal matching algorithm. The logic judgment module is configured to store four-layer behavior model parameters and standard behavior paths, and compare the behavior node matching results with the standard behavior paths to perform behavior deviation detection and classification. The prompt generation module is configured to determine the prompt level based on the deviation detection result, retrieve matching prompt templates from the hierarchical prompt library and generate real-time prompt information, and transmit the real-time prompt information to the MR interaction module for output; The evaluation module is configured to record the behavior execution chain, and uses a six-dimensional weighted evaluation algorithm to calculate the evaluation score and output the evaluation result.