Tactical action specification AI correction system and method thereof

Through the combination of multimodal perception and intelligent discrimination, the problems of subjective evaluation bias and strong environmental dependence in traditional tactical action training have been solved, and high-precision, real-time and personalized tactical action evaluation and correction have been achieved, improving training effectiveness and evaluation objectivity.

CN120627804APending Publication Date: 2025-09-12CHINESE PEOPLES LIBERATION ARMY ARMY MEDICAL UNIV BORDER DEFENSE TRAINING BRIGADE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510753567.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional tactical action training has problems such as subjective assessment bias, strong environmental dependence and severe feedback delay, and cannot achieve high-precision, real-time and personalized action assessment and correction.

Method used

It adopts a combination of multimodal perception layer, intelligent discrimination layer and interactive feedback layer, collects data through millimeter wave radar array, infrared thermal imager and wide-angle stereo vision module, performs spatiotemporal alignment and fusion, identifies tactical-specific skeletal nodes, generates personalized action standard thresholds, and corrects the trainee's movements through multi-channel feedback.

Benefits of technology

It has achieved improved training efficiency, enhanced objectivity of evaluation, strong environmental adaptability and real-time feedback correction. The pass rate of recruits' tactical actions has been improved, the consistency of system scoring has been improved, the feedback delay has been controlled within 400ms, and the personalized training effect has been significant.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120627804A_ABST
    Figure CN120627804A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a tactical action specification AI correction system and method, and the method comprises the steps that a multi-mode sensing layer collects tactical action data of a trainee through a millimeter wave radar array, an infrared thermal imager and a wide-angle stereoscopic vision module; the intelligent discrimination layer comprises a heterogeneous data fusion module, a tactical special skeleton recognition module, an action matching module, a personalized standard generation module and a priority evaluation module, and realizes high-precision processing and analysis of action data; the interaction feedback layer corrects actions of trainees in real time through multi-channel voice prompts, augmented reality tactical sand tables and tactile feedback, solves the problem of environmental dependence by fusing a multi-mode perception technology, improves recognition precision in a complex environment by adopting a 32-point extended skeleton model and adversarial training, dynamically generates personalized standards based on individual features, and improves recognition accuracy. The problems of subjective evaluation deviation, high environmental dependence, serious feedback delay and the like of traditional tactical training are solved, and the training efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and specifically to an AI correction system and method for tactical action standardization, which is suitable for action standardization assessment and real-time correction in scenarios such as military training, special operations training, and police tactical training. Background Art

[0002] Traditional tactical action training relies primarily on manual observation, scoring, and two-dimensional video playback analysis, which presents numerous limitations. First, manual assessments are subject to subjective bias. For example, the error in interpreting the torso's height above the ground during a "low-posture crawl" can reach ±5cm, and it is impossible to quantify three-dimensional trajectory parameters. Second, traditional training methods are highly dependent on the environment. Conventional optical cameras are constrained by lighting conditions and are less effective at night or in smoky environments. Furthermore, the frame rate (≤60fps) makes it difficult to capture transient tactical movements. Furthermore, action correction relies on post-training review, which results in significant feedback delays (≥2 hours), making it impossible to immediately correct incorrect action patterns during training execution.

[0003] Existing technologies attempt to incorporate inertial sensors or monocular vision algorithms for motion analysis, but these still suffer from issues such as insufficient sensor fusion and poor algorithm robustness. While a single IMU device can measure limb angular velocity, it lacks global spatial positioning capabilities. The RGB camera-based OpenPose algorithm has a joint recognition error rate as high as 25% in complex backgrounds (where camouflage uniforms and vegetation colors are mixed).

[0004] The main flaws of the current technical solution include: the traditional skeletal model contains only 17 joints and cannot analyze the details of tactical movements; the clock synchronization error of millimeter-wave radar, infrared and visual data is greater than 50ms, resulting in the failure of spatiotemporal alignment; the preset "standard action" is a rigid threshold (such as the leap step length is fixed at 1.2m), which does not take into account individual differences between soldiers and the dynamic battlefield environment. Summary of the Invention

[0005] The purpose of the present invention is to provide an AI correction system and method for tactical action standardization. Through the organic combination of multimodal perception, intelligent discrimination and real-time feedback, it overcomes the problems of subjective evaluation bias, strong environmental dependence and severe feedback delay in traditional training methods, and realizes high-precision, real-time and personalized evaluation and correction of tactical actions.

[0006] The present invention proposes an AI correction system for tactical action standardization, comprising:

[0007] A multimodal perception layer, used to collect trainee tactical movement data, comprising a millimeter-wave radar array, an infrared thermal imager, and a wide-angle stereo vision module;

[0008] An intelligent discrimination layer, connected to the multimodal perception layer, is configured to receive tactical action data collected by the multimodal perception layer, identify action errors in the tactical action data, and generate correction instructions. The intelligent discrimination layer includes:

[0009] a heterogeneous data fusion module for performing spatiotemporal alignment and fusion of the data collected by the millimeter-wave radar array, the infrared thermal imager, and the wide-angle stereo vision module to generate a spatiotemporal cube representation;

[0010] a tactical-specific skeleton recognition module, connected to the heterogeneous data fusion module, for identifying the positions of the trainee's skeleton joints based on the space-time cube representation and generating an extended skeleton model containing tactical-specific nodes;

[0011] An action matching module, connected to the tactical-specific skeleton recognition module, is used to extract multi-dimensional tactical action features from the extended skeleton model and perform real-time matching and comparison between the multi-dimensional tactical action features and a standard action library;

[0012] A personalized standard generation module is used to generate personalized action standard thresholds based on the individual characteristics of the trainee;

[0013] a priority evaluation module, connected to the action matching module and the personalized standard generation module, configured to receive the matching results of the action matching module, evaluate the severity of action errors based on the personalized action standard threshold, prioritize multiple errors, and generate correction instructions;

[0014] The interactive feedback layer is connected to the intelligent discrimination layer, and is used to receive the correction instructions generated by the intelligent discrimination layer and feed back the correction instructions to the trainee in a multi-channel manner.

[0015] Preferably, the heterogeneous data fusion module includes:

[0016] A multi-source data preprocessing unit, configured to perform noise reduction, filtering, and feature enhancement processing on the raw data collected by the millimeter-wave radar array, the infrared thermal imager, and the wide-angle stereo vision module;

[0017] a time synchronization unit, connected to the multi-source data pre-processing unit, for performing time alignment on different sensor data within a 50 millisecond time window;

[0018] a spatial alignment unit, connected to the time synchronization unit, for mapping different sensor data into a unified spatial coordinate system;

[0019] The multimodal feature fusion unit is connected to the spatial alignment unit and is used to fuse the multimodal features after spatiotemporal alignment through an attention mechanism to generate a spatiotemporal cube representation with a temporal resolution of 10 milliseconds and a spatial resolution of 3 millimeters.

[0020] Preferably, the tactical-specific skeleton recognition module includes:

[0021] Extended joint structure unit, used to expand the standard 17-point skeleton model to 32 points, including:

[0022] 24 human skeleton nodes, including finely divided neck, shoulder and hand joints;

[0023] There are 8 points associated with tactical equipment, including the muzzle pointing point, the buttstock fitting point, the tactical vest center point, and the helmet edge point;

[0024] A multi-scene adaptation network unit, connected to the extended joint structure unit, for extracting multi-scale features through a deep residual network and a feature pyramid structure;

[0025] The adversarial training enhancement unit is connected to the multi-scenario adaptation network unit and is used to improve the joint point recognition accuracy in a camouflage environment by generating an adversarial network, thereby increasing the joint point recognition accuracy to 92%.

[0026] Preferably, the action matching module includes:

[0027] A multi-dimensional feature extraction unit, configured to extract, based on the extended skeleton model:

[0028] Geometric features, including key angle parameters, key distance parameters and contour features;

[0029] Dynamic characteristics, including acceleration parameters, angular velocity parameters and impulse characteristics;

[0030] Synergistic characteristics, including limb coordination parameters, balance parameters, and rhythm parameters;

[0031] A dynamic time warping matching unit, connected to the multi-dimensional feature extraction unit, is used to calculate the similarity between the real-time action and the standard action library through a multi-level sampling mechanism and feature weight adaptation technology;

[0032] A real-time computing pipeline unit is connected to the dynamic time warping matching unit and is used to achieve 20 complete action evaluations per second based on a sliding window mechanism and a parallel computing architecture.

[0033] Preferably, the personalized standard generation module includes:

[0034] Soldier Characteristic Modeling Unit, which records and analyzes trainee physiological characteristics, competency ratings, and training history data;

[0035] a parameter mapping unit, connected to the soldier characteristic modeling unit, for generating personalized mapping rules for movement standards based on the trainee's height, weight, limb proportions, and ability rating;

[0036] The dynamic threshold generation unit is connected to the parameter mapping unit and is used to generate a four-level evaluation standard including an ideal value, an acceptable range, a warning threshold and an error threshold according to the individual characteristics and training progress of the trainee.

[0037] Preferably, the priority evaluation module includes:

[0038] An error severity assessment unit, which is used to calculate the severity of an error based on its impact on tactical effectiveness and potential risk of injury;

[0039] a prioritization unit, connected to the error severity evaluation unit, for prioritizing the plurality of errors detected and generating a processing queue;

[0040] a correlation error merging unit, connected to the priority sorting unit, for identifying multiple errors with causal relationships and merging them into a single comprehensive correction instruction;

[0041] The training process adjustment unit is connected to the associated error merging unit and is used to dynamically adjust the training difficulty and content based on the performance evaluation results of the trainee.

[0042] Preferably, the interactive feedback layer includes:

[0043] A multi-channel voice prompt unit, used to trigger voice commands based on the severity of the error, including:

[0044] Level 1 prompt, used to indicate mild deviations in a gentle tone;

[0045] Level 2 prompt, used to quickly warn of serious errors;

[0046] An augmented reality tactical sandbox unit is connected in parallel with the multi-channel voice prompt unit and is used to superimpose virtual cover and standard action paths through a head-mounted display to assist trainees in establishing a spatial action cognitive model;

[0047] The tactile feedback unit is connected in parallel with the multi-channel voice prompt unit and the augmented reality tactical sandbox unit, and is used to provide directional tactile prompts through the vibration element on the tactical vest.

[0048] Preferably, the tactical action data includes:

[0049] Leaping movement data, including step length dispersion, trunk height, and thermal radiation distribution of left and right leg muscles;

[0050] Prone action data, including prone speed, posture stability and final concealment effect;

[0051] Shooting posture data, including gun grip stability, muzzle pointing angle, and body support point distribution;

[0052] Cover utilization data, including body exposure area, cover occlusion angle, and position adjustment trajectory.

[0053] Preferably, the system further comprises:

[0054] A training data management module, connected to the intelligent discrimination layer, is used to store and analyze the trainee's historical training data and generate training progress reports and skill mastery assessments;

[0055] An environmental adaptation module, connected to the multimodal perception layer, is used to dynamically adjust sensor parameters and processing strategies based on the current training environment conditions to ensure system stability under different lighting, weather and terrain conditions;

[0056] The system self-diagnosis module is connected to the multimodal perception layer, the intelligent judgment layer and the interactive feedback layer, and is used to monitor the working status of each component of the system in real time, perform regular self-calibration, and locate faults when an anomaly is detected.

[0057] AI correction methods for tactical action standardization include:

[0058] Collecting the trainee's tactical movement data through a multimodal perception layer, which includes a millimeter-wave radar array, an infrared thermal imager, and a wide-angle stereo vision module;

[0059] The data collected by the millimeter wave radar array, the infrared thermal imager and the wide-angle stereo vision module are subjected to spatiotemporal alignment and fusion by a heterogeneous data fusion module to generate a spatiotemporal cube representation;

[0060] Identifying the positions of the trainee's skeleton joints based on the space-time cube representation through a tactical-specific skeleton recognition module, and generating an extended skeleton model including tactical-specific nodes;

[0061] Extracting multi-dimensional tactical action features from the extended skeleton model through an action matching module, and performing real-time matching and comparison between the multi-dimensional tactical action features and a standard action library;

[0062] Generate personalized action standard thresholds based on the trainee's individual characteristics through a personalized standard generation module;

[0063] Receiving the matching results of the action matching module through the priority evaluation module, evaluating the severity of the action error based on the personalized action standard threshold, prioritizing multiple errors, and generating correction instructions;

[0064] The correction instruction is received through the interactive feedback layer, and the correction instruction is fed back to the trainee in a multi-channel manner, wherein the multi-channel manner includes voice prompts, augmented reality sandbox display and tactile feedback.

[0065] The present invention has the following beneficial effects:

[0066] 1. Significant improvement in training effectiveness: Actual measurement data from an army training base (N=300 personnel, three-month cycle) showed that the pass rate for tactical maneuvers among new recruits increased from 65% to 88%, and the training cycle was shortened from 120 hours to 78 hours (a 35% reduction).

[0067] 2. Enhanced objectivity of evaluation: The consistency coefficient (Cohen's Kappa coefficient) between system scores and manual scores by experienced instructors has been improved from 0.62 to 0.89. Quantitative evaluation indicators have been supported, such as cover utilization efficiency (= effective shielding time / total exposure time), which has been optimized from 0.71 to 0.93.

[0068] 3. Strong environmental adaptability: Under harsh conditions such as heavy rain and at night, the system relies on millimeter-wave radar and infrared thermal imaging to maintain a joint point recognition rate of ≥85%, breaking through the environmental limitations of traditional vision systems.

[0069] 4. Personalized training implementation: Dynamically generate standard parameters suitable for individual characteristics. For example, the standard range of leap length for a soldier with a height of 175 cm is 1.1-1.3 meters, avoiding the unfair evaluation of soldiers of different body shapes by traditional fixed standards.

[0070] 5. Real-time feedback correction: The full-link delay from perception to feedback is controlled within 400ms. The system can trigger correction within an average of 4.3 seconds after the first occurrence of an erroneous action, effectively preventing the erroneous action from becoming fixed. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 This is a schematic diagram of the overall architecture of the AI ​​correction system for tactical action standardization of the present invention;

[0072] Figure 2 Schematic diagram of the structure of the multimodal perception layer of the present invention;

[0073] Figure 3 Schematic diagram of the structure of the intelligent discrimination layer of the present invention;

[0074] Figure 4 Schematic diagram of the process of the heterogeneous data fusion module of the present invention;

[0075] Figure 5 This is a schematic diagram of the extended joint point structure of the tactical skeleton recognition module of the present invention;

[0076] Figure 6 Schematic diagram of the workflow of the action matching module of the present invention;

[0077] Figure 7 Schematic diagram of parameter mapping relationship of the personalized standard generation module of the present invention;

[0078] Figure 8 Schematic diagram of the error severity assessment process of the priority assessment module of the present invention;

[0079] Figure 9 This is a schematic diagram of multi-channel feedback of the interactive feedback layer of the present invention;

[0080] Figure 10 Schematic diagram of the flow of the AI ​​correction method for tactical action standardization of the present invention. DETAILED DESCRIPTION

[0081] Please refer to the attached Figure 1-10 , the specific implementation of the present invention is further described in detail below with reference to the accompanying drawings.

[0082] like Figure 1 As shown in the figure, the AI ​​correction system for tactical action standardization provided by the present invention includes a multimodal perception layer 1, an intelligent discrimination layer 2, and an interactive feedback layer 3. The layers are connected through data flow to form a complete perception-discrimination-feedback closed-loop system.

[0083] like Figure 2 As shown, the multimodal perception layer 1 includes a millimeter wave radar array 11, an infrared thermal imager 12 and a wide-angle stereo vision module 13.

[0084] The millimeter-wave radar array 11 uses a 77GHz frequency band radar with a resolution of 0.5° and a detection range of 0.2-30m. It generates high-precision point clouds through MIMO (Multiple-Input Multiple-Output) technology, with a point cloud density of ≥200 points / m 2 , used to track the human body's center of mass in real time with an accuracy of ±2cm. Preferably, four millimeter-wave radars are deployed at the four corners of the training ground, forming an array structure that covers a 30m x 30m training area, ensuring comprehensive monitoring. For example, when a soldier performs a low-posture crawl, the millimeter-wave radar array can accurately track changes in torso height, maintaining high-precision monitoring even in smoke or at night.

[0085] The infrared thermal imager 12 is equipped with an uncooled microbolometer with a resolution of 640×480 and a thermal sensitivity of ≤50mK. It is used to identify thermal signatures during tactical movements, such as changes in temperature gradients on the contact surface of cover and the distribution of hot zones in muscle activity. Preferably, the infrared thermal imager has an acquisition frequency of 60Hz to ensure effective capture of rapidly changing thermal signatures. In actual application, when a soldier performs a leap, the infrared thermal imager can display differences in the thermal radiation distribution of the left and right leg muscles—for example, the right side may be 1.2°C higher than the left—indicating an imbalance in force application, thus providing important evidence for subsequent corrections.

[0086] The wide-angle stereo vision module 13 includes a dual-camera structure, a frame rate of 120fps, a baseline distance of 20cm, and is combined with an active near-infrared fill light system to achieve a skeletal joint positioning error of less than 5mm in a low-light environment. Preferably, the wide-angle stereo vision module uses a global exposure sensor to effectively avoid motion blur problems under high-speed movements. For example, when a soldier performs a side roll (the peak angular velocity can reach 300° / s), a traditional 60fps camera is prone to motion blur, while the 120fps high frame rate camera of this system can clearly capture the details of the movement, providing a basis for accurate evaluation.

[0087] Data collected by the multimodal perception layer 1 is transmitted to the intelligent judgment layer 2 via 10Gbps industrial Ethernet. The IEEE 1588 precision time protocol is used to ensure data timestamp synchronization with an accuracy of <1ms, resolving the asynchrony of multi-source data in traditional systems. In practical applications, this high-precision time synchronization is particularly critical for evaluating continuous movements such as "leap forward and then lie down," enabling precise analysis of the coordination of transitions.

[0088] like Figure 3 As shown, the intelligent discrimination layer 2 includes a heterogeneous data fusion module 21, a tactical-specific skeleton recognition module 22, an action matching module 23, a personalized standard generation module 24 and a priority evaluation module 25.

[0089] like Figure 4 As shown, the heterogeneous data fusion module 21 includes a multi-source data preprocessing unit 211 , a time synchronization unit 212 , a spatial alignment unit 213 and a multimodal feature fusion unit 214 .

[0090] The multi-source data preprocessing unit 211 designs independent preprocessing channels for different sensor data: voxel downsampling (grid size 3cm) and statistical outlier filtering (neighborhood radius 10cm, standard deviation threshold 2.0) are applied to radar point clouds; Gaussian smoothing (σ = 1.2) and adaptive binarization (window size 15×15) are applied to infrared thermal images; and the contrast-limited adaptive histogram equalization (CLAHE) algorithm is applied to RGB images to enhance texture details in low-light environments. For example, when soldiers conduct tactical maneuver training in morning fog, the RGB images often have low contrast. After processing with the CLAHE algorithm, the contrast can be increased from an initial 0.4 to 0.75, significantly improving the accuracy of subsequent skeleton recognition.

[0091] The time synchronization unit 212 designs a data alignment mechanism within a 50ms time window and uses a bidirectional long short-term memory network (LSTM, number of hidden units = 256) to process the temporal relationship of sensor data with different sampling rates. The forward propagation process of the LSTM network can be expressed as:

[0092] f t =σ(W f ·[h t-1 ,x t ]+b f ),

[0093] i t =σ(W i ·[h t-1 ,x t ]+b i ),

[0094]

[0095] o t =σ(W o ·[h t-1 ,x t ]+b o ),

[0096] h t =o t *tanh(C t ),

[0097] Where: f t is the forget gate control vector, which determines whether to retain or discard cell state information, and its value range is [0,1]; i t The input gate control vector determines which information to update, and its value range is [0,1]; is the candidate value of the new state, and its value range is [-1,1]; C t It is the cell state, which stores long-term memory information; tis the output gate control vector, with a value range of [0,1]; h t is the hidden state output, which is the network output of the current time step; W f 、W i 、W C 、W o is the weight matrix, corresponding to the weight parameters of the forget gate, input gate, candidate state and output gate respectively; b f 、b i 、b C 、b o is the bias vector, corresponding to the bias parameters of each gate; σ is the sigmoid activation function, which maps the input to the [0,1] interval; tanh is the hyperbolic tangent activation function, which maps the input to the [-1,1] interval; [h t-1 ,x t ] means to change the hidden state h of the previous time step t-1 With the current input x t Concatenate on the feature dimension; * indicates element-wise multiplication operation.

[0098] The LSTM network's timing processing capabilities are particularly important in tactical maneuver assessment. For example, when evaluating the "leap-and-lie" continuous maneuver, the data sampling rates of millimeter-wave radar (100Hz sampling rate), infrared thermal imaging (60Hz sampling rate), and RGB image (120Hz sampling rate) are inconsistent. The time synchronization unit, leveraging the memory properties of the LSTM network, effectively handles this timing inconsistency, ensuring the consistency and accuracy of the maneuver assessment.

[0099] The spatial alignment unit 213 designs a spatial transformer network (STN) to automatically learn the spatial mapping relationship between different sensor data. STN achieves spatial alignment by predicting the affine transformation matrix θ:

[0100]

[0101] For the input coordinates (x i ,y i ), the transformed coordinates (x o ,y o ) is calculated as follows:

[0102]

[0103] Where: θ is a 2×3 affine transformation matrix that controls spatial transformations such as scaling, rotation, and translation; θ 11 and θ 22 Controls the scaling in the x and y directions; θ 12 and θ 21 Controls rotation and shear; θ 13 and θ 23Controls translation in the x and y directions; (x i ,y i ) is the original coordinate point; (x o ,y o ) is the coordinate point after transformation.

[0104] In practice, different sensors have different coordinate systems. For example, the coordinate origin of a millimeter-wave radar is located at the center of the radar, while the coordinate origin of a stereo vision module is located at the optical center of the left camera. The spatial alignment unit maps all sensor data to a unified world coordinate system by learning the optimal transformation matrix θ, achieving spatial consistency. When a soldier performs a tactical maneuver, this technology ensures that the spatial positions of all limbs in the fused data are consistent, with an accuracy of ±3mm, laying the foundation for subsequent precise assessment.

[0105] The multimodal feature fusion unit 214 uses an attention mechanism to fuse the spatiotemporally aligned multimodal features. First, the radar feature vector (128 dimensions) and the heat map feature vector (96 dimensions) are initially fused using a Cross-Modal Transformer (8 heads). This is then deeply fused with the RGB feature vector (256 dimensions) using a graph convolutional network (GCN).

[0106] The attention calculation formula of Cross-ModalTransformer is as follows:

[0107]

[0108] Where: Q is the query matrix, the dimension is (n q ,d k ), represents the feature that needs to be compared with other features; K is the key matrix, the dimension is (n k ,d k ), represents the queried feature; V is the value matrix, the dimension is (n k ,d v ), represents the actual feature content to be extracted; QK T Represents the similarity matrix between query and key, with dimension (n q ,n k );d k is the dimension of the key, which serves as a scaling factor to prevent the gradient from disappearing; softmax is the softmax normalization function, which converts the similarity into a probability distribution; n q and n k are the sequence lengths of query and key respectively; d v The dimension of the value.

[0109] The calculation process of the multi-head attention mechanism is as follows:

[0110] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O ,

[0111] in: represents the output of the i-th attention head; are the linear transformation matrices of the query, key, and value of the i-th attention head respectively; W O is the output linear transformation matrix; Concat represents the concatenation operation on the feature dimension; h is the number of attention heads, which is set to 8 in this system.

[0112] In tactical action assessment, the value of multimodal feature fusion lies in its ability to combine the strengths of different sensors. For example, when a soldier is taking cover, radar features can provide precise spatial position information, thermal imaging features can reveal the distribution of heat signatures in exposed areas, and RGB features can provide detailed posture information. Through the Cross-Modal Transformer's attention mechanism, the system automatically learns how to optimally fuse these features in different scenarios. Actual tests have shown that during nighttime training, this fusion mechanism automatically increases the weights of radar and thermal imaging (approximately 0.45 and 0.4, respectively) while decreasing the weight of RGB features (approximately 0.15), thereby maintaining recognition accuracy.

[0113] Finally, the multimodal feature fusion unit 214 generates a spatiotemporal cube representation with a temporal resolution of 10ms and a spatial resolution of 3mm, containing the position, velocity, and acceleration information of 32 joints. This high-precision spatiotemporal cube provides a rich data foundation for subsequent motion analysis, capable of capturing subtle motion features such as muzzle movements and supporting elbow joint stability.

[0114] like Figure 5 As shown, the tactical-specific skeleton recognition module 22 includes an extended joint point structure unit 221, a multi-scenario adaptation network unit 222 and an adversarial training enhancement unit 223.

[0115] The extended joint structure unit 221 is expanded to 32 points based on the standard 17-point skeleton model, including 24 human skeleton nodes and 8 tactical equipment-related nodes. The human skeleton nodes include finely divided neck (upper neck, middle neck, lower neck), shoulder (front shoulder point, shoulder midpoint, shoulder back point) and hand (inner wrist point, tiger's mouth point) joints. Tactical equipment-related nodes include the muzzle pointing point, the buttstock fitting point, the tactical vest center point and the helmet edge point. This extended skeleton model can accurately describe the key details of tactical movements. For example, in the assessment of gun posture, the traditional model can only provide a rough arm position, while the extended model can accurately measure key parameters such as the "buttstock-shoulder" fitting distance and the "muzzle-line of sight" coordination.

[0116] In one embodiment of the present invention, the extended joint point structure follows a hierarchical tree structure, defining parent-child node relationships. For example, the parent node of the "butt joint" is the "shoulder front point," constraining the distance between the two to no more than 5 cm; the parent node of the "muzzle pointing point" is the "left wrist point," constraining the relative positions of the two when holding the gun. This hierarchical structure helps improve the stability and accuracy of skeleton recognition. For example, when a soldier is partially occluded, even if some nodes are difficult to directly identify, the system can make reasonable inferences through the parent-child node relationship to maintain the integrity of the skeleton model.

[0117] The multi-scene adaptation network unit 222 uses an improved deep residual network (ResNet) structure and a feature pyramid network (FPN) to extract multi-scale features. The residual block calculation formula of ResNet is as follows:

[0118] y l =h(x l )+F(x l ,W l ),

[0119] x l+1 =f(y l ),

[0120] Where: x l is the input feature map of the lth layer; x l+1 is the output feature map of the l+1 layer; F(x l ,W l ) is the residual function, which is composed of convolutional layer and batch normalization layer, W l is the weight parameter of this layer; h(x l ) is the identity mapping, which is a direct connection operation in this system; f is the ReLU activation function, which introduces nonlinearity; y l is the intermediate output of the residual block.

[0121] The feature pyramid network constructs a five-level feature pyramid with scales of 1 / 8, 1 / 16, 1 / 32, 1 / 64, and 1 / 128 of the original input, respectively. It achieves the fusion of features of different scales through top-down and bottom-up bidirectional information flow. This multi-scale feature extraction structure is particularly important for tactical training scenarios because soldiers may perform actions at different distances. For example, at close range (within 5m), fine-grained features are needed to identify finger positions, at medium distance (5-15m), the focus is on overall posture, and at long distance (15-30m), the main focus is on identifying large-scale action contours. Through the feature pyramid structure, the system can maintain stable recognition performance across the entire distance range.

[0122] The adversarial training enhancement unit 223 uses a generative adversarial network (GAN) to improve the accuracy of joint point recognition in a camouflage environment. GAN includes a generator G and a discriminator D, and the training objectives are:

[0123]

[0124] Where: G is the generator network, the input is random noise z, and the output is the synthetic image; D is the discriminator network, the input is the image, and the output is the probability that the image is a real image; p data (x) is the real image distribution; p z (z) is the random noise distribution, usually uniform distribution or Gaussian distribution; Indicates expected operation; min G max D Represents a game optimization process where the generator tries to minimize the objective function, while the discriminator tries to maximize it.

[0125] In tactical action recognition scenarios, camouflage environments pose a severe challenge to traditional vision algorithms. For example, the OpenPose algorithm achieves 95% accuracy in joint recognition under standard conditions, but this accuracy drops to 55% in scenes where camouflage uniforms are mixed with a vegetative background. Through adversarial training, the generator continuously creates more challenging camouflage-environmental mixed scenes (such as camouflage uniforms in the jungle, white equipment in the snow, etc.), while the discriminator learns how to accurately identify joints in these scenes. Experiments show that after training with 16,000 samples, the system's joint recognition accuracy in complex camouflage environments has increased to 92%, significantly surpassing traditional algorithms.

[0126] like Figure 6 As shown, the action matching module 23 includes a multi-dimensional feature extraction unit 231, a dynamic time warping matching unit 232 and a real-time computing pipeline unit 233.

[0127] The multidimensional feature extraction unit 231 extracts three types of features based on the extended skeleton model: geometric features, dynamic features, and synergistic features. Geometric features include 15 sets of key angle parameters (such as the "gun-body" angle and the "torso-ground" angle), 10 sets of key distance parameters (such as the "stock-shoulder" fit distance and the "jump stride"), and 5 sets of contour features. Dynamic features include 6 sets of acceleration parameters, 8 sets of angular velocity parameters, and 4 sets of impulse features. Synergistic features include 4 sets of limb coordination parameters, 3 sets of balance parameters, and 4 sets of rhythm parameters.

[0128] Preferably, the angle parameter calculation adopts the method of determining the angle by three points:

[0129]

[0130] Where: θ is a vector and The angle between them, in radians; and are two three-dimensional vectors; Represents the dot product of two vectors; and They represent the modulus of the two vectors respectively; arccos is the inverse cosine function.

[0131] Angle parameters are crucial in tactical action assessment. For example, when evaluating shooting posture, the "gun-to-body" angle should be maintained within 70°±5°. Excessive angles can affect stability. The "torso-ground" angle in a prone shooting position should be maintained within 10°±3°; excessive angles increase exposure. The system's high-precision angle calculations enable it to detect even the slightest deviation in posture and promptly correct it.

[0132] The dynamic time warping matching unit 232 calculates the similarity between the real-time action and the standard action based on the improved dynamic time warping (DTW) algorithm. The standard DTW algorithm defines a cumulative distance matrix D, where D(i, j) represents the minimum distance between the first i elements of sequence X and the first j elements of sequence Y:

[0133] D(i,j)=d(x i ,y j )+min{D(i-1,j),D(i,j-1),D(i-1,i-1)},

[0134] Where: D(i,j) is the element of the cumulative distance matrix, which represents the cumulative cost of sequence matching; d(x i ,y j ) is the element x i and y jThe distance measurement function between them is usually Euclidean distance; min{D(i-1,j),D(i,j-1),D(i-1,j-1)} means choosing the one with the lowest cost among the three possible paths, corresponding to insertion, deletion and matching operations respectively.

[0135] The improved DTW algorithm of this invention adopts a multi-level sampling mechanism, first performing global matching at a coarse granularity (50ms) to find the approximate matching area, and then performing local precise matching at a fine granularity (10ms). In addition, a feature weight adaptive mechanism is introduced to dynamically adjust the weights of different features according to the action type:

[0136]

[0137] Where: d(x i ,y j ) is the weighted distance metric function; K is the number of feature categories; w k is the weight of the k-th feature, and the sum of all weights is 1; is the distance metric function of the k-th feature; and are the k-th eigenvalues ​​at the i-th and j-th time points in sequences X and Y, respectively.

[0138] Preferably, the leaping action focuses on the stride length parameters (weight 0.3) and torso height (weight 0.25); the prone action focuses on speed control (weight 0.35) and final posture (weight 0.3); the shooting posture focuses on gun stability (weight 0.4) and posture angle (weight 0.35). This dynamic weight adjustment mechanism enables the system to accurately evaluate different types of actions. For example, when evaluating leaping actions, stride length and torso height are key indicators - a stride length that is too large will lead to a longer exposure time, and a torso height that is too high will increase the risk of being discovered. By giving these parameters higher weights, the system can prioritize correcting these key issues.

[0139] The real-time computation pipeline unit 233 uses a sliding window mechanism (main window 200ms, step size 50ms) to continuously evaluate action segments. It also simultaneously evaluates eight standard action templates via an 8-channel parallel DTW processing unit. Through matrix operations optimized by the ARM Neon instruction set, single-core processing speed is increased by 3.6 times, achieving a complete action evaluation frequency of 20 times per second. This high-frequency evaluation is particularly important for rapid tactical maneuvers. For example, the "roll-and-shoot" combination typically completes within one second. Traditional systems can only evaluate it 3-5 times per second, easily missing critical moments. This system's 20Hz evaluation frequency ensures that all key action details are captured.

[0140] like Figure 7As shown, the personalized standard generation module 24 includes a soldier feature modeling unit 241, a parameter mapping unit 242 and a dynamic threshold generation unit 243.

[0141] The soldier characteristic modeling unit 241 records and analyzes the trainee's physiological characteristics (eight parameters, including height H, weight W, and limb proportions L), ability ratings (six parameters, including strength S, endurance E, flexibility F, and coordination C, each rated 1-5), and training history data (four parameters, including number of training days T and skill mastery M). This multi-dimensional characteristic modeling is crucial for personalized training, as differences in physical conditions can significantly affect the execution of tactical actions. For example, when a soldier 1.90 meters tall performs a low-post crawl, their torso is typically 3-5 cm higher off the ground than a soldier 1.65 meters tall. Using a uniform standard for evaluation would result in unfairness for taller soldiers.

[0142] The parameter mapping unit 242 generates personalized mapping rules for the movement standard based on the individual characteristics of the trainee. Taking the leap length standard as an example, its calculation formula is:

[0143] Step size standard interval = 0.7H×(1±(0.1-0.02×S-0.01×E)),

[0144] Among them: H is height (unit: meter), which indicates the soldier's height value; S is strength level (level 1-5), which indicates the soldier's strength level; E is endurance level (level 1-5), which indicates the soldier's endurance level; 0.7H represents the ratio of basic stride length to height; (0.1-0.02×S-0.01×E) represents the floating range, which decreases as the strength and endurance levels increase, requiring higher-level soldiers to maintain more precise stride length control.

[0145] For example, for a trainee 1.75m tall, with a strength level of 3 and an endurance level of 4, the standard range for leap length is 1.75 x 0.7 x (1 ± 0.03) = 1.23 ± 0.04m, or 1.19-1.27m. In contrast, for a recruit of the same height but with a strength level of 1 and an endurance level of 2, the standard range is 1.75 x 0.7 x (1 ± 0.07) = 1.23 ± 0.09m, or 1.14-1.32m, providing a wider margin for error. This personalized standard effectively balances training challenge and rationality, improving training motivation and effectiveness.

[0146] The calculation formula for the torso height threshold is:

[0147] Trunk height threshold = 0.08H × (1 + 0.3 × (1-S / 5) × (1-F / 5)),

[0148] Where: H is height (meters); S is strength level (levels 1-5); F is flexibility level (levels 1-5); 0.08H represents the ratio of base torso height to height; (1 + 0.3 × (1-S / 5) × (1-F / 5)) represents the adjustment factor; the lower the strength and flexibility, the more relaxed the allowable height. For example, during a crawl, for a soldier with both strength and flexibility level 5, the torso height off the ground should not exceed 0.08H; for a soldier with both strength and flexibility level 2, the threshold can be relaxed to 0.11H.

[0149] The calculation formula for the shelter utilization angle is:

[0150] Shelter utilization angle = 60°×(1-0.15×(1-C / 5)×(1-M / 5)),

[0151] Where: C is the coordination level (levels 1-5), indicating the soldier's physical coordination ability; M is the skill mastery level (levels 1-5), indicating the soldier's training proficiency; 60° is the baseline cover angle; (1-0.15 × (1-C / 5) × (1-M / 5)) is the adjustment factor. The lower the coordination and proficiency, the more relaxed the cover angle requirement. For example, for veterans with coordination and proficiency levels of 5, the cover angle should be 60°; for beginners (with coordination and proficiency levels of 2), the standard can be appropriately lowered to 54°.

[0152] Preferably, the parameter mapping unit 242 further includes an adaptive learning mechanism to dynamically adjust the mapping rule parameters based on training performance:

[0153]

[0154] in: and are the i-th mapping parameters before and after the update, respectively; η is the learning rate, with an initial value of 0.05 and a decay of 10% after every 100 training cycles, which controls the amplitude of parameter updates; performance score - expected score is the performance gap, which indicates the difference between the actual performance and the expected target; is the parameter sensitivity, which indicates the impact of parameter changes on the score.

[0155] Through this adaptive learning mechanism, the system can optimize parameter mapping rules for different soldier types. For example, if it discovers that soldiers of a certain body type perform systematically differently than expected on a specific parameter, it will automatically adjust the corresponding mapping parameters to make the standards more fair and reasonable. In actual application, after a three-month training cycle, the system made differentiated adjustments to the stride length parameters of soldiers of different height groups, significantly improving the fairness of scoring.

[0156] The dynamic threshold generation unit 243 generates four levels of evaluation criteria based on the trainee's individual characteristics and training progress: ideal value, acceptable range, warning threshold, and error threshold. Preferably, a loose standard (acceptable range: ideal value ±15%) is adopted in the initial training phase (first 20 lessons), a slightly tighter standard (acceptable range: ideal value ±10%) is adopted in the middle phase (21-60 lessons), and a stricter standard (acceptable range: ideal value ±7%) is implemented in the advanced phase (>60 lessons). This gradual adjustment of standards avoids the frustration of recruits facing overly stringent standards while ensuring that requirements are continuously improved as training progresses.

[0157] In addition, the dynamic threshold generation unit 243 also considers the training environment factors to adjust the threshold:

[0158] Adjusted threshold = basic threshold × (1 + w1 × terrain coefficient + w2 × equipment coefficient),

[0159] Among them: the basic threshold is the parameter threshold under standard training conditions; w1 and w2 are weight coefficients, determined by regression analysis, with typical values ​​of 0.25 and 0.2, respectively, which control the influence weights of different factors; the terrain coefficient is the terrain difficulty coefficient, which is 0 for flat ground, 0.3 for a 15° slope, and 0.6 for a 30° slope, indicating the influence of terrain on the difficulty of the movement; the equipment coefficient is the equipment load coefficient, which is 0 for standard equipment and increases by 0.1 for every 5 kg increase in load, indicating the influence of load on the execution of the movement.

[0160] For example, if a soldier performs a leaping exercise on a 15° slope while carrying an additional 10kg of gear, their leaping length standard range will automatically adjust from the base range of 1.19-1.27m to 1.19-1.27 × (1 + 0.25 × 0.3 + 0.2 × 0.2) = 1.19-1.36m, creating a more relaxed upper limit to accommodate the increased difficulty brought on by the terrain and the load. This environmental adaptation adjustment makes training assessments more aligned with actual combat needs.

[0161] like Figure 8 As shown, the priority evaluation module 25 includes an error severity evaluation unit 251 , a priority sorting unit 252 , an associated error merging unit 253 and a training process adjustment unit 254 .

[0162] The error severity assessment unit 251 assesses the severity of the error based on the impact of the action error on tactical effectiveness and potential injury risk. The tactical impact assessment uses fuzzy logic calculation:

[0163] Tactical impact = α × exposure risk + β × firepower effectiveness loss + γ × mobility loss,

[0164] Among them: α, β, and γ are weight coefficients, with typical values ​​of 0.5, 0.3, and 0.2, respectively. They are determined by expert experience and reflect the relative importance of each factor; exposure risk refers to the increased probability of being discovered or hit due to incorrect actions, ranging from 0 to 1; firepower effectiveness loss refers to the degree of decrease in shooting accuracy caused by incorrect posture, ranging from 0 to 1; mobility loss refers to the degree of decrease in moving speed or flexibility caused by incorrect actions, ranging from 0 to 1.

[0165] For example, excessive torso height during a leap (exceeding the standard by more than 5 cm) significantly increases exposure risk (score 0.8), but has relatively minor impacts on firepower and mobility (0.2 and 0.3, respectively). The combined tactical impact is 0.5 × 0.8 + 0.3 × 0.2 + 0.2 × 0.3 = 0.52. Meanwhile, improper elbow support during a gun-holding stance, while associated with a lower exposure risk (0.2), significantly impacts firepower effectiveness (0.9), resulting in a combined impact of 0.5 × 0.2 + 0.3 × 0.9 + 0.2 × 0.3 = 0.43. Through this quantitative assessment, the system can scientifically determine the tactical impact of different errors.

[0166] Injury risk assessment is based on biomechanical model calculations:

[0167]

[0168] Where: N is the number of joints, indicating the total number of joints to be evaluated; i Should be the stress value currently borne by the i-th joint; joint i The safety threshold is the upper limit of the safe stress of the joint, which is determined based on biomechanical data; the duration coefficient is proportional to the duration of the incorrect action, which is 1.0 within 1 second and increases by 0.2 for every additional second, reflecting the cumulative damage caused by continuous bad posture.

[0169] In practical applications, this model can effectively prevent training injuries. For example, improper knee landing during a prone position (such as a hard landing) can cause knee joint stress to approach 90% of the safety threshold. This may not cause noticeable discomfort in the short term, but it can lead to knee damage during long-term training. Upon detecting this high-risk action, the system immediately triggers a corrective instruction: "Cushion your knee before landing to avoid a hard landing," effectively reducing the risk of training injuries.

[0170] The priority sorting unit 252 prioritizes the detected errors and calculates a priority score:

[0171] Priority score = 0.4 × severity + 0.3 × difficulty of correction + 0.2 × scope of correction + 0.1 × frequency of recurrence,

[0172] Among them: all indicators are normalized to the range of 0-1; severity is output by the error severity assessment unit, indicating the degree of harm caused by the error; the correction difficulty coefficient reflects the complexity of error correction, with simple posture adjustment being 0.2 and complex coordinated movements being 0.8, affecting the priority of correction; the correction impact range indicates the degree of impact of correcting the error on the overall movement, with local adjustment being 0.3 and overall posture change being 0.9; the recurrence frequency indicates the frequency of occurrence of the error in recent training, with the first occurrence being 0.1 and frequent occurrence being 0.9, reflecting the stubbornness of the error.

[0173] For example, when the system detects two errors simultaneously: "Excessive step length" (severity 0.7, difficulty of correction 0.5, impact range 0.6, repetition frequency 0.2) and "Unstable muzzle pointing" (severity 0.8, difficulty of correction 0.3, impact range 0.4, repetition frequency 0.8), the calculated priority scores are 0.57 and 0.59, respectively. The system will prioritize correcting the muzzle pointing issue over the step length issue. This scientific prioritization mechanism ensures that trainees address the most critical issues first.

[0174] The Correlated Error Merging Unit 253 identifies multiple causally related errors and merges them into a single, comprehensive corrective instruction. For example, the system detects a causal relationship between "poor butt-shoulder contact" and "unstable aiming posture"—the root cause of unstable aiming is the incorrect butt resting against the shoulder. Instead of presenting each issue separately, the system generates a comprehensive instruction: "Adjust the butt to keep it close to the shoulder socket and maintain stable aim." This avoids the confusion caused by excessive instructions. This intelligent merging significantly improves feedback efficiency. In practice, the number of instructions has been reduced by approximately 40%, while increasing the effectiveness of the correction.

[0175] The training process adjustment unit 254 dynamically adjusts the training difficulty and content based on the trainee's performance evaluation results. After each training unit (approximately 5 minutes), a stage performance evaluation is conducted, calculating key indicators such as movement accuracy, reaction speed, coherence, and tactical effectiveness, and generating a comprehensive score (0-100 points). Difficulty adjustment uses the following formula:

[0176]

[0177] Among them: the current score is the trainee's most recent comprehensive performance score (0-100 points); the target score is the expected level at the current stage, usually set to 75 points, indicating a good level of mastery; tanh is the hyperbolic tangent function, which maps the input to the interval [-1,1] to limit the adjustment range; 0.2 is the upper limit of the adjustment coefficient, which limits the adjustment range to within ±20% to ensure training continuity.

[0178] For example, if a participant excels in rifle-shooting training, scoring 90 points, significantly higher than the target of 75, the system will calculate a difficulty adjustment coefficient of 1 + 0.2 × tanh ((90 - 75) / 20) = 1 + 0.2 × 0.61 ≈ 1.12, correspondingly increasing the difficulty by 12%—possibly by narrowing the acceptable range of angular deviation, increasing the time required to maintain a stable posture, and so on. This dynamic difficulty adjustment mechanism ensures that training remains within the appropriate challenge range, preventing participants from losing motivation due to being too easy, or from feeling frustrated due to being too difficult.

[0179] like Figure 9 As shown, the interactive feedback layer 3 includes a multi-channel voice prompt unit 31, an augmented reality tactical sandbox unit 32 and a tactile feedback unit 33.

[0180] The multi-channel voice prompt unit 31 triggers voice commands based on the severity of the error, including level 1 prompts (calm tone, indicating minor deviations, such as "Right knee bend angle less than 3°, adjustment recommended") and level 2 prompts (urgent warnings, indicating serious errors, such as "Muzzle elevation angle exceeds limit! Correct immediately to below 30°!"). This graded prompt mechanism is consistent with human cognitive characteristics. Minor errors are addressed with a calm tone to avoid tension, while serious errors are addressed with an urgent tone to attract immediate attention, thereby improving correction efficiency.

[0181] Preferably, voice prompts utilize directional sound technology to ensure only the trainee hears them, avoiding interference with other personnel. Voice synthesis delay is controlled within 100ms to ensure timely feedback. In actual training, voice prompts are particularly effective for slower-paced training exercises (such as shooting stance training), with feedback accuracy exceeding 95%. However, they are less effective for faster-paced movements (such as rapid leaps), as feedback may not keep pace with the movements.

[0182] The augmented reality tactical sandbox unit 32 overlays virtual cover and standard action paths on a head-mounted display. For example, a red dashed line indicates the ideal leap route, and a green area indicates the location of safe cover. The AR display, with a 60fps refresh rate and a 120° field of view, provides immersive visual feedback. This intuitive visual guidance is particularly suitable for spatial position training. For example, during cover utilization training, the AR display clearly marks safe and dangerous areas, helping trainees develop accurate spatial awareness.

[0183] For example, during urban combat training, an AR sandbox displays real-time blind spots in buildings, areas of obstructed vision, and optimal movement routes. This intuitive feedback allows trainees to quickly master how to utilize cover to achieve the tactical advantage of "seeing without being seen." Actual tests have shown that groups using AR-assisted training scored 27% higher in cover utilization efficiency tests than those using traditional training.

[0184] The tactile feedback unit 33 provides directional tactile cues through vibration elements on the tactical vest. For example, vibration on the left shoulder indicates a leftward adjustment, and the intensity of the vibration indicates the adjustment amount. Tactile feedback is particularly useful in noisy environments or training scenarios requiring silence. The tactile vest is equipped with 16 vibration points (4 on the chest, 4 on the back, 2 on each shoulder, and 2 on each arm) to provide precise directional information.

[0185] Tactile feedback demonstrates unique advantages in tactical silent training. For example, during nighttime infiltration training, which requires absolute silence, voice prompts are inadequate. However, tactile feedback can silently guide posture adjustments. If a shooter's center of gravity shifts to the left, a vibration point on the right side activates, prompting them to adjust to the right. The vibration intensity increases as the deviation increases, providing precise information on the adjustment range. Field tests show that after five hours of acclimatization training, soldiers can understand and respond to over 95% of tactile commands, with a reaction time of less than 1.2 seconds.

[0186] like Figure 1 As shown, the system of the present invention further includes a training data management module 4 , an environment adaptation module 5 and a system self-diagnosis module 6 .

[0187] The training data management module 4 is connected to the intelligent discrimination layer 2 and is used to store and analyze the trainee's historical training data, generating training progress reports and skill mastery assessments. This module utilizes a hierarchical data storage architecture, storing real-time data (the last 30 days) in a cache and historical data in long-term storage. Data analysis functions include learning curve modeling, skill weakness identification, and generation of personalized training recommendations.

[0188] For example, the system can identify a soldier's weak links in the leap-prone-shooting sequence. By analyzing historical data, it was found that the soldier performed well in the leap and shooting training (scores > 90 points), but performed poorly in the prone phase (scores < 70 points), especially the prone speed control and final posture stability parameters. Based on this analysis, the system automatically recommends adding specialized training for the prone action and focusing on the performance of the prone phase in comprehensive training. This targeted training suggestion can significantly improve training efficiency and accelerate the improvement of weaknesses.

[0189] The environmental adaptation module 5 is connected to the multimodal perception layer 1 and dynamically adjusts sensor parameters and processing strategies based on the current training environment conditions. For example, in strong light conditions, it reduces the weight of RGB visual processing and increases the weight of radar and thermal imaging. In rainy and foggy environments, it adjusts radar filter parameters to optimize point cloud quality. The module includes an environmental recognition unit (based on sensor self-test and environmental feature analysis) and a parameter adjustment unit (using a preset rule base to achieve adaptive adjustments).

[0190] In practical applications, the system's environmental adaptability greatly expands its use cases. For example, in heavy rain, the recognition rate of conventional vision systems drops below 30%. However, this system maintains a joint point recognition rate above 85% by automatically adjusting its strategy (lowering the visual weight to 0.15, increasing the radar weight to 0.6, and the thermal imaging weight to 0.25). Similarly, during nighttime training, the system automatically switches to an operating mode that primarily uses infrared thermal imaging, supplemented by millimeter-wave radar, to ensure all-weather training capabilities.

[0191] The system self-diagnosis module 6 is connected to the multimodal perception layer 1, the intelligent judgment layer 2, and the interactive feedback layer 3. It monitors the operating status of each system component in real time, performs regular self-calibration, and locates faults when anomalies are detected. Self-diagnosis consists of four levels: sensor status check (signal quality, noise level), data processing verification (algorithm output rationality), model performance evaluation (recognition accuracy, response latency), and overall system health status.

[0192] Optimally, the system automatically performs a complete self-check process (approximately 3 minutes) before startup each day, and a quick check (approximately 20 seconds) every 4 hours to ensure continuous and stable system operation. The self-diagnosis function significantly improves the reliability and maintainability of the system. For example, when the signal quality of a millimeter-wave radar unit degrades, the system can automatically detect and adjust the operating mode, while issuing a maintenance reminder. When environmental factors (such as extreme temperatures) may affect system performance, it will issue an early warning and recommend appropriate adjustments to the training plan. Actual tests have shown that the self-diagnosis function reduces training interruptions caused by system failures by 78%, significantly improving training continuity and equipment utilization.

[0193] like Figure 10 As shown, the present invention also provides a tactical action standardization AI correction method, comprising the following steps:

[0194] S1: Collecting the trainee's tactical movement data through a multimodal perception layer, which includes a millimeter-wave radar array, an infrared thermal imager, and a wide-angle stereo vision module;

[0195] S2: performing spatiotemporal alignment and fusion of the data collected by the millimeter-wave radar array, the infrared thermal imager, and the wide-angle stereo vision module through a heterogeneous data fusion module to generate a spatiotemporal cube representation;

[0196] S3: using a tactical skeleton recognition module to identify the positions of the trainee's skeleton joints based on the space-time cube representation, and generating an extended skeleton model including tactical nodes;

[0197] S4: extracting multi-dimensional tactical action features from the extended skeleton model through an action matching module, and performing real-time matching and comparison between the multi-dimensional tactical action features and a standard action library;

[0198] S5: Generate personalized action standard thresholds based on the trainee’s individual characteristics through a personalized standard generation module;

[0199] S6: receiving the matching result of the action matching module through the priority evaluation module, evaluating the severity of the action error based on the personalized action standard threshold, prioritizing multiple errors, and generating correction instructions;

[0200] S7: Receive the correction instruction through the interactive feedback layer, and feed back the correction instruction to the trainee in a multi-channel manner, wherein the multi-channel manner includes voice prompts, augmented reality sandbox display and tactile feedback.

[0201] In a preferred embodiment of the present invention, the spatiotemporal alignment and fusion process in step S2 includes: first, preprocessing the three modal data, denoising, enhancement and preliminary feature extraction; then performing time synchronization through the LSTM network within a 50ms time window; then achieving three-dimensional space alignment through the spatial transformation network; finally, completing multimodal feature fusion through the Cross-ModalTransformer and graph convolutional network to generate a spatiotemporal cube representation.

[0202] For example, when executing a tactical leap, the data from the three sensors is preprocessed to first resolve the inconsistent sampling rates (100Hz for millimeter-wave radar, 60Hz for infrared thermal imaging, and 120Hz for RGB images). The LSTM network uses its memory properties to interpolate or extrapolate data points within a 50ms window to generate a time-aligned data sequence. Next, the spatial transformer network maps the three modal data to a unified world coordinate system by learning the optimal transformation matrix, achieving spatial alignment. Finally, multimodal feature fusion integrates information from different feature dimensions into a unified space-time cube representation, providing a complete data foundation for subsequent analysis.

[0203] In step S3, tactical skeleton recognition is achieved using an expanded 32-point skeleton model and a deep neural network enhanced by adversarial training. This adversarial training employs a recurrent generative adversarial network to generate over 10,000 camouflage-environment hybrid scenes, driving continuous optimization of the recognition network. Ultimately, the system achieves 92% accuracy in joint point recognition against complex backgrounds.

[0204] The introduction of specialized tactical nodes is crucial to improving assessment accuracy. For example, traditional skeletal models cannot accurately assess gun grip posture. However, this system's specialized nodes, such as the "muzzle pointing point" and "stock contact point," can precisely measure the relative position of the firearm to the body. During shooting posture training, the system can calculate key parameters such as the "stock-to-shoulder" fit gap (should be <2cm) and "muzzle-to-sight" coordination (deviation should be <5°), significantly improving assessment accuracy. In practical applications, the expanded skeletal model has increased the accuracy of shooting posture assessment from 72% with traditional systems to 94%.

[0205] In step S4, motion matching first extracts three types of features (geometric, dynamic, and collaborative). Then, motion similarity is calculated using an improved DTW algorithm. Improvements include a multi-level sampling mechanism (coarse-grained 50ms, fine-grained 10ms) and adaptive feature weighting, significantly improving matching efficiency and accuracy.

[0206] The feature weighting adaptive mechanism handles different action types differently. For example, when evaluating a "tactical roll," the system focuses more on dynamic characteristics (angular velocity parameter weight 0.45, acceleration parameter weight 0.3) and synergy characteristics (rhythm parameter weight 0.15); while when evaluating a "sniper shooting" posture, it prioritizes geometric characteristics (angle parameter weight 0.4, distance parameter weight 0.3) and synergy characteristics (balance parameter weight 0.2). This dynamic weighting adjustment enables the system to accurately evaluate a variety of tactical actions, achieving an overall matching accuracy of over 90%, 21 percentage points higher than traditional fixed-weight systems.

[0207] In step S5, personalized standards are generated using a mapping function based on the trainee's height, weight, ability rating, and other parameters. For example, the formula for calculating the jump length standard interval is 0.7H×(1±(0.1-0.02×S-0.01×E)), where H is height, S is strength level, and E is endurance level.

[0208] Personalized standards have greatly improved the fairness of training evaluations. For example, if two soldiers in the same squad with large height differences (168cm and 192cm respectively) were to perform the same tactical maneuvers, using a unified standard for evaluation would be extremely unfair to the taller soldier - his stride would naturally be larger and his torso height would be higher. By generating personalized standards, the system sets a more suitable parameter range for the 192cm soldier: the leap length is 1.34±0.06m (while it is 1.18±0.05m for a 168cm soldier), and the maximum crawling torso height is 17cm (while it is 15cm for a 168cm soldier). This differentiated standard makes the evaluation more objective and fair, and stimulates training enthusiasm.

[0209] In step S6, the priority assessment first calculates the severity of the error (taking into account tactical impact and injury risk), then sorts the errors based on the priority score (0.4×severity+0.3×correction difficulty+0.2×scope of impact+0.1×repetition frequency), merges the related errors, and generates the final correction instructions.

[0210] In actual training, the value of priority assessment is reflected in the situation where multiple errors occur simultaneously. For example, a recruit encounters three problems simultaneously during the leap-and-shoot combination action: the leap length is too large (increased exposure risk), the butt position is incorrect after lying down (affecting shooting stability), and improper breathing control (affecting accuracy). The system comprehensively evaluates the severity, correction difficulty, and correlation of each error to determine the optimal correction order: first correct the butt position (the greatest tactical impact and easy to correct), then the leap length (the tactical impact is large but the correction difficulty is moderate), and finally breathing control (relatively minor and requires long-term practice). This optimized correction order improves training efficiency by approximately 35%.

[0211] In step S7, interactive feedback is provided through three channels: voice prompts (triggered in stages, with a latency of <100ms), AR sandbox display (60fps, 120° field of view), and tactile feedback (via the tactical vest's vibration element). These three feedback methods work together to ensure that correction instructions are effectively conveyed in various environmental conditions.

[0212] The advantage of multi-channel feedback lies in its adaptability to different training scenarios and individual learning characteristics. For example, in indoor precision shooting training, voice and AR visual feedback are most effective; in outdoor tactical maneuvering training, AR sandbox displays provide intuitive guidance on movement routes; and in special training at night or in conditions requiring silence, tactile feedback becomes the primary feedback channel. The system automatically adjusts the proportion of each channel used based on the training subject and environmental conditions to maximize feedback effectiveness.

[0213] Preferably, the full-link delay of the entire method (from perception to feedback) is controlled within 400ms to achieve true real-time correction. The system triggers correction within an average of 4.3 seconds after the first occurrence of the incorrect action, effectively preventing the solidification of the incorrect action. This rapid feedback mechanism is crucial to preventing the formation of incorrect action habits. In traditional training, instructors usually provide feedback only after the action is completed (the delay can be several minutes or longer). At this time, the incorrect action pattern has been performed many times and is difficult to correct. The instant feedback of this system can intervene as soon as the error is formed, greatly improving the efficiency of correction.

[0214] Through this method, the present invention achieves high-precision, real-time, and personalized assessment and correction of tactical maneuvers, effectively improving training effectiveness and efficiency. Practical application verification at an army training base showed that the pass rate for tactical maneuvers for new recruits increased from 65% to 88%, the training cycle was shortened from 120 hours to 78 hours, and the consistency coefficient (Cohen's Kappa coefficient) between the system's scoring and the manual scoring of experienced instructors increased from 0.62 to 0.89, fully demonstrating the practical value of the present invention.

[0215] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. Tactical action standardization AI correction system, characterized by: include: A multimodal perception layer, used to collect trainee tactical movement data, comprising a millimeter-wave radar array, an infrared thermal imager, and a wide-angle stereo vision module; An intelligent discrimination layer, connected to the multimodal perception layer, is configured to receive tactical action data collected by the multimodal perception layer, identify action errors in the tactical action data, and generate correction instructions. The intelligent discrimination layer includes: a heterogeneous data fusion module for performing spatiotemporal alignment and fusion of the data collected by the millimeter-wave radar array, the infrared thermal imager, and the wide-angle stereo vision module to generate a spatiotemporal cube representation; a tactical-specific skeleton recognition module, connected to the heterogeneous data fusion module, for identifying the positions of the trainee's skeleton joints based on the space-time cube representation and generating an extended skeleton model containing tactical-specific nodes; An action matching module, connected to the tactical-specific skeleton recognition module, is used to extract multi-dimensional tactical action features from the extended skeleton model and perform real-time matching and comparison between the multi-dimensional tactical action features and a standard action library; A personalized standard generation module is used to generate personalized action standard thresholds based on the individual characteristics of the trainee; a priority evaluation module, connected to the action matching module and the personalized standard generation module, configured to receive the matching results of the action matching module, evaluate the severity of action errors based on the personalized action standard threshold, prioritize multiple errors, and generate correction instructions; The interactive feedback layer is connected to the intelligent discrimination layer, and is used to receive the correction instructions generated by the intelligent discrimination layer and feed back the correction instructions to the trainee in a multi-channel manner.

2. The system according to claim 1, wherein: The heterogeneous data fusion module includes: A multi-source data preprocessing unit, configured to perform noise reduction, filtering, and feature enhancement processing on the raw data collected by the millimeter-wave radar array, the infrared thermal imager, and the wide-angle stereo vision module; a time synchronization unit, connected to the multi-source data pre-processing unit, for performing time alignment on different sensor data within a 50 millisecond time window; a spatial alignment unit, connected to the time synchronization unit, for mapping different sensor data into a unified spatial coordinate system; The multimodal feature fusion unit is connected to the spatial alignment unit and is used to fuse the multimodal features after spatiotemporal alignment through an attention mechanism to generate a spatiotemporal cube representation with a temporal resolution of 10 milliseconds and a spatial resolution of 3 millimeters.

3. The system according to claim 1, wherein: The tactical-specific skeleton recognition module includes: Extended joint structure unit, used to expand the standard 17-point skeleton model to 32 points, including: 24 human skeleton nodes, including finely divided neck, shoulder and hand joints; There are 8 points associated with tactical equipment, including the muzzle pointing point, the buttstock fitting point, the tactical vest center point, and the helmet edge point; A multi-scene adaptation network unit, connected to the extended joint structure unit, for extracting multi-scale features through a deep residual network and a feature pyramid structure; The adversarial training enhancement unit is connected to the multi-scenario adaptation network unit and is used to improve the joint point recognition accuracy in a camouflage environment by generating an adversarial network, thereby increasing the joint point recognition accuracy to 92%.

4. The system according to claim 1, wherein: The action matching module includes: A multi-dimensional feature extraction unit, configured to extract, based on the extended skeleton model: Geometric features, including key angle parameters, key distance parameters and contour features; Dynamic characteristics, including acceleration parameters, angular velocity parameters and impulse characteristics; Synergistic characteristics, including limb coordination parameters, balance parameters, and rhythm parameters; A dynamic time warping matching unit, connected to the multi-dimensional feature extraction unit, is used to calculate the similarity between the real-time action and the standard action library through a multi-level sampling mechanism and feature weight adaptation technology; A real-time computing pipeline unit is connected to the dynamic time warping matching unit and is used to achieve 20 complete action evaluations per second based on a sliding window mechanism and a parallel computing architecture.

5. The system according to claim 1, wherein: The personalized standard generation module includes: Soldier Characteristic Modeling Unit, which records and analyzes trainee physiological characteristics, competency ratings, and training history data; a parameter mapping unit, connected to the soldier characteristic modeling unit, for generating personalized mapping rules for movement standards based on the trainee's height, weight, limb proportions, and ability rating; The dynamic threshold generation unit is connected to the parameter mapping unit and is used to generate a four-level evaluation standard including an ideal value, an acceptable range, a warning threshold and an error threshold according to the individual characteristics and training progress of the trainee.

6. The system according to claim 1, wherein: The priority evaluation module includes: An error severity assessment unit, which is used to calculate the severity of an error based on its impact on tactical effectiveness and potential risk of injury; a prioritization unit, connected to the error severity evaluation unit, for prioritizing the plurality of errors detected and generating a processing queue; a correlation error merging unit, connected to the priority sorting unit, for identifying multiple errors with causal relationships and merging them into a single comprehensive correction instruction; The training process adjustment unit is connected to the associated error merging unit and is used to dynamically adjust the training difficulty and content based on the performance evaluation results of the trainee.

7. The system according to claim 1, wherein: The interactive feedback layer includes: A multi-channel voice prompt unit, used to trigger voice commands based on the severity of the error, including: Level 1 prompt, used to indicate mild deviations in a gentle tone; Level 2 prompt, used to quickly warn of serious errors; An augmented reality tactical sandbox unit is connected in parallel with the multi-channel voice prompt unit and is used to superimpose virtual cover and standard action paths through a head-mounted display to assist trainees in establishing a spatial action cognitive model; The tactile feedback unit is connected in parallel with the multi-channel voice prompt unit and the augmented reality tactical sandbox unit, and is used to provide directional tactile prompts through the vibration element on the tactical vest.

8. The system according to claim 1, wherein: The tactical action data includes: Leaping movement data, including step length dispersion, trunk height, and thermal radiation distribution of left and right leg muscles; Prone action data, including prone speed, posture stability and final concealment effect; Shooting posture data, including gun grip stability, muzzle pointing angle, and body support point distribution; Cover utilization data, including body exposure area, cover occlusion angle, and position adjustment trajectory.

9. The system according to claim 1, wherein: The system further comprises: A training data management module, connected to the intelligent discrimination layer, is used to store and analyze the trainee's historical training data and generate training progress reports and skill mastery assessments; An environmental adaptation module, connected to the multimodal perception layer, is used to dynamically adjust sensor parameters and processing strategies based on the current training environment conditions to ensure system stability under different lighting, weather and terrain conditions; The system self-diagnosis module is connected to the multimodal perception layer, the intelligent judgment layer and the interactive feedback layer, and is used to monitor the working status of each component of the system in real time, perform regular self-calibration, and locate faults when an anomaly is detected.

10. The AI ​​correction method for tactical action standardization is characterized by: include: Collecting the trainee's tactical movement data through a multimodal perception layer, which includes a millimeter-wave radar array, an infrared thermal imager, and a wide-angle stereo vision module; The data collected by the millimeter wave radar array, the infrared thermal imager and the wide-angle stereo vision module are subjected to spatiotemporal alignment and fusion by a heterogeneous data fusion module to generate a spatiotemporal cube representation; Identifying the positions of the trainee's skeleton joints based on the space-time cube representation through a tactical-specific skeleton recognition module, and generating an extended skeleton model including tactical-specific nodes; Extracting multi-dimensional tactical action features from the extended skeleton model through an action matching module, and performing real-time matching and comparison between the multi-dimensional tactical action features and a standard action library; Generate personalized action standard thresholds based on the trainee's individual characteristics through a personalized standard generation module; Receiving the matching results of the action matching module through the priority evaluation module, evaluating the severity of the action error based on the personalized action standard threshold, prioritizing multiple errors, and generating correction instructions; The correction instruction is received through the interactive feedback layer, and the correction instruction is fed back to the trainee in a multi-channel manner, wherein the multi-channel manner includes voice prompts, augmented reality sandbox display and tactile feedback.