A guide-control evaluation system and method based on multi-modal data fusion

By using multimodal data fusion technology, spatiotemporal synchronization and dynamic perception of multi-source information are achieved, solving the problem of insufficient spatiotemporal alignment accuracy between modalities, improving the objectivity and traceability of the adjudication results, providing a high-confidence decision-making basis, and ensuring the consistency of evaluation standards.

CN120998093BActive Publication Date: 2025-12-12NANJING KONGCHI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511519130.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-12-12
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

In existing technologies, multimodal data fusion in guidance and control assessment systems suffers from problems such as insufficient spatiotemporal alignment accuracy between modalities, reliance on human experience for feature extraction, and low transparency of anomaly detection logic. This makes the assessment results susceptible to subjective interference and difficult to backtrack and verify, affecting the objectivity and scientific nature of training effect evaluation.

Method used

By collecting multimodal data in real time, performing spatiotemporal alignment processing, combining a sliding window mechanism to conduct time-series analysis of dynamic features, and introducing a confidence assessment system to quantify and fuse static features, a comprehensive adjudication score is generated. By combining the backtracking of dynamic features to determine the occurrence path and nodes of abnormal behavior, a traceable adjudication evidence chain is constructed.

Benefits of technology

It significantly improves the objectivity and traceability of the adjudication results, ensures the consistency of the evaluation standards, provides a highly reliable basis for decision-making, effectively avoids interference from subjective factors, and generates a complete adjudication report.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998093B_ABST
    Figure CN120998093B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of training simulation, and particularly relates to a guide control evaluation system and method based on multi-modal data fusion. The present application constructs a multi-modal data segment synchronized in time and space, combines a sliding window mechanism to perform time series analysis on dynamic characteristics, and introduces a confidence evaluation system to quantitatively fuse static characteristics, thereby significantly improving the objectivity and traceability of the evaluation result. Moreover, the weight distribution of the characteristic parameters is corrected in real time according to the historical data benchmark value, so as to ensure that the comprehensive decision score can reflect the current behavior characteristics and maintain the consistency of the evaluation standard. In the abnormal behavior recognition link, the change path of the dynamic characteristics is traced back, and the multi-modal data is cross-validated to generate a complete evaluation report containing time node positioning, characteristic evidence chain and pattern matching result, thereby effectively avoiding subjective factor interference and providing a high-confidence decision basis for training effect evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of training and simulation technology, specifically relating to a guidance, control, and evaluation system and method based on multimodal data fusion. Background Technology

[0002] With the increasing complexity and real-time requirements of modern military and emergency drills, higher demands have been placed on the objectivity, accuracy, and traceability of drill guidance, control, and assessment. The rapid development of information technology has also spurred the widespread application of multi-source heterogeneous data in the field of guidance, control, and assessment. Specifically, multimodal data fusion technology plays a key role in integrating multi-source information such as images, videos, and sensors to achieve comprehensive dynamic perception and refined analysis of the drill process.

[0003] While some guidance and control assessment schemes based on multimodal data exist in the existing technology, they generally suffer from problems such as insufficient spatiotemporal alignment accuracy between modalities, reliance on human experience for feature extraction, and low transparency of anomaly detection logic. This makes the assessment results susceptible to subjective interference and difficult to backtrack and verify, which undoubtedly leads to a decrease in the credibility of the assessment conclusions. Consequently, it also affects the objectivity and scientific nature of the training effect evaluation and makes it difficult to meet the requirements of high-confidence assessment. Based on this, this solution proposes a guidance and control assessment method based on multimodal data fusion to solve the above problems. Summary of the Invention

[0004] The purpose of this invention is to provide a guidance and control assessment system and method based on multimodal data fusion, which can effectively improve the objectivity and traceability of behavior judgment during training and exercise, and make the assessment results more accurate and reliable.

[0005] The specific technical solution adopted by this invention is as follows:

[0006] A guidance and control assessment method based on multimodal data fusion includes:

[0007] Real-time acquisition of multimodal data during the training process, including image feedback data and video feedback data;

[0008] The image feedback data is preprocessed and combined with the time series of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminating the temporal offset and spatial distortion between different modalities, and obtaining the data to be verified.

[0009] The data to be verified is matched and compared with a preset standard template to extract the associated feature parameters used as the basis for adjudication. The associated feature parameters include dynamic features and static features.

[0010] The dynamic characteristics are analyzed over time using a sliding window mechanism to determine their changing trends, and the confidence level of the static characteristics is assessed. The dynamic characteristics are then combined with their changing trends for quantitative fusion to generate a comprehensive adjudication score.

[0011] Based on the comprehensive adjudication score, abnormal behaviors that do not comply with the behavior compliance during the training process are identified. By tracing back the dynamic characteristics, the occurrence path and nodes of abnormal behaviors are determined. Cross-validation is performed using spatiotemporally aligned multimodal data to generate a traceable adjudication evidence chain.

[0012] In a preferred embodiment, when acquiring multimodal data during the real-time training process, image feedback data is acquired by a combination of an infrared thermal imaging sensor and a visible light camera to collect the target's thermal radiation information and optical images, while video feedback data is acquired by a combination of a panoramic monitoring device and a first-person perspective wearable device to collect the target's panoramic situation and trajectory.

[0013] In a preferred embodiment, the step of combining the time series of video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminating temporal offsets and spatial distortions between different modalities, and obtaining the data to be verified includes:

[0014] Extract the timestamps of consecutive frames from the video feedback data as a reference time axis, and map the image feedback data onto the reference time axis through interpolation.

[0015] Based on the spatial coordinate information of image feedback data, the spatial position coordinates of the same target in video feedback data are aligned through feature point matching and geometric transformation;

[0016] The video frame at the same time and position as the image feedback data is extracted from the video feedback data as a reference frame. Then, the reference frame and the image feedback data are registered at the pixel level, and the registration error matrix is ​​output. Based on the error matrix, the spatial distortion of the image feedback data is corrected to complete the spatiotemporal alignment.

[0017] The spatiotemporally aligned image feedback data is fused with the video feedback data within the corresponding time window to form a spatiotemporally synchronized multimodal data segment, which is then recorded as data to be verified.

[0018] In a preferred embodiment, the step of fusing the spatiotemporally aligned image feedback data with the video feedback data within the corresponding time window to form a spatiotemporally synchronized multimodal data segment includes:

[0019] The thermal radiation intensity features in the image feedback data and the motion trajectory features in the video feedback data are weighted and fused to generate a multi-dimensional state vector of the target.

[0020] The dynamic contour features of the target are extracted from the video data, and the dynamic contour features are superimposed with the thermal radiation distribution features to generate comprehensive characterization information including the target shape and thermal feature distribution.

[0021] Using the acquisition time of the image feedback data as a reference point, the video feedback data is traced back to obtain the traceback time window;

[0022] Acquire video frame sequences within the retrospective time window and identify target locations within the video frame sequences;

[0023] The target's motion direction is determined by the changes in the target's position in the video frame sequence. When the motion direction remains unchanged and the target position in the image feedback data is approached, the video frame sequence is sampled at fixed time intervals to form a spatiotemporally synchronized multimodal data segment.

[0024] When the direction of motion changes or is not oriented toward the target position in the image feedback data, the target position before the last change of direction is used as the reference to extract the video frame sequence up to the current moment, and the sampling interval is set according to the preset minimum number of output frames, and the spatiotemporally synchronized multimodal data segments are output.

[0025] In a preferred embodiment, the step of matching and comparing the data to be verified with a preset standard template to extract the relevant feature parameters used as the basis for adjudication includes:

[0026] Obtain the multi-dimensional state vector from the data to be verified, and the baseline state vector from the standard template;

[0027] The similarity between the data to be verified and the baseline state vector of the standard template under the same target type is calculated to obtain the state matching degree.

[0028] The state matching degree is compared with a preset matching threshold;

[0029] If the state matching degree is greater than or equal to the preset threshold, it is determined that the data to be verified matches the standard template successfully, and the dynamic contour change rate, thermal radiation intensity gradient and motion trajectory continuity in the data to be verified are extracted as dynamic features, and the target existence duration, multimodal data consistency and environmental interference degree in the data to be verified are extracted as static features.

[0030] If the state matching degree is less than the preset threshold, it is determined that the data to be verified fails to match the standard template, and a secondary verification mechanism is initiated to retrieve the multimodal data of the same target that has been successfully matched in the past as an auxiliary reference.

[0031] Based on the spatiotemporal continuity constraint, the motion trajectory of the current data to be verified is interpolated and completed. The reasonable extension path of the current trajectory is predicted by combining the motion trend in historical data, and its matching degree with the standard template is recalculated. If the matching degree after correction is still lower than the threshold, the target behavior is judged to be abnormal.

[0032] In a preferred embodiment, the steps of performing time-series analysis on dynamic features based on a sliding window mechanism to determine the changing trend of dynamic features, and assessing the confidence level of static features, include:

[0033] The time length of the sliding window is preset, and dynamic feature sequences are collected by sliding at equal time intervals;

[0034] Calculate the mean and variance of the dynamic contour change rate within each sliding window, and calculate the trend coefficient of the dynamic contour change rate based on the preset trend function.

[0035] Calculate the cumulative slope of the thermal radiation intensity gradient within the sliding window, and determine the direction of change of thermal radiation intensity based on the sign of the cumulative slope. If the slope is positive, it is determined that the thermal radiation intensity is increasing; if the slope is negative, it is determined that it is decreasing.

[0036] The continuity index is obtained by calculating the proportion of path interruption distance to the total path distance based on the motion trajectory coordinate sequence.

[0037] The trend coefficient, direction of change, and consistency index are weighted and fused to obtain the time series stability score of dynamic characteristics;

[0038] By comparing the overlap ratio of the time periods in which the target appears in thermal imaging and video, a multimodal data consistency score is obtained;

[0039] The geometric similarity of the target contour in infrared and visible light dual-mode images is measured to obtain a geometric similarity score;

[0040] The environmental interference score is obtained by calculating the proportion of background noise interference during the period when the target is present by comparing signal strength.

[0041] The consistency score, geometric similarity score, and environmental interference score of the multimodal data are normalized, and the normalized consistency score, geometric similarity score, and environmental interference score are weighted and fused to obtain the comprehensive confidence score of the static features.

[0042] In a preferred embodiment, the step of quantitatively fusing the changing trends of dynamic characteristics to generate a comprehensive adjudication score includes:

[0043] Obtain the initial weights for the time series stability score and the overall confidence score;

[0044] Obtain the historical averages of the time series stability score and the overall confidence score, and record them as the stability baseline value and the confidence baseline value, respectively;

[0045] Calculate the deviation ratio between the current time series stability score and the stability benchmark value, and record it as the first adjustment factor;

[0046] Calculate the percentage deviation between the current overall confidence score and the confidence baseline, and record it as the second adjustment factor;

[0047] The initial weights are dynamically adjusted based on the first and second adjustment factors to obtain the updated time series stability score weights and confidence score weights.

[0048] The updated stability score weights and confidence score weights are multiplied by the current time-series stability score and overall confidence score, respectively, and the result is output as the overall adjudication score.

[0049] In a preferred embodiment, the steps of identifying abnormal behaviors that do not comply with the training process based on a comprehensive adjudication score, determining the occurrence path and nodes of abnormal behaviors by tracing back dynamic characteristics, and cross-validating them using spatiotemporally aligned multimodal data to generate a traceable chain of adjudication evidence include:

[0050] Target behaviors with a comprehensive adjudication score below a preset evaluation threshold are identified as abnormal behaviors.

[0051] The stability benchmark value is shifted downward to form a risk-sensitive threshold.

[0052] Starting from the moment when abnormal behavior is triggered, the time-series stability score sequence of dynamic characteristics is traced back along the time axis to identify the moment when the time-series stability score first falls below the risk sensitivity threshold as the node where the abnormality occurs.

[0053] Extract the multimodal data segment of the current moment of the anomaly occurrence node, analyze the change in target thermal radiation intensity gradient from image feedback data, extract the motion direction offset from video feedback data, and verify the consistency of target contour through geometric similarity.

[0054] The changes in thermal radiation intensity gradient, the offset of motion direction, and the contour consistency results are spatiotemporally aligned and fused to construct an abnormal behavior feature vector.

[0055] The abnormal behavior feature vector is compared with a preset library of typical violation patterns to match and output the abnormal type. Then, based on the timestamp, the original data source is associated to generate a complete chain of adjudication evidence that includes node location, abnormal behavior feature evidence, and pattern matching results.

[0056] This invention also provides a guidance and adjudication system based on multimodal data fusion, using the aforementioned guidance and adjudication method based on multimodal data fusion, comprising:

[0057] The data acquisition module is used to collect multimodal data in real time during the training process, including image feedback data and video feedback data.

[0058] The data processing module is used to preprocess the image feedback data and combine it with the time series of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminate the time offset and spatial distortion between different modalities, and obtain the data to be verified.

[0059] The parameter extraction module is used to match and compare the data to be verified with the preset standard template to extract the associated feature parameters used as the basis for adjudication. The associated feature parameters include dynamic features and static features.

[0060] The quantitative analysis module is used to perform time-series analysis on dynamic features based on the sliding window mechanism, determine the changing trend of dynamic features, and evaluate the confidence of static features. Then, it combines the changing trend of dynamic features for quantitative fusion to generate a comprehensive adjudication score.

[0061] The anomaly identification module is used to identify abnormal behaviors that do not comply with the behavior compliance during the training and exercise process based on the comprehensive adjudication score. It determines the occurrence path and node of the abnormal behavior by tracing back the dynamic characteristics, and performs cross-validation by combining spatiotemporally aligned multimodal data to generate a traceable adjudication evidence chain.

[0062] And, an electronic device, the electronic device comprising:

[0063] At least one processor;

[0064] and a memory communicatively connected to the at least one processor;

[0065] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the aforementioned guidance and assessment method based on multimodal data fusion.

[0066] The technical effects achieved by this invention are as follows:

[0067] This invention achieves comprehensive integration and dynamic perception of multi-source information such as images, videos, and sensors during training and exercises through multi-modal data fusion technology. It effectively solves the problem of insufficient spatiotemporal alignment accuracy between modalities in traditional guidance and control adjudication schemes. By constructing spatiotemporally synchronized multi-modal data segments, combining a sliding window mechanism to perform time-series analysis of dynamic features, and introducing a confidence assessment system to quantitatively fuse static features, the objectivity and traceability of adjudication results are significantly improved. Furthermore, the weight allocation of feature parameters is adjusted in real time based on historical data benchmarks to ensure that the comprehensive adjudication score reflects current behavioral characteristics while maintaining the consistency of evaluation standards. In the abnormal behavior identification stage, by tracing the change path of dynamic features and combining cross-validation of multi-modal data, a complete adjudication report containing time node positioning, feature evidence chain, and pattern matching results is generated, effectively avoiding interference from subjective factors and providing a high-confidence decision-making basis for evaluating training and exercises effectiveness. Attached Figure Description

[0068] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0069] Figure 2 This is a schematic diagram of the system modules of the present invention;

[0070] Figure 3 This is a schematic diagram of the electronic device structure of the present invention. Detailed Implementation

[0071] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0072] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0073] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in a preferred embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that mutually excludes other embodiments.

[0074] Please see Figure 1 As shown, this invention provides a guidance and control assessment method based on multimodal data fusion, comprising:

[0075] S1. Real-time acquisition of multimodal data during the training process, including image feedback data and video feedback data;

[0076] In step S1, during combat exercises, it is necessary not only to capture the target's movement trajectory and attitude changes, but also to simultaneously collect its thermal radiation information. This is to avoid misjudgments of behavior caused by environmental obstruction or signal interference during the training process. To ensure the effectiveness of data feedback, image feedback and video feedback will be combined to form a continuous dynamic characterization of the target's behavior and improve the spatiotemporal resolution of situational awareness. Specifically, when collecting multimodal data in real time during the training process, image feedback data is collected by infrared thermal imaging sensors and visible light cameras to collect the target's thermal radiation information and optical images. Video feedback data is collected by panoramic monitoring equipment and first-person perspective wearable devices to collect the target's panoramic situation and movement trajectory. This achieves spatiotemporal alignment and complementary enhancement of multi-dimensional information, ensuring the comprehensiveness and accuracy of guidance and control assessment.

[0077] S2. Preprocess the image feedback data and combine it with the time series of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminate the time offset and spatial distortion between different modalities, and obtain the data to be verified.

[0078] In step S2, after the image feedback data is output, to ensure image quality, it undergoes denoising, enhancement, and registration to improve image clarity and contrast. Then, it is combined with the timestamp information of the video feedback data for frame-level synchronization and spatial coordinate mapping to eliminate acquisition differences between different devices, thereby obtaining spatiotemporally aligned data to be verified. This lays the foundation for subsequent data fusion and feature extraction. The step of combining the time series of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminating temporal offsets and spatial distortions between different modalities, to obtain the data to be verified includes:

[0079] Extract the timestamps of consecutive frames from the video feedback data as a reference time axis, and map the image feedback data onto the reference time axis through interpolation.

[0080] Based on the spatial coordinate information of image feedback data, the spatial position coordinates of the same target in video feedback data are aligned through feature point matching and geometric transformation;

[0081] The video frame at the same time and position as the image feedback data is extracted from the video feedback data as a reference frame. Then, the reference frame and the image feedback data are registered at the pixel level, and the registration error matrix is ​​output. Based on the error matrix, the spatial distortion of the image feedback data is corrected to complete the spatiotemporal alignment.

[0082] The spatiotemporally aligned image feedback data is fused with the video feedback data within the corresponding time window to form a spatiotemporally synchronized multimodal data segment, which is then recorded as data to be verified.

[0083] Specifically, when outputting the data to be verified, the first step is to extract the continuous frame timestamps of the video feedback data as a reference time axis. Using bilinear interpolation or cubic spline interpolation, the time coordinates of each pixel in the preprocessed image feedback data are mapped to this reference time axis, ensuring strict alignment between the image and video frames in the time dimension. Then, based on the target spatial coordinates marked in the image feedback data, the SIFT feature point detection algorithm is used to extract co-occurring feature points in the image and video frames. The RANSAC algorithm is used to filter interior points and calculate the affine transformation matrix, mapping the target position in the video frame to the image coordinate system to achieve consistent spatial alignment. Furthermore, features that are consistent with the image feedback data are selected from the video feedback data. The N frames before and after the closest timestamps are used as reference frame sequences (the value of N is dynamically set according to the actual scene, usually 3 to 5 frames, to ensure temporal proximity and motion continuity). Then, a pixel-level registration method based on optical flow is used to calculate the sub-pixel displacement field between the image feedback data and the reference frames, which generates a registration error matrix. Based on the registration error matrix, spatial distortion correction is performed on the image feedback data to eliminate geometric distortion caused by differences in sensor viewing angles. Then, the corrected image feedback data is superimposed and fused with the video frames within the corresponding time window to generate a spatiotemporally synchronized multimodal data segment containing thermal radiation intensity, motion trajectory, and morphological features, which serves as the verification data for subsequent feature extraction and adjudication analysis.

[0084] It should be noted that the steps of fusing the spatiotemporally aligned image feedback data with the video feedback data within the corresponding time window to form a spatiotemporally synchronized multimodal data segment include:

[0085] The thermal radiation intensity features in the image feedback data and the motion trajectory features in the video feedback data are weighted and fused to generate a multi-dimensional state vector of the target.

[0086] The dynamic contour features of the target are extracted from the video data, and the dynamic contour features are superimposed with the thermal radiation distribution features to generate comprehensive characterization information including the target shape and thermal feature distribution.

[0087] Using the acquisition time of the image feedback data as a reference point, the video feedback data is traced back to obtain the traceback time window;

[0088] Acquire video frame sequences within the retrospective time window and identify target locations within the video frame sequences;

[0089] The target's motion direction is determined by the changes in the target's position in the video frame sequence. When the motion direction remains unchanged and the target position in the image feedback data is approached, the video frame sequence is sampled at fixed time intervals to form a spatiotemporally synchronized multimodal data segment.

[0090] When the direction of motion changes or is not towards the target position in the image feedback data, the target position before the last change of direction is used as the reference to extract the video frame sequence up to the current moment, and the sampling interval is set according to the preset minimum number of output frames, and the spatiotemporally synchronized multimodal data segments are output.

[0091] In the above process, when outputting multimodal data segments, the thermal radiation intensity features in the image feedback data and the motion trajectory features in the video feedback data are first weighted and processed. This comprehensively considers the influence of thermal radiation intensity on the target state and the ability of the motion trajectory to reflect the target's dynamics, generating a multi-dimensional state vector that fully characterizes the target state. Simultaneously, the dynamic contour features of the target are accurately extracted from the video data. These dynamic contour features can present the morphological changes of the target during movement. Then, the dynamic contour features are superimposed and fused with the thermal radiation distribution features to generate comprehensive characterization information containing both the target's morphology and thermal feature distribution. This comprehensive characterization information can more accurately reflect the actual state of the target during training. Finally, using the acquisition time point of the image feedback data as a reference point, a backtracking operation is performed on the video feedback data to determine the backtracking time window. The size of the backtracking time window is determined according to actual needs. The system dynamically sets parameters based on the characteristics of the training scenario, collects video frame sequences within a retrospective time window, identifies target positions within the video frame sequences, and accurately determines the target's movement direction based on changes in the target position within the video frame sequences. When the movement direction remains unchanged and the target position in the image feedback data is approaching, the video frame sequence is sampled at fixed time intervals to form spatiotemporally synchronized multimodal data segments, thereby recording the target's movement and state information at this stage. If the movement direction changes or the target position in the image feedback data is not approaching, the video frame sequence up to the current moment is extracted based on the target position before the last change in direction, and the sampling interval is set according to the preset minimum number of output frames. Finally, spatiotemporally synchronized multimodal data segments are output. The purpose of this is to provide sufficient data support for subsequent target behavior analysis and ensure the accuracy and reliability of the analysis results.

[0092] S3. Match and compare the data to be verified with the preset standard template, and extract the associated feature parameters used as the basis for adjudication. The associated feature parameters include dynamic features and static features.

[0093] In step S3, after the data to be verified is output, it is compared with a standard template constructed from historical training data. Based on this comparison, key feature parameters in the data to be verified are extracted to provide data support for subsequent abnormal behavior identification. The step of matching and comparing the data to be verified with the preset standard template to extract the associated feature parameters used as the basis for adjudication includes:

[0094] Obtain the multi-dimensional state vector from the data to be verified, and the baseline state vector from the standard template;

[0095] The similarity between the data to be verified and the baseline state vector of the standard template under the same target type is calculated to obtain the state matching degree.

[0096] The state matching degree is compared with a preset matching threshold;

[0097] If the state matching degree is greater than or equal to the preset threshold, it is determined that the data to be verified matches the standard template successfully, and the dynamic contour change rate, thermal radiation intensity gradient and motion trajectory continuity in the data to be verified are extracted as dynamic features, and the target existence duration, multimodal data consistency and environmental interference degree in the data to be verified are extracted as static features.

[0098] If the state matching degree is less than the preset threshold, it is determined that the data to be verified fails to match the standard template, and a secondary verification mechanism is initiated to retrieve the multimodal data of the same target that has been successfully matched in the past as an auxiliary reference.

[0099] Based on the spatiotemporal continuity constraint, the motion trajectory of the current data to be verified is interpolated and completed. The reasonable extension path of the current trajectory is predicted by combining the motion trend in historical data, and its matching degree with the standard template is recalculated. If the matching degree after correction is still lower than the threshold, the target behavior is judged to be abnormal.

[0100] Specifically, when matching the data to be verified with the standard template, multi-dimensional state vectors are first collected from both the data to be verified and the standard template. These multi-dimensional state vectors encompass various state information of the target during training. Then, the similarity between the baseline state vectors of the data to be verified and the standard template under the same target type is calculated to obtain the state matching degree. This similarity calculation can be implemented using cosine similarity or dynamic time warping algorithms to quantify the closeness of the two in temporal evolution and spatial distribution. The state matching degree is then compared with a preset matching threshold to determine whether the data to be verified has successfully matched the standard template. The matching threshold is dynamically adjusted according to the security level requirements of the actual scenario, typically set between 0.75 and 0.92. If the state matching degree is greater than or equal to the preset threshold, the data to be verified is considered to have successfully matched the standard template. At this point, dynamic contour changes in the data to be verified are extracted. Rate, thermal radiation intensity gradient, and trajectory continuity are used as dynamic features, which reflect the dynamic changes of the target during the training process. At the same time, the target existence duration, multimodal data consistency, and environmental interference are extracted from the data to be verified as static features, which reflect the static attributes of the target during the training process. If the state matching degree is less than the preset threshold, it is determined that the data to be verified fails to match the standard template. At this time, a secondary verification mechanism is initiated, which retrieves multimodal data of similar targets that have been successfully matched in the past as an auxiliary reference. Based on the spatiotemporal continuity constraint, the trajectory of the current data to be verified is interpolated and completed. Combined with the movement trend in the historical data, the reasonable extension path of the current trajectory is predicted, and its state matching degree with the standard template is recalculated. If the state matching degree after correction is still lower than the threshold, the target behavior is determined to be abnormal, thereby accurately identifying the abnormal behavior of the target.

[0101] S4. Based on the sliding window mechanism, perform time series analysis on dynamic features to determine the changing trend of dynamic features, and conduct confidence assessment on static features. Then, combine the changing trend of dynamic features for quantitative fusion to generate a comprehensive adjudication score.

[0102] In step S4, after both dynamic and static features are output, the dynamic and static features are fused and quantified into a comprehensive adjudication score for evaluating the consistency of the target behavior. This provides an intuitive basis for subsequent identification of abnormal behavior. The steps of performing time-series analysis on the dynamic features based on a sliding window mechanism to determine the changing trend of the dynamic features, and assessing the confidence level of the static features, include:

[0103] The time length of the sliding window is preset, and dynamic feature sequences are collected by sliding at equal time intervals;

[0104] Calculate the mean and variance of the dynamic contour change rate within each sliding window, and calculate the trend coefficient of the dynamic contour change rate based on the preset trend function.

[0105] Calculate the cumulative slope of the thermal radiation intensity gradient within the sliding window, and determine the direction of change of thermal radiation intensity based on the sign of the cumulative slope. If the slope is positive, it is determined that the thermal radiation intensity is increasing; if the slope is negative, it is determined that it is decreasing.

[0106] The continuity index is obtained by calculating the proportion of path interruption distance to the total path distance based on the motion trajectory coordinate sequence.

[0107] The trend coefficient, direction of change, and consistency index are weighted and fused to obtain the time series stability score of dynamic characteristics;

[0108] By comparing the overlap ratio of the time periods in which the target appears in thermal imaging and video, a multimodal data consistency score is obtained;

[0109] The geometric similarity of the target contour in infrared and visible light dual-mode images is measured to obtain a geometric similarity score;

[0110] The environmental interference score is obtained by calculating the proportion of background noise interference during the period when the target is present by comparing signal strength.

[0111] The consistency score, geometric similarity score, and environmental interference score of the multimodal data are normalized, and the normalized consistency score, geometric similarity score, and environmental interference score are weighted and fused to obtain the comprehensive confidence score of the static features.

[0112] Specifically, after the dynamic features are output, a dynamic feature sequence is collected based on the time span of the sliding window. The time span of the sliding window needs to be dynamically adjusted based on the matching relationship between the target's motion cycle and the sensor's sampling frequency. For example, for fast-moving targets, a shorter time span is used to improve response sensitivity, while for slow-moving or stationary targets, the window time is appropriately extended to enhance feature stability. Then, the mean and variance of the dynamic contour change rate within each sliding window are calculated. The mean reflects the average degree of change of the dynamic contour within that time period, while the variance reflects the fluctuation of the change. Finally, the trend coefficient of the dynamic contour change rate is calculated based on a preset trend function. The trend coefficient can intuitively represent the direction and intensity of the dynamic contour change. The expression for the trend function is:

[0113] Simultaneously, the cumulative slope of the thermal radiation intensity gradient within the sliding window is calculated. By analyzing the sign of the cumulative slope, a positive slope indicates an upward trend in thermal radiation intensity within the window, suggesting the target may be active or heating up. A negative slope indicates a downward trend, potentially indicating cooling or reduced activity. Based on the motion trajectory coordinate sequence, the proportion of path interruption distance to the total path is calculated to obtain a coherence index. This coherence index reflects the continuity of the target's motion trajectory. The trend coefficient, direction of change, and coherence index are weighted and fused, with the weights of different indices set according to their importance in judging target behavior, resulting in a temporal stability score for dynamic characteristics. This temporal stability score comprehensively reflects the stability of the target's dynamic characteristics over a period of time. Regarding static characteristics, the overlap ratio of the target's appearance in thermal imaging and video is compared. A high overlap ratio indicates strong temporal consistency between the two modalities, resulting in a multi-modal... The static data consistency score measures the geometric similarity of the target contour in both infrared and visible light dual-mode images. A higher geometric similarity indicates better consistency of the target morphology in the two images, resulting in a geometric similarity score. The geometric similarity score is calculated using a normalized cross-correlation method to ensure the result is between 0 and 1. The interference ratio of background noise during the target's presence is calculated by comparing signal intensity; a smaller interference ratio indicates less influence of the environment on the target signal, resulting in an environmental interference score. The multimodal data consistency score, geometric similarity score, and environmental interference score are normalized to ensure they fall within the same numerical range for easier comparison and fusion. Then, the normalized multimodal data consistency score, geometric similarity score, and environmental interference score are weighted and fused. The weights are determined based on their contribution to the static feature assessment, ultimately yielding a comprehensive confidence score for the static features, thus comprehensively reflecting their reliability.

[0114] Secondly, the steps for generating a comprehensive adjudication score by quantitatively integrating the changing trends of dynamic characteristics include:

[0115] Obtain the initial weights for the time series stability score and the overall confidence score;

[0116] Obtain the historical averages of the time series stability score and the overall confidence score, and record them as the stability baseline value and the confidence baseline value, respectively;

[0117] Calculate the deviation ratio between the current time series stability score and the stability benchmark value, and record it as the first adjustment factor;

[0118] Calculate the percentage deviation between the current overall confidence score and the confidence baseline, and record it as the second adjustment factor;

[0119] The initial weights are dynamically adjusted based on the first and second adjustment factors to obtain the updated time series stability score weights and confidence score weights.

[0120] The updated stability score weights and confidence score weights are multiplied by the current time-series stability score and the overall confidence score, respectively, and the result is output as the overall adjudication score.

[0121] In the above process, after both the time series stability score and the overall confidence score are output, the quantization fusion stage begins. Specifically, the initial weights of the time series stability score and the overall confidence score are first obtained.

[0122] Initial weights are typically preset based on historical data statistics and the needs of actual application scenarios to ensure that the relative importance of the time series stability score and the overall confidence score is reflected in the initial state. Subsequently, the historical averages of the time series stability score and the overall confidence score are obtained and recorded as the stability benchmark and confidence benchmark, respectively, for use in subsequent deviation calculations. During the calculation process, the deviation ratio of the current time series stability score from the stability benchmark is first determined and recorded as the first adjustment factor, reflecting the degree of deviation of the current time series stability from the historical average. Similarly, the deviation ratio of the current overall confidence score from the confidence benchmark is calculated and recorded as the second adjustment factor, used to measure the current overall confidence score. The initial weights are dynamically adjusted based on the changes in the overall confidence level, according to the first and second adjustment factors. This process aims to adjust the weights of the time-series stability score and the overall confidence score based on the actual situation of the current data, so that the overall adjudication score can more accurately reflect the consistency of the target behavior. The specific adjustment process can be achieved through an exponential decay function or a linear weighting method, ensuring that the weights of factors that deviate significantly from the benchmark value are reduced accordingly. The adjusted weights are called the updated time-series stability score weights and confidence score weights, respectively. Then, the updated stability score weights and confidence score weights are multiplied by the current time-series stability score and the overall confidence score, respectively, and the calculation results are added together to output the overall adjudication score.

[0123] S5. Identify abnormal behaviors that do not comply with the compliance of the training process based on the comprehensive adjudication score, and determine the occurrence path and abnormal occurrence node of the abnormal behavior by backtracking the dynamic characteristics. Combine the spatiotemporally aligned multimodal data for cross-validation to generate a traceable adjudication evidence chain.

[0124] In step S5, after the comprehensive adjudication score is output, it is compared with a preset evaluation threshold. The evaluation threshold is determined based on the statistical analysis of historical behavior data and the compliance boundaries set by the training rules. Specifically, it needs to consider the compliance requirements of the actual training scenario. When the comprehensive adjudication score exceeds the evaluation threshold, the corresponding behavior is judged to be abnormal. At this time, the dynamic characteristics of this abnormal behavior are back-analyzed to determine the abnormal occurrence node. By analyzing the time series fluctuations and confidence change trends before and after the abnormal node, key deviation moments are identified. Multimodal cross-validation is performed using spatiotemporally aligned video, log, and sensor data to ensure the accuracy and traceability of the abnormal judgment. The steps of identifying abnormal behaviors that do not comply with the behavior compliance during the training process based on the comprehensive adjudication score, determining the occurrence path and abnormal occurrence node of the abnormal behavior by back-tracing the dynamic characteristics, and performing cross-validation using spatiotemporally aligned multimodal data to generate a traceable adjudication evidence chain include:

[0125] Target behaviors with a comprehensive adjudication score below a preset evaluation threshold are identified as abnormal behaviors.

[0126] The stability benchmark value is shifted downward to form a risk-sensitive threshold.

[0127] Starting from the moment when abnormal behavior is triggered, the time-series stability score sequence of dynamic characteristics is traced back along the time axis to identify the moment when the time-series stability score first falls below the risk sensitivity threshold as the node where the abnormality occurs.

[0128] Extract the multimodal data segment of the current moment of the anomaly occurrence node, analyze the change in target thermal radiation intensity gradient from image feedback data, extract the motion direction offset from video feedback data, and verify the consistency of target contour through geometric similarity.

[0129] The changes in thermal radiation intensity gradient, the offset of motion direction, and the contour consistency results are spatiotemporally aligned and fused to construct an abnormal behavior feature vector.

[0130] Compare the abnormal behavior feature vector with the preset typical violation pattern library, match and output the abnormal type, and then associate the original data source with the timestamp to generate a complete judgment evidence chain that includes node location, abnormal behavior feature evidence and pattern matching results.

[0131] Specifically, after identifying abnormal behavior, the system first extracts target behaviors with a comprehensive adjudication score below a preset threshold as abnormal behaviors. Then, the stability benchmark value is shifted downwards to form a more sensitive risk assessment threshold, enhancing the ability to capture early signs of abnormality. Generally, the risk assessment threshold is 85% to 90% of the stability benchmark value, with the specific value adjusted based on the complexity of the scene. Next, starting from the moment the abnormal behavior is triggered, the system traces back along the timeline to obtain the temporal stability score sequence of dynamic characteristics. By comparing the temporal stability score with the risk sensitivity threshold point by point, the moment when the temporal stability score first falls below the risk sensitivity threshold is identified as the anomaly occurrence node. Subsequently, multimodal data segments from the anomaly occurrence node to the current moment are extracted. The change in the target's thermal radiation intensity gradient is analyzed from the image feedback data to reflect the adjustment process of the target's surface temperature. Finally, the motion direction offset is extracted from the video feedback data to quantify the target... Anomalies in spatial location are detected, and the consistency of the target contour in infrared and visible dual-mode images is verified through geometric similarity to ensure the temporal synchronization of multimodal data. This geometric similarity verification method is consistent with the method used for multi-source data alignment mentioned above. Then, the changes in thermal radiation intensity gradient, the offset of motion direction, and the contour consistency results are spatiotemporally aligned and fused to construct anomaly behavior feature vectors containing three dimensions: time, space, and physical features. The anomaly behavior feature vectors can comprehensively characterize the dynamic characteristics of anomalies. The anomaly behavior feature vectors are compared with a preset library of typical violation patterns. The feature matching degree is calculated using cosine similarity or dynamic time warping algorithms. When the matching degree exceeds a preset threshold, the anomaly type is matched and output. Then, the original data source is associated with the timestamp to generate a complete judgment evidence chain containing node location, heterogeneous behavior feature evidence, and pattern matching results, realizing full-process traceability from data collection and feature extraction to behavior judgment.

[0132] Please see Figure 2 A guidance and adjudication system based on multimodal data fusion, using the aforementioned guidance and adjudication method based on multimodal data fusion, includes:

[0133] The data acquisition module is used to collect multimodal data in real time during the training process. The multimodal data includes image feedback data and video feedback data.

[0134] The data processing module is used to preprocess the image feedback data and combine it with the time series of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminate the time offset and spatial distortion between different modalities, and obtain the data to be verified.

[0135] The parameter extraction module is used to match and compare the data to be verified with the preset standard template to extract the associated feature parameters used as the basis for adjudication. The associated feature parameters include dynamic features and static features.

[0136] The quantitative analysis module is used to perform time-series analysis on dynamic features based on the sliding window mechanism, determine the changing trend of dynamic features, and evaluate the confidence of static features. Then, it combines the changing trend of dynamic features for quantitative fusion to generate a comprehensive adjudication score.

[0137] The anomaly identification module is used to identify abnormal behaviors that do not comply with the behavior compliance during the training and exercise process based on the comprehensive adjudication score. It determines the occurrence path and node of the abnormal behavior by tracing back the dynamic characteristics, and performs cross-validation by combining spatiotemporally aligned multimodal data to generate a traceable adjudication evidence chain.

[0138] The execution process of the aforementioned guidance and assessment system corresponds exactly to the execution process of the guidance and assessment method based on multimodal data fusion, so it will not be repeated here.

[0139] Please see Figure 3 An electronic device, comprising:

[0140] At least one processor;

[0141] and memory that is communicatively connected to at least one processor;

[0142] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that at least one processor can execute the above-mentioned guidance and assessment method based on multimodal data fusion.

[0143] The processor of the aforementioned electronic device can be a chip with data processing capabilities, such as a central processing unit (CPU), a graphics processing unit (GPU), or a neural network processor (NPU). The memory can be a storage medium such as random access memory (RAM), read-only memory (ROM), or solid-state drive (SSD). In addition, the electronic device may also include an arithmetic logic unit, such as an arithmetic logic unit (ALU) or a floating-point unit (FPU), as well as input devices and output devices. The input devices can be such as a touch screen, keyboard, or mouse, used to receive user operation instructions, while the output devices can be such as a monitor or printer, used to display or output processing results.

[0144] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention are implemented according to conventional methods in the art unless otherwise specified or limited.

Claims

1. A guidance and control assessment method based on multimodal data fusion, characterized in that: include: Real-time acquisition of multimodal data during the training process, including image feedback data and video feedback data; The image feedback data is preprocessed and combined with the time series of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminating the temporal offset and spatial distortion between different modalities, and obtaining the data to be verified. The data to be verified is matched and compared with a preset standard template to extract the associated feature parameters used as the basis for adjudication. The associated feature parameters include dynamic features and static features. The dynamic characteristics are analyzed over time using a sliding window mechanism to determine their changing trends, and the confidence level of the static characteristics is assessed. The dynamic characteristics are then combined with their changing trends for quantitative fusion to generate a comprehensive adjudication score. Based on the comprehensive adjudication score, abnormal behaviors that do not comply with the behavior compliance during the training process are identified. By tracing back the dynamic characteristics, the occurrence path and nodes of abnormal behaviors are determined. Cross-validation is performed using spatiotemporally aligned multimodal data to generate a traceable adjudication evidence chain.

2. The guidance and assessment method based on multimodal data fusion according to claim 1, characterized in that: When collecting multimodal data in real time during the training process, the image feedback data is collected by the infrared thermal imaging sensor and the visible light camera to collect the target's thermal radiation information and optical images, and the video feedback data is collected by the panoramic monitoring equipment and the first-person perspective wearable device to collect the target's panoramic situation and running trajectory.

3. The guidance and assessment method based on multimodal data fusion according to claim 1, characterized in that: The step of combining the time series of video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminating temporal offsets and spatial distortions between different modalities, and obtaining the data to be verified includes: Extract the timestamps of consecutive frames from the video feedback data as a reference time axis, and map the image feedback data onto the reference time axis through interpolation. Based on the spatial coordinate information of image feedback data, the spatial position coordinates of the same target in video feedback data are aligned through feature point matching and geometric transformation; The video frame at the same time and position as the image feedback data is extracted from the video feedback data as a reference frame. Then, the reference frame and the image feedback data are registered at the pixel level, and the registration error matrix is ​​output. Based on the error matrix, the spatial distortion of the image feedback data is corrected to complete the spatiotemporal alignment. The spatiotemporally aligned image feedback data is fused with the video feedback data within the corresponding time window to form a spatiotemporally synchronized multimodal data segment, which is then recorded as data to be verified.

4. The guidance and assessment method based on multimodal data fusion according to claim 3, characterized in that: The step of fusing the spatiotemporally aligned image feedback data with the video feedback data within the corresponding time window to form a spatiotemporally synchronized multimodal data segment includes: The thermal radiation intensity features in the image feedback data and the motion trajectory features in the video feedback data are weighted and fused to generate a multi-dimensional state vector of the target. The dynamic contour features of the target are extracted from the video data, and the dynamic contour features are superimposed with the thermal radiation distribution features to generate comprehensive characterization information including the target shape and thermal feature distribution. Using the acquisition time of the image feedback data as a reference point, the video feedback data is traced back to obtain the traceback time window; Acquire video frame sequences within the retrospective time window and identify target locations within the video frame sequences; The target's motion direction is determined by the changes in the target's position in the video frame sequence. When the motion direction remains unchanged and the target position in the image feedback data is approached, the video frame sequence is sampled at fixed time intervals to form a spatiotemporally synchronized multimodal data segment. When the direction of motion changes or is not oriented toward the target position in the image feedback data, the target position before the last change of direction is used as the reference to extract the video frame sequence up to the current moment, and the sampling interval is set according to the preset minimum number of output frames, and the spatiotemporally synchronized multimodal data segments are output.

5. The guidance and assessment method based on multimodal data fusion according to claim 1, characterized in that: The step of matching and comparing the data to be verified with a preset standard template to extract the associated feature parameters used as the basis for adjudication includes: Obtain the multi-dimensional state vector from the data to be verified, and the baseline state vector from the standard template; The similarity between the data to be verified and the baseline state vector of the standard template under the same target type is calculated to obtain the state matching degree. The state matching degree is compared with a preset matching threshold; If the state matching degree is greater than or equal to the preset threshold, it is determined that the data to be verified matches the standard template successfully, and the dynamic contour change rate, thermal radiation intensity gradient and motion trajectory continuity in the data to be verified are extracted as dynamic features, and the target existence duration, multimodal data consistency and environmental interference degree in the data to be verified are extracted as static features. If the state matching degree is less than the preset threshold, it is determined that the data to be verified fails to match the standard template, and a secondary verification mechanism is initiated to retrieve the multimodal data of the same target that has been successfully matched in the past as an auxiliary reference. Based on the spatiotemporal continuity constraint, the motion trajectory of the current data to be verified is interpolated and completed. The reasonable extension path of the current trajectory is predicted by combining the motion trend in historical data, and its matching degree with the standard template is recalculated. If the matching degree after correction is still lower than the threshold, the target behavior is judged to be abnormal.

6. The guidance and assessment method based on multimodal data fusion according to claim 5, characterized in that: The steps of performing time-series analysis on dynamic features based on the sliding window mechanism to determine the changing trend of dynamic features, and assessing the confidence level of static features, include: The time length of the sliding window is preset, and dynamic feature sequences are collected by sliding at equal time intervals; Calculate the mean and variance of the dynamic contour change rate within each sliding window, and calculate the trend coefficient of the dynamic contour change rate based on the preset trend function. Calculate the cumulative slope of the thermal radiation intensity gradient within the sliding window, and determine the direction of change of thermal radiation intensity based on the sign of the cumulative slope. If the slope is positive, it is determined that the thermal radiation intensity is increasing; if the slope is negative, it is determined that it is decreasing. The continuity index is obtained by calculating the proportion of path interruption distance to the total path distance based on the motion trajectory coordinate sequence. The trend coefficient, direction of change, and consistency index are weighted and fused to obtain the time series stability score of dynamic characteristics; By comparing the overlap ratio of the time periods in which the target appears in thermal imaging and video, a multimodal data consistency score is obtained; The geometric similarity of the target contour in infrared and visible light dual-mode images is measured to obtain a geometric similarity score; The environmental interference score is obtained by calculating the proportion of background noise interference during the period when the target is present by comparing signal strength. The consistency score, geometric similarity score, and environmental interference score of the multimodal data are normalized, and the normalized consistency score, geometric similarity score, and environmental interference score are weighted and fused to obtain the comprehensive confidence score of the static features.

7. The guidance and assessment method based on multimodal data fusion according to claim 6, characterized in that: The step of quantitatively fusing the changing trends of dynamic characteristics to generate a comprehensive adjudication score includes: Obtain the initial weights for the time series stability score and the overall confidence score; Obtain the historical averages of the time series stability score and the overall confidence score, and record them as the stability baseline value and the confidence baseline value, respectively; Calculate the deviation ratio between the current time series stability score and the stability benchmark value, and record it as the first adjustment factor; Calculate the percentage deviation between the current overall confidence score and the confidence baseline, and record it as the second adjustment factor; The initial weights are dynamically adjusted based on the first and second adjustment factors to obtain the updated time series stability score weights and confidence score weights. The updated stability score weights and confidence score weights are multiplied by the current time-series stability score and overall confidence score, respectively, and the result is output as the overall adjudication score.

8. The guidance and assessment method based on multimodal data fusion according to claim 1, characterized in that: The steps of identifying abnormal behaviors that do not comply with the training process based on comprehensive adjudication scores, determining the occurrence path and nodes of abnormal behaviors by tracing back dynamic characteristics, and cross-validating them with spatiotemporally aligned multimodal data to generate a traceable chain of adjudication evidence include: Target behaviors with a comprehensive adjudication score below a preset evaluation threshold are identified as abnormal behaviors. The stability benchmark value is shifted downward to form a risk-sensitive threshold. Starting from the moment when abnormal behavior is triggered, the time-series stability score sequence of dynamic characteristics is traced back along the time axis to identify the moment when the time-series stability score first falls below the risk sensitivity threshold as the node where the abnormality occurs. Extract the multimodal data segment of the current moment of the anomaly occurrence node, analyze the change in target thermal radiation intensity gradient from image feedback data, extract the motion direction offset from video feedback data, and verify the consistency of target contour through geometric similarity. The changes in thermal radiation intensity gradient, the offset of motion direction, and the contour consistency results are spatiotemporally aligned and fused to construct an abnormal behavior feature vector. The abnormal behavior feature vector is compared with a preset library of typical violation patterns to match and output the abnormal type. Then, based on the timestamp, the original data source is associated to generate a complete chain of adjudication evidence that includes node location, abnormal behavior feature evidence, and pattern matching results.

9. A guidance and assessment system based on multimodal data fusion, characterized in that: The guidance and assessment method based on multimodal data fusion according to any one of claims 1 to 8 includes: The data acquisition module is used to collect multimodal data in real time during the training process, including image feedback data and video feedback data. The data processing module is used to preprocess the image feedback data and combine it with the time series of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminate the time offset and spatial distortion between different modalities, and obtain the data to be verified. The parameter extraction module is used to match and compare the data to be verified with the preset standard template to extract the associated feature parameters used as the basis for adjudication. The associated feature parameters include dynamic features and static features. The quantitative analysis module is used to perform time-series analysis on dynamic features based on the sliding window mechanism, determine the changing trend of dynamic features, and evaluate the confidence of static features. Then, it combines the changing trend of dynamic features for quantitative fusion to generate a comprehensive adjudication score. The anomaly identification module is used to identify abnormal behaviors that do not comply with the behavior compliance during the training and exercise process based on the comprehensive adjudication score. It determines the occurrence path and node of the abnormal behavior by tracing back the dynamic characteristics, and performs cross-validation by combining spatiotemporally aligned multimodal data to generate a traceable adjudication evidence chain.

10. An electronic device, characterized in that: The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the guidance and assessment method based on multimodal data fusion as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Behavior analysis method based on video recognition, processor and storage medium

    CN120182766A

  • Immersive VR psychological detection system and method based on multi-modal AI

    CN120436642A