Guidance and control evaluation system and method based on multi-modal data fusion
By using multimodal data fusion technology, spatiotemporal alignment and feature extraction of multi-source information are achieved, solving the problem of insufficient accuracy between modalities, improving the objectivity and traceability of the adjudication results, and providing high-confidence decision support.
Patent Information
- Application Number
- CN202511519130.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-23
AI Technical Summary
In existing technologies, multimodal data fusion in guidance and control assessment systems suffers from problems such as insufficient spatiotemporal alignment accuracy between modalities, reliance on human experience for feature extraction, and low transparency of anomaly detection logic. This makes the assessment results susceptible to subjective interference and difficult to backtrack and verify, affecting the objectivity and scientific nature of training effect evaluation.
By collecting multimodal data in real time, performing spatiotemporal alignment and preprocessing, extracting associated feature parameters, combining a sliding window mechanism for time series analysis and confidence assessment, generating a comprehensive adjudication score, and using dynamic features to trace back abnormal behavior, a traceable chain of adjudication evidence is constructed.
It significantly improves the objectivity and traceability of the adjudication results, ensures the consistency of the evaluation standards, provides a highly reliable basis for decision-making, effectively avoids interference from subjective factors, and generates a complete adjudication report.
Smart Images

Figure CN120998093A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of simulation training, and particularly relates to a guide control and evaluation system and method based on multi-modal data fusion. BACKGROUND
[0002] With the complexity of modern military and emergency drill scale and the improvement of real-time requirements, people have higher requirements for the objectivity, accuracy and traceability of the guide control and evaluation of the simulation training, and with the rapid development of information technology, the wide application of multi-source heterogeneous data in the guide control and evaluation direction is also generated, and the multi-modal data fusion technology plays a key role, which realizes the dynamic perception and fine analysis of the simulation training process by integrating multi-source information such as images, videos and sensors.
[0003] In the prior art, although there are some guide control and evaluation schemes based on multi-modal data, there are generally problems such as insufficient spatiotemporal alignment accuracy between modalities, feature extraction depending on artificial experience, and low transparency of abnormality discrimination logic, which leads to the evaluation results being easily disturbed by subjective factors and being difficult to verify, and thus undoubtedly leads to the decrease of the credibility of the evaluation conclusion, and accordingly, the objectivity and scientificity of the simulation training effect evaluation are affected, and it is difficult to meet the high-confidence evaluation requirement, based on which, the present scheme proposes a guide control and evaluation method based on multi-modal data fusion to solve the above problems. SUMMARY
[0004] The purpose of the present application is to provide a guide control and evaluation system and method based on multi-modal data fusion, which can effectively improve the objectivity and traceability of behavior discrimination in the simulation training process, and make the evaluation results more accurate and reliable.
[0005] The technical scheme adopted by the present application is as follows: A guide control and evaluation method based on multi-modal data fusion, comprising: real-time acquisition of multi-modal data in the simulation training process, the multi-modal data comprising image feedback data and video feedback data; preprocessing the image feedback data, and combining the time sequence of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminate the time offset and spatial distortion between different modalities, and obtain to-be-verified data; matching and comparing the to-be-verified data with a preset standard template, and extracting associated feature parameters for decision basis, wherein the associated feature parameters comprise dynamic features and static features; performing time sequence analysis on the dynamic features based on a sliding window mechanism to determine the change trend of the dynamic features, and performing confidence evaluation on the static features, and then combining the change trend of the dynamic features to perform quantitative fusion and generate a comprehensive decision score; Abnormal behaviors not conforming to the behavior compliance in the training process are identified according to the comprehensive evaluation score, and the occurrence path and the abnormal occurrence node of the abnormal behavior are determined by backtracking the dynamic characteristics, cross verification is performed on the multi-modal data after time-space alignment, and a traceable evaluation evidence chain is generated.
[0006] In a preferred solution, when the multi-modal data in the training process is collected in real time, the image feedback data is collected by the infrared thermal imaging sensor and the visible light camera to cooperatively collect the thermal radiation information and the optical image of the target, and the video feedback data is collected by the panoramic monitoring device and the first perspective wearable device to cooperatively collect the panoramic situation and the running track of the target.
[0007] In a preferred solution, the step of performing time-space alignment on the preprocessed image feedback data in combination with the time sequence of the video feedback data to eliminate the time offset and the spatial distortion between different modalities to obtain the to-be-verified data, includes: extracting the continuous frame timestamps of the video feedback data as a reference time axis, and mapping the image feedback data to the reference time axis through interpolation processing; based on the spatial coordinate information of the image feedback data, aligning the spatial position coordinates of the same target in the video feedback data through feature point matching and geometric transformation; extracting a video frame at the same time and the same position as the picture feedback data from the video feedback data as a reference frame, and then performing pixel-level registration on the reference frame and the image feedback data to output a registration error matrix, and correcting the spatial distortion of the image feedback data according to the error matrix to complete the time-space alignment; fusing the image feedback data after time-space alignment and the video feedback data in the corresponding time window to form a time-space synchronous multi-modal data segment, and recording it as to-be-verified data.
[0008] In a preferred solution, the step of fusing the image feedback data after time-space alignment and the video feedback data in the corresponding time window to form a time-space synchronous multi-modal data segment, includes: weighting and fusing the thermal radiation intensity features in the image feedback data and the motion track features in the video feedback data to generate a multi-dimensional state vector of the target; extracting the dynamic contour features of the target from the video data, and superimposing the dynamic contour features and the thermal radiation distribution features to generate comprehensive representation information including the target morphology and the thermal feature distribution; taking the collection time point of the image feedback data as a reference point, backtracking the video feedback data to obtain a backtracking time window; collecting a video frame sequence in the backtracking time window, and identifying the target position in the video frame sequence; The motion direction of the target is determined according to the position change of the target in the sequence of video frames, and when the motion direction is not changed and approaches the position of the target in the image feedback data, the sequence of video frames is sampled at a fixed time interval to form a spatiotemporal synchronous multi-modal data segment; When the motion direction is changed or not towards the position of the target in the image feedback data, the sequence of video frames from the last position change to the current time is intercepted, and the sampling interval is set according to the preset minimum output frame number, and the spatiotemporal synchronous multi-modal data segment is output.
[0009] In a preferred scheme, the step of matching the to-be-verified data with the preset standard template to extract the associated feature parameters for decision-making includes: Obtaining a multi-dimensional state vector in the to-be-verified data and a reference state vector in the standard template; Calculating the similarity of the reference state vectors of the to-be-verified data and the standard template under the same target type to obtain a state matching degree; Comparing the state matching degree with a preset matching threshold; If the state matching degree is greater than or equal to the preset threshold, it is determined that the to-be-verified data and the standard template are matched successfully, and the dynamic contour change rate, the thermal radiation intensity gradient and the motion trajectory continuity in the to-be-verified data are extracted as dynamic features, and the target existing time length, the multi-modal data consistency and the environmental interference degree are extracted as static features; If the state matching degree is less than the preset threshold, it is determined that the to-be-verified data and the standard template are matched unsuccessfully, and a secondary verification mechanism is started to retrieve multi-modal data of the same target as auxiliary reference; Based on the spatiotemporal continuity constraint condition, the motion trajectory of the current to-be-verified data is interpolated and completed, the reasonable extension path of the current trajectory is predicted combining the motion trend in the historical data, and the matching degree with the standard template is recalculated, and if the matching degree after correction is still lower than the threshold, it is determined that the target behavior is abnormal.
[0010] In a preferred scheme, the step of performing time series analysis on the dynamic features based on the sliding window mechanism to determine the change trend of the dynamic features and performing confidence evaluation on the static features includes: The time length of the sliding window is preset, and the dynamic feature sequence is collected at equal time intervals; The mean and variance of the dynamic contour change rate in each sliding window are calculated, and the trend coefficient of the dynamic contour change rate is calculated according to a preset change trend function; a cumulative change slope of the thermal radiation intensity gradient in the sliding window is calculated, and a change direction of the thermal radiation intensity is determined according to a sign of the cumulative change slope, if the slope is positive, it is determined that the thermal radiation intensity has an upward trend, and if the slope is negative, it is determined that the thermal radiation intensity has a downward trend; a proportion of the path interruption distance in the total path is calculated based on the motion trajectory coordinate sequence, to obtain a continuity index; the trend coefficient, the change direction and the continuity index are weighted and fused to obtain a time sequence stability score of the dynamic feature; an overlap proportion of the appearance time period of the target in the thermal imaging and the video is compared to obtain a multi-modal data consistency score; a geometric similarity score is obtained by measuring the geometric similarity of the target contour in the infrared and visible light dual-mode images; an environmental interference degree score is obtained by calculating the interference proportion of the background noise in the target existing period through signal intensity comparison; the multi-modal data consistency score, the geometric similarity score and the environmental interference degree score are normalized, and the normalized multi-modal data consistency score, the geometric similarity score and the environmental interference degree score are weighted and fused to obtain a comprehensive confidence score of the static feature.
[0011] In a preferred scheme, the step of quantitatively fusing the change trend of the dynamic feature to generate a comprehensive decision score comprises: initialization weights of the time sequence stability score and the comprehensive confidence score are obtained; a history mean of the time sequence stability score and the comprehensive confidence score is obtained, and is recorded as a stability reference value and a confidence reference value respectively; a deviation proportion of the current time sequence stability score and the stability reference value is calculated, and is recorded as a first adjustment factor; a deviation proportion of the current comprehensive confidence score and the confidence reference value is calculated, and is recorded as a second adjustment factor; the initialization weights are dynamically corrected according to the first adjustment factor and the second adjustment factor to obtain updated time sequence stability score weights and confidence score weights; the updated stability score weights and the confidence score weights are multiplied by the current time sequence stability score and the comprehensive confidence score respectively, and the calculation results are output as the comprehensive decision score.
[0012] In a preferred scheme, the step of identifying abnormal behaviors that do not conform to the behavior compliance in the training process according to the comprehensive decision score, and determining the occurrence path and the abnormal occurrence node of the abnormal behavior by backtracking the dynamic feature, and cross-verifying the multi-modal data after the space-time alignment to generate a traceable decision evidence chain comprises: extracting the target behavior with a comprehensive decision score lower than a preset evaluation threshold as an abnormal behavior; downwardly shifting the stability reference value to form a risk-sensitive threshold; Taking the abnormal behavior triggering moment as the starting point, the time sequence stability score sequence of the dynamic feature is traced back along the time axis in reverse direction, and the moment when the time sequence stability score is first lower than the risk-sensitive threshold is identified as the abnormal occurrence node; Extracting the multi-modal data segment of the current moment of the abnormal occurrence node, and analyzing the target thermal radiation intensity gradient change from the image feedback data, extracting the motion direction offset from the video feedback data, and verifying the consistency of the target contour through geometric similarity; The thermal radiation intensity gradient change, the motion direction offset and the contour consistency result are spatio-temporally aligned and fused to construct an abnormal behavior feature vector; The abnormal behavior feature vector is compared with a preset typical violation mode library, and the abnormal type is matched and output, and then the original data source is associated according to the time stamp to generate a complete evaluation evidence chain including node positioning, abnormal behavior feature evidence and mode matching result.
[0013] The application also provides a guide and control evaluation system based on multi-modal data fusion, which uses the guide and control evaluation method based on multi-modal data fusion. The data acquisition module is used for real-time acquisition of multi-modal data in the process of performance training, and the multi-modal data includes image feedback data and video feedback data; The data processing module is used for pre-processing the image feedback data, and spatio-temporally aligning the pre-processed image feedback data in combination with the time sequence of the video feedback data, eliminating the time offset and spatial distortion between different modalities, and obtaining the to-be-verified data; The parameter extraction module is used for matching and comparing the to-be-verified data with a preset standard template, and extracting associated feature parameters for decision basis, wherein the associated feature parameters include dynamic features and static features; The quantitative analysis module is used for performing time sequence analysis on the dynamic features based on a sliding window mechanism, determining the change trend of the dynamic features, and performing confidence evaluation on the static features, and then performing quantitative fusion in combination with the change trend of the dynamic features to generate a comprehensive decision score; The abnormality recognition module is used for identifying abnormal behaviors that do not conform to the behavior compliance in the performance training process according to the comprehensive decision score, determining the occurrence path and abnormal occurrence node of the abnormal behavior by backtracking the dynamic features, and cross-verifying the multi-modal data after spatio-temporal alignment to generate a traceable evaluation evidence chain.
[0014] And an electronic device, the electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; The computer program is executed by the at least one processor to enable the at least one processor to perform the multi-modal data fusion-based guiding, controlling and evaluation method.
[0015] The present application achieves the following technical effects: The present application realizes comprehensive integration and dynamic perception of multi-source information such as images, videos and sensors in the process of performance training through multi-modal data fusion technology, effectively solves the problem of insufficient inter-modal space-time alignment accuracy in the traditional guiding, controlling and evaluation scheme, constructs a multi-modal data segment with space-time synchronization, combines a sliding window mechanism to perform time series analysis on dynamic characteristics, and introduces a confidence evaluation system to quantitatively fuse static characteristics, thereby significantly improving the objectivity and traceability of the evaluation results, and according to the historical data benchmark value, the weight distribution of the feature parameters is corrected in real time to ensure that the comprehensive decision score can reflect the current behavior characteristics and maintain the consistency of the evaluation standard. In the abnormal behavior recognition link, the change path of the dynamic characteristics is traced back, combined with the cross-validation of multi-modal data, a complete evaluation report containing time node positioning, feature evidence chain and pattern matching result is generated, subjective factors are effectively avoided, and a high-confidence decision basis is provided for performance training effect evaluation. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a method flowchart of the present application; Figure 2 is a system module schematic diagram of the present application; Figure 3 is an electronic device structure schematic diagram of the present application. DETAILED DESCRIPTION
[0017] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0018] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the scope of the present application, therefore the present application is not limited to the specific embodiments disclosed below.
[0019] Secondly, the "one embodiment" or "embodiments" referred to herein can include a specific feature, structure, or characteristic in at least one implementation of the application. The "in one preferred embodiment" appearing in various places in the specification are not all referring to the same embodiment, nor are they necessarily all referring to a single, preferred embodiment, or to an alternative embodiment, which can be separate, or which can be combined with other embodiments.
[0020] Referring to Figure 1 As shown in the drawings, the application provides a guide control evaluation method based on multi-modal data fusion, comprising: S1, real-time acquisition of multi-modal data in the process of the exercise, the multi-modal data comprising image feedback data and video feedback data; In the step S1, in the combat exercise, not only the motion trajectory and attitude change of the target need to be captured, but also the thermal radiation information of the target need to be synchronously collected, so as to avoid the behavior misjudgment caused by environmental shielding or signal interference in the process of the exercise. In order to ensure the effectiveness of the data feedback, the image feedback and the video feedback are combined to form the continuity and dynamic description of the target behavior, and the spatio-temporal resolution of the situation awareness is improved. When the multi-modal data in the process of the exercise is collected in real time, the thermal radiation information and the optical image of the target are collected by the infrared thermal imaging sensor and the visible light camera in cooperation, and the panoramic situation and the running track of the target are collected by the panoramic monitoring device and the first perspective wearable device in cooperation, so as to realize the spatio-temporal alignment and complementary enhancement of multi-dimensional information, and ensure the comprehensiveness and accuracy of the guide control evaluation.
[0021] S2, pre-processing the image feedback data, and combining the time sequence of the video feedback data, the pre-processed image feedback data is spatio-temporally aligned, the time offset and the spatial distortion between different modalities are eliminated, and the to-be-verified data is obtained; In the step S2, after the image feedback data is output, in order to ensure the image quality, the image feedback data is subjected to processing operations such as denoising, enhancement and registration, so as to improve the definition and contrast of the image. Then, the time stamp information of the video feedback data is combined with the video feedback data for frame-level synchronization and spatial coordinate mapping, so as to eliminate the collection difference between different devices, thereby obtaining the to-be-verified data after spatio-temporal alignment, and laying a foundation for subsequent data fusion and feature extraction. The step of combining the time sequence of the video feedback data, the pre-processed image feedback data is spatio-temporally aligned, the time offset and the spatial distortion between different modalities are eliminated, and the to-be-verified data is obtained, comprising: The continuous frame time stamp of the video feedback data is extracted as a reference time axis, and the image feedback data is mapped to the reference time axis through interpolation processing; Based on the spatial coordinate information of the image feedback data, the spatial position coordinates of the same target in the video feedback data are aligned through feature point matching and geometric transformation; extracting a video frame at the same time and the same position as the picture feedback data from the video feedback data as a reference frame, performing pixel-level registration on the reference frame and the image feedback data, outputting a registration error matrix, and correcting the spatial distortion of the image feedback data according to the error matrix to complete the space-time alignment; fusing the image feedback data after the space-time alignment and the video feedback data in the corresponding time window to form a space-time synchronous multi-modal data segment, and recording it as the to-be-verified data; Specifically, when outputting the to-be-verified data, first, the continuous frame timestamps of the video feedback data are extracted as a reference time axis, the time coordinates of each pixel point of the preprocessed image feedback data are mapped to the reference time axis by a bilinear interpolation or a cubic spline interpolation method, the image frame and the video frame are strictly aligned in the time dimension, then based on the target spatial coordinates marked in the image feedback data, the SIFT feature point detection algorithm is used to extract the co-occurring feature points in the image and the video frame, the RANSAC algorithm is used to filter the inliers and calculate the affine transformation matrix, the target position in the video frame is mapped to the image coordinate system to realize the consistency alignment of the spatial position, further, the N frames before and after the image feedback data timestamp are selected from the video feedback data as a reference frame sequence (the value of N is dynamically set according to the actual scene, usually 3 to 5 frames to ensure time proximity and motion continuity), then a pixel-level registration method based on optical flow is used to calculate the sub-pixel level displacement field between the image feedback data and the reference frame, that is, a registration error matrix is generated, and the image feedback data is corrected for spatial distortion according to the registration error matrix to eliminate the geometric distortion caused by the difference in sensor viewing angle, then the corrected image feedback data is superimposed and fused with the video frame in the corresponding time window, and a space-time synchronous multi-modal data segment containing thermal radiation intensity, motion trajectory and morphological features is generated as the to-be-verified data for subsequent feature extraction and decision analysis.
[0022] It should be noted that the step of fusing the image feedback data after the space-time alignment and the video feedback data in the corresponding time window to form a space-time synchronous multi-modal data segment includes: weighting and fusing the thermal radiation intensity feature in the image feedback data and the motion trajectory feature in the video feedback data to generate a multi-dimensional state vector of the target; extracting the dynamic contour feature of the target from the video data, and superimposing the dynamic contour feature and the thermal radiation distribution feature to generate comprehensive representation information including the target morphology and thermal feature distribution; taking the image feedback data acquisition time point as a reference point, backtracking the video feedback data to obtain a backtracking time window; collecting a video frame sequence in the backtracking time window, and identifying the target position in the video frame sequence; The target's motion direction is determined by the changes in the target's position in the video frame sequence. When the motion direction remains unchanged and the target position in the image feedback data is approached, the video frame sequence is sampled at fixed time intervals to form a spatiotemporally synchronized multimodal data segment. When the direction of motion changes or is not towards the target position in the image feedback data, the target position before the last change of direction is used as the reference to extract the video frame sequence up to the current moment, and the sampling interval is set according to the preset minimum number of output frames, and the spatiotemporally synchronized multimodal data segments are output. In the above process, when outputting multimodal data segments, the thermal radiation intensity features in the image feedback data and the motion trajectory features in the video feedback data are first weighted and processed. This comprehensively considers the influence of thermal radiation intensity on the target state and the ability of the motion trajectory to reflect the target's dynamics, generating a multi-dimensional state vector that fully characterizes the target state. Simultaneously, the dynamic contour features of the target are accurately extracted from the video data. These dynamic contour features can present the morphological changes of the target during movement. Then, the dynamic contour features are superimposed and fused with the thermal radiation distribution features to generate comprehensive characterization information containing both the target's morphology and thermal feature distribution. This comprehensive characterization information can more accurately reflect the actual state of the target during training. Finally, using the acquisition time point of the image feedback data as a reference point, a backtracking operation is performed on the video feedback data to determine the backtracking time window. The size of the backtracking time window is determined according to actual needs. The system dynamically sets parameters based on the characteristics of the training scenario, collects video frame sequences within a retrospective time window, identifies target positions within the video frame sequences, and accurately determines the target's movement direction based on changes in the target position within the video frame sequences. When the movement direction remains unchanged and the target position in the image feedback data is approaching, the video frame sequence is sampled at fixed time intervals to form spatiotemporally synchronized multimodal data segments, thereby recording the target's movement and state information at this stage. If the movement direction changes or the target position in the image feedback data is not approaching, the video frame sequence up to the current moment is extracted based on the target position before the last change in direction, and the sampling interval is set according to the preset minimum number of output frames. Finally, spatiotemporally synchronized multimodal data segments are output. The purpose of this is to provide sufficient data support for subsequent target behavior analysis and ensure the accuracy and reliability of the analysis results.
[0023] S3. Match and compare the data to be verified with the preset standard template, and extract the associated feature parameters used as the basis for adjudication. The associated feature parameters include dynamic features and static features. In the step S3, after the to-be-verified data is output, the to-be-verified data is compared with the standard template constructed by the historical training data, and key feature parameters in the to-be-verified data are extracted on the basis to provide data support for subsequent abnormal behavior identification. The matching comparison of the to-be-verified data with the preset standard template and the extraction of the associated feature parameters for the basis of the judgment include: obtaining a multi-dimensional state vector in the to-be-verified data and a reference state vector in the standard template; performing similarity calculation on the reference state vectors of the to-be-verified data and the standard template under the same target type to obtain a state matching degree; comparing the state matching degree with a preset matching threshold; if the state matching degree is greater than or equal to the preset threshold, it is determined that the to-be-verified data and the standard template are matched successfully, and a dynamic contour change rate, a thermal radiation intensity gradient and a motion trajectory continuity in the to-be-verified data are extracted as dynamic features, and a target existing time length, a multi-modal data consistency and an environmental interference degree in the to-be-verified data are extracted as static features; if the state matching degree is less than the preset threshold, it is determined that the to-be-verified data and the standard template are matched unsuccessfully, a secondary verification mechanism is started, and historical multi-modal data of the same type of target matched successfully are called as auxiliary references; based on a space-time continuity constraint condition, a motion trajectory of the current to-be-verified data is interpolated and completed, a reasonable extension path of the current trajectory is predicted in combination with a motion trend in the historical data, a matching degree of the current trajectory with the standard template is recalculated, and if the matching degree after the correction is still lower than the threshold, it is determined that the target behavior is abnormal; Specifically, in the matching of the to-be-verified data and the standard template, first, multi-dimensional state vectors in the to-be-verified data and the standard template are collected, the multi-dimensional state vectors cover various state information of the target in the training process, then, a state matching degree is obtained by calculating the similarity between the to-be-verified data and the reference state vector of the standard template under the same target type, the similarity calculation can be realized by using cosine similarity or dynamic time warping algorithm to quantify the closeness of the two in the time evolution and spatial distribution, and then the state matching degree is compared with a preset matching threshold to determine whether the to-be-verified data matches the standard template successfully, the matching threshold is dynamically adjusted according to the safety level requirement of the actual scene, and is usually set between 0.75 and 0.92, if the state matching degree is greater than or equal to the preset threshold, it is determined that the to-be-verified data matches the standard template successfully, at this time, the dynamic contour change rate, the thermal radiation intensity gradient and the motion trajectory continuity in the to-be-verified data are extracted as dynamic features, the dynamic features can reflect the dynamic change of the target in the training process, and the target existing time length, the multi-modal data consistency and the environmental interference degree in the to-be-verified data are extracted as static features, the static features can reflect the static properties of the target in the training process, if the state matching degree is less than the preset threshold, it is determined that the to-be-verified data fails to match the standard template, at this time, a secondary verification mechanism is started, the multi-modal data of the same target which has been successfully matched in history is called as auxiliary reference, the motion trajectory of the current to-be-verified data is interpolated and completed based on the spatiotemporal continuity constraint condition, the reasonable extension path of the current trajectory is predicted combined with the motion trend in the historical data, and the state matching degree with the standard template is recalculated, if the state matching degree after the correction is still lower than the threshold, it is determined that the target behavior is abnormal, so that the target abnormal behavior is accurately identified.
[0024] S4, time series analysis of the dynamic features is performed based on a sliding window mechanism to determine the change trend of the dynamic features, and the confidence of the static features is evaluated, and then the change trend of the dynamic features is quantitatively fused to generate a comprehensive decision score; In the step S4, after the dynamic features and the static features are output, the dynamic features and the static features are fused and quantified into a comprehensive decision score for evaluating the consistency of the target behavior, which provides an intuitive basis for subsequent identification of abnormal behavior, wherein the steps of performing time series analysis of the dynamic features based on a sliding window mechanism to determine the change trend of the dynamic features, and evaluating the confidence of the static features, include: The time length of the sliding window is preset, and the dynamic feature sequence is collected at equal time intervals; The mean and variance of the dynamic contour change rate in each sliding window are calculated, and the trend coefficient of the dynamic contour change rate is calculated according to a preset change trend function; a cumulative change slope of the thermal radiation intensity gradient in the sliding window is calculated, and a change direction of the thermal radiation intensity is determined according to a sign of the cumulative change slope, if the slope is positive, it is determined that the thermal radiation intensity has an upward trend, and if the slope is negative, it is determined that the thermal radiation intensity has a downward trend; a proportion of the path interruption distance in the total path is calculated based on the motion trajectory coordinate sequence, to obtain a continuity index; the trend coefficient, the change direction and the continuity index are weighted and fused to obtain a time sequence stability score of the dynamic feature; an overlap proportion of the appearance time period of the target in the thermal imaging and the video is compared to obtain a multi-modal data consistency score; a geometric similarity score is obtained by measuring the geometric similarity of the target contour in the infrared and visible light dual-mode images; an environmental interference degree score is obtained by calculating the interference proportion of the background noise in the target existing period through signal intensity comparison; the multi-modal data consistency score, the geometric similarity score and the environmental interference degree score are normalized, and the normalized multi-modal data consistency score, the geometric similarity score and the environmental interference degree score are weighted and fused to obtain a comprehensive confidence score of the static feature; Specifically, after the dynamic feature is output, a dynamic feature sequence is collected according to the time span of the sliding window. The time span of the sliding window needs to be dynamically adjusted according to the matching relationship between the target motion period and the sensor sampling frequency when setting. For example, for fast-moving targets, a shorter time span is used to improve response sensitivity, and for slow or stationary targets, the window time is appropriately lengthened to enhance feature stability. Then the mean and variance of the dynamic contour change rate in each sliding window are calculated. The mean reflects the average change degree of the dynamic contour in this time period, and the variance reflects the fluctuation of the change. Then the trend coefficient of the dynamic contour change rate is calculated according to the preset change trend function. The trend coefficient can directly indicate the trend and intensity of the dynamic contour change. The expression of the change trend function is: Meanwhile, a cumulative change slope of the thermal radiation intensity gradient in the sliding window is calculated, by analyzing the sign of the cumulative change slope, if the slope is positive, it indicates that the thermal radiation intensity is in an upward trend in the window, which means that the target may be in an active or warming state, if the slope is negative, it is determined to be a downward trend, which may indicate that the target is cooling or the activity is weakening, based on the coordinate sequence of the motion trajectory, the proportion of the path interruption distance in the total path is calculated to obtain a continuity index, which reflects the continuity of the target motion trajectory, the trend coefficient, the change direction and the continuity index are weighted and fused, the weights of different indexes are set according to the importance of the target behavior judgment, the time sequence stability score of the dynamic feature is obtained, which can comprehensively reflect the stability of the dynamic feature of the target in a period of time, in the static feature aspect, the overlap ratio of the target appearing period in the thermal imaging and the video is compared, if the overlap ratio is high, it indicates that the consistency of the two modal data in time is strong, the multi-modal data consistency score is obtained, the geometric similarity of the target contour in the infrared and visible light dual-mode image is measured, the higher the geometric similarity, the better the consistency of the target form in the two images, the geometric similarity score is obtained, the geometric similarity is calculated by using the normalized cross correlation method, to ensure that the result is between 0 and 1, the interference proportion of the background noise in the target existing period is calculated by comparing the signal intensity, the smaller the interference proportion, the smaller the influence of the environment on the target signal, the environmental interference degree score is obtained, the multi-modal data consistency score, the geometric similarity score and the environmental interference degree score are normalized to make different scores in the same numerical range, which is convenient for comparison and fusion, then the normalized multi-modal data consistency score, the geometric similarity score and the environmental interference degree score are weighted and fused, the weight is determined according to the contribution degree of the static feature evaluation, and then the comprehensive confidence score of the static feature can be finally obtained, so as to comprehensively reflect the reliability of the static feature.
[0025] Secondly, the change trend of the dynamic feature is quantitatively fused to generate a comprehensive decision score, comprising: obtaining the initialization weight of the time sequence stability score and the comprehensive confidence score; obtaining the historical mean of the time sequence stability score and the comprehensive confidence score, and recording them as the stability reference value and the confidence reference value respectively; calculating the deviation proportion of the current time sequence stability score and the stability reference value, and recording it as the first adjustment factor; calculating the deviation proportion of the current comprehensive confidence score and the confidence reference value, and recording it as the second adjustment factor; dynamically correcting the initialization weight according to the first adjustment factor and the second adjustment factor to obtain the updated time sequence stability score weight and the confidence score weight; The updated stability score weight and the updated confidence score weight are multiplied by the current time sequence stability score and the comprehensive confidence score respectively, and the calculation result is output as a comprehensive decision score; In the above, after the time sequence stability score and the comprehensive confidence score are output, a quantitative fusion stage is entered, and specifically, the initialization weights of the time sequence stability score and the comprehensive confidence score are first obtained, The initialization weights are usually preset according to historical data statistics and the needs of actual application scenarios, so as to ensure that the relative importance of the time sequence stability score and the comprehensive confidence score can be reflected in the initial state. Then, the historical average values of the time sequence stability score and the comprehensive confidence score are obtained and recorded as stability reference value and confidence reference value respectively, which are used for subsequent deviation calculation. In the calculation process, the deviation proportion of the current time sequence stability score and the stability reference value is first determined and recorded as a first adjustment factor, which reflects the deviation degree of the current time sequence stability from the historical average level. Similarly, the deviation proportion of the current comprehensive confidence score and the confidence reference value is calculated and recorded as a second adjustment factor, which is used to measure the change of the current comprehensive confidence. According to the first adjustment factor and the second adjustment factor, the initialization weights are dynamically corrected. This process aims to adjust the weights of the time sequence stability score and the comprehensive confidence score according to the actual situation of the current data, so that the comprehensive decision score can more accurately reflect the consistency of the target behavior. The specific adjustment process can be realized by an exponential decay function or a linear weighting method, so as to ensure that the weight of the factor deviating from the reference value is correspondingly reduced. The corrected weights are respectively referred to as updated time sequence stability score weight and updated confidence score weight. Then, the updated stability score weight and the updated confidence score weight are multiplied by the current time sequence stability score and the comprehensive confidence score respectively, and the calculation result is added, so as to output the comprehensive decision score.
[0026] S5, according to the comprehensive decision score, identifying abnormal behaviors that do not conform to the behavior compliance in the training process, and determining the occurrence path and the abnormal occurrence node of the abnormal behaviors by backtracking the dynamic characteristics, cross verifying the multi-modal data after the time and space alignment, and generating a traceable decision evidence chain; In the step S5, after the comprehensive decision score output, it is compared with the preset evaluation threshold, and the evaluation threshold is determined according to the historical behavior data statistical analysis and the compliance boundary set according to the training rules, and the compliance requirements of the actual training scene need to be considered, when the comprehensive decision score exceeds the evaluation threshold, it is determined that the corresponding behavior is abnormal, at this time, the dynamic characteristics of the abnormal behavior are analyzed, the abnormal occurrence node is determined, the time sequence fluctuation and confidence change trend before and after the abnormal node are analyzed, the key deviation time is identified, and the video, log and sensor data are combined with the space-time alignment for multi-modal cross verification, so that the accuracy and traceability of the abnormality determination are ensured, wherein, the abnormal behavior that does not conform to the behavior compliance in the training process is identified according to the comprehensive decision score, the occurrence path and the abnormal occurrence node of the abnormal behavior are determined by backtracking the dynamic characteristics, the multi-modal data after space-time alignment are cross verified, and the traceable evaluation evidence chain is generated, including: extract the target behavior with the comprehensive decision score lower than the preset evaluation threshold as the abnormal behavior; downwardly offset the stability reference value to form a risk sensitive threshold; taking the abnormal behavior triggering time as the starting point, the time sequence stability score sequence of the dynamic characteristics is backtracked along the time axis, and the time when the time sequence stability score first falls below the risk sensitive threshold is identified as the abnormal occurrence node; extract the multi-modal data segment of the current time of the abnormal occurrence node, and parse the target thermal radiation intensity gradient change from the image feedback data, extract the motion direction offset from the video feedback data, and verify the consistency of the target outline through geometric similarity; spatially and temporally align the thermal radiation intensity gradient change, the motion direction offset and the outline consistency result to construct an abnormal behavior feature vector; compare the abnormal behavior feature vector with the preset typical violation mode library, match and output the abnormal type, and then associate the original data source according to the time stamp to generate a complete evaluation evidence chain including node positioning, abnormal behavior feature evidence and mode matching result; Specifically, after the abnormal behavior is determined, the target behavior with a comprehensive decision score lower than a preset threshold is first extracted as the abnormal behavior, and then the stability reference value is downwardly offset to form a more sensitive risk determination threshold to enhance the ability to capture early abnormal signs. Generally, the risk determination threshold is 85% to 90% of the stability reference value, and the specific value is adjusted according to the scene complexity. Then, starting from the time when the abnormal behavior is triggered, the time sequence stability score sequence of the dynamic feature is backtracked along the time axis, and the time when the time sequence stability score first falls below the risk sensitive threshold is identified as the abnormal occurrence node by comparing the time sequence stability score with the risk sensitive threshold point by point. Subsequently, the multi-modal data segment from the abnormal occurrence node to the current time is extracted, the target thermal radiation intensity gradient change is analyzed from the image feedback data to reflect the adjustment process of the target surface temperature, and then the motion direction offset is extracted from the video feedback data to quantify the abnormal offset of the target spatial position. The geometric similarity verification is performed on the consistency of the target outline in the infrared and visible light dual-mode images to ensure the time synchronization of the multi-modal data. The geometric similarity verification method is consistent with the method for multi-source data alignment described above. Then, the thermal radiation intensity gradient change, the motion direction offset and the outline consistency result are spatio-temporally aligned and fused to construct an abnormal behavior feature vector containing three dimensions of time, space and physical characteristics. The abnormal behavior feature vector can fully depict the dynamic characteristics of the abnormal behavior. By comparing the abnormal behavior feature vector with a preset typical violation mode library, the feature matching degree is calculated by the cosine similarity or dynamic time warping algorithm. When the matching degree exceeds a preset threshold, the abnormal type is matched and output. Then, the original data source is associated according to the time stamp to generate a complete evaluation evidence chain containing node positioning, abnormal behavior feature evidence and mode matching result, realizing the whole process tracing from data acquisition, feature extraction to behavior determination.
[0027] Please refer to Figure 2 A guide and control evaluation system based on multi-modal data fusion, using the above-mentioned guide and control evaluation method based on multi-modal data fusion, comprising: A data acquisition module for real-time acquisition of multi-modal data during the training process, the multi-modal data including image feedback data and video feedback data; A data processing module for pre-processing the image feedback data and spatio-temporally aligning the pre-processed image feedback data with the time sequence of the video feedback data to eliminate the time offset and spatial distortion between different modalities to obtain the to-be-verified data; A parameter extraction module for matching and comparing the to-be-verified data with a preset standard template to extract associated feature parameters for decision basis, wherein the associated feature parameters include dynamic features and static features; The quantification analysis module is configured to perform time series analysis on the dynamic features based on a sliding window mechanism, determine a change trend of the dynamic features, and perform confidence evaluation on the static features, and then perform quantitative fusion in combination with the change trend of the dynamic features to generate a comprehensive decision score. The anomaly identification module is configured to identify abnormal behavior that does not conform to behavior compliance in the training process according to the comprehensive decision score, determine an occurrence path and an abnormal occurrence node of the abnormal behavior by backtracking the dynamic features, and perform cross verification on the multi-modal data after time and space alignment to generate a traceable decision evaluation evidence chain.
[0028] The execution process of the above-mentioned guidance and control decision evaluation system completely corresponds to the execution process of the guidance and control decision evaluation method based on multi-modal data fusion, and thus repetitive and redundant descriptions will not be repeated here.
[0029] Please refer to Figure 3 An electronic device includes: at least one processor; and a memory connected in communication with the at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the above-mentioned guidance and control decision evaluation method based on multi-modal data fusion.
[0030] The processor of the above-mentioned electronic device can be a chip with data processing capability such as a central processing unit (CPU), a graphics processing unit (GPU), or a neural network processing unit (NPU), and the memory can use a storage medium such as a random access memory (RAM), a read-only memory (ROM), or a solid state drive (SSD). In addition, the electronic device can also include an arithmetic logic unit (ALU) or a floating point unit (FPU) for operation, as well as an input device and an output device. The input device can be a touch screen, a keyboard, or a mouse, etc., for receiving user operation instructions, and the output device can be a display or a printer, etc., for displaying or outputting processing results.
[0031] The above-mentioned only is the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application. The structures, devices and operation methods not specifically described and explained in the present application, such as no special description and limitation, are implemented according to the conventional means in the art.
Claims
1. A multi-modal data fusion based guide-control evaluation method, characterized in that: include: Real-time acquisition of multimodal data during the training process, including image feedback data and video feedback data; The image feedback data is preprocessed and combined with the time series of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminating the temporal offset and spatial distortion between different modalities, and obtaining the data to be verified. The data to be verified is matched and compared with a preset standard template to extract the associated feature parameters used as the basis for adjudication. The associated feature parameters include dynamic features and static features. The dynamic characteristics are analyzed over time using a sliding window mechanism to determine their changing trends, and the confidence level of the static characteristics is assessed. The dynamic characteristics are then combined with their changing trends for quantitative fusion to generate a comprehensive adjudication score. Based on the comprehensive adjudication score, abnormal behaviors that do not comply with the behavior compliance during the training process are identified. By tracing back the dynamic characteristics, the occurrence path and nodes of abnormal behaviors are determined. Cross-validation is performed using spatiotemporally aligned multimodal data to generate a traceable adjudication evidence chain.
2. The guidance and assessment method based on multimodal data fusion according to claim 1, characterized in that: When collecting multimodal data in real time during the training process, the image feedback data is collected by the infrared thermal imaging sensor and the visible light camera to collect the target's thermal radiation information and optical images, and the video feedback data is collected by the panoramic monitoring equipment and the first-person perspective wearable device to collect the target's panoramic situation and running trajectory.
3. The guidance and assessment method based on multimodal data fusion according to claim 1, characterized in that: The step of combining the time series of video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminating temporal offsets and spatial distortions between different modalities, and obtaining the data to be verified includes: Extract the timestamps of consecutive frames from the video feedback data as a reference time axis, and map the image feedback data onto the reference time axis through interpolation. Based on the spatial coordinate information of image feedback data, the spatial position coordinates of the same target in video feedback data are aligned through feature point matching and geometric transformation; The video frame at the same time and position as the image feedback data is extracted from the video feedback data as a reference frame. Then, the reference frame and the image feedback data are registered at the pixel level, and the registration error matrix is output. Based on the error matrix, the spatial distortion of the image feedback data is corrected to complete the spatiotemporal alignment. The spatiotemporally aligned image feedback data is fused with the video feedback data within the corresponding time window to form a spatiotemporally synchronized multimodal data segment, which is then recorded as data to be verified.
4. The guidance and assessment method based on multimodal data fusion according to claim 1, characterized in that: The step of fusing the spatiotemporally aligned image feedback data with the video feedback data within the corresponding time window to form a spatiotemporally synchronized multimodal data segment includes: The thermal radiation intensity features in the image feedback data and the motion trajectory features in the video feedback data are weighted and fused to generate a multi-dimensional state vector of the target. The dynamic contour features of the target are extracted from the video data, and the dynamic contour features are superimposed with the thermal radiation distribution features to generate comprehensive characterization information including the target shape and thermal feature distribution. Using the acquisition time of the image feedback data as a reference point, the video feedback data is traced back to obtain the traceback time window; Acquire video frame sequences within the retrospective time window and identify target locations within the video frame sequences; The target's motion direction is determined by the changes in the target's position in the video frame sequence. When the motion direction remains unchanged and the target position in the image feedback data is approached, the video frame sequence is sampled at fixed time intervals to form a spatiotemporally synchronized multimodal data segment. When the direction of motion changes or is not oriented toward the target position in the image feedback data, the target position before the last change of direction is used as the reference to extract the video frame sequence up to the current moment, and the sampling interval is set according to the preset minimum number of output frames, and the spatiotemporally synchronized multimodal data segments are output.
5. The guidance and assessment method based on multimodal data fusion according to claim 1, characterized in that: The step of matching and comparing the data to be verified with a preset standard template to extract the associated feature parameters used as the basis for adjudication includes: Obtain the multi-dimensional state vector from the data to be verified, and the baseline state vector from the standard template; The similarity between the data to be verified and the baseline state vector of the standard template under the same target type is calculated to obtain the state matching degree. The state matching degree is compared with a preset matching threshold; If the state matching degree is greater than or equal to the preset threshold, it is determined that the data to be verified matches the standard template successfully, and the dynamic contour change rate, thermal radiation intensity gradient and motion trajectory continuity in the data to be verified are extracted as dynamic features, and the target existence duration, multimodal data consistency and environmental interference degree in the data to be verified are extracted as static features. If the state matching degree is less than the preset threshold, it is determined that the data to be verified fails to match the standard template, and a secondary verification mechanism is initiated to retrieve the multimodal data of the same target that has been successfully matched in the past as an auxiliary reference. Based on the spatiotemporal continuity constraint, the motion trajectory of the current data to be verified is interpolated and completed. The reasonable extension path of the current trajectory is predicted by combining the motion trend in historical data, and its matching degree with the standard template is recalculated. If the matching degree after correction is still lower than the threshold, the target behavior is judged to be abnormal.
6. The guidance and assessment method based on multimodal data fusion according to claim 5, characterized in that: The steps of performing time-series analysis on dynamic features based on the sliding window mechanism to determine the changing trend of dynamic features, and assessing the confidence level of static features, include: The time length of the sliding window is preset, and dynamic feature sequences are collected by sliding at equal time intervals; Calculate the mean and variance of the dynamic contour change rate within each sliding window, and calculate the trend coefficient of the dynamic contour change rate based on the preset trend function. Calculate the cumulative slope of the thermal radiation intensity gradient within the sliding window, and determine the direction of change of thermal radiation intensity based on the sign of the cumulative slope. If the slope is positive, it is determined that the thermal radiation intensity is increasing; if the slope is negative, it is determined that it is decreasing. The continuity index is obtained by calculating the proportion of path interruption distance to the total path distance based on the motion trajectory coordinate sequence. The trend coefficient, direction of change, and consistency index are weighted and fused to obtain the time series stability score of dynamic characteristics; By comparing the overlap ratio of the time periods in which the target appears in thermal imaging and video, a multimodal data consistency score is obtained; The geometric similarity of the target contour in infrared and visible light dual-mode images is measured to obtain a geometric similarity score; The environmental interference score is obtained by calculating the proportion of background noise interference during the period when the target is present by comparing signal strength. The consistency score, geometric similarity score, and environmental interference score of the multimodal data are normalized, and the normalized consistency score, geometric similarity score, and environmental interference score are weighted and fused to obtain the comprehensive confidence score of the static features.
7. The guidance and assessment method based on multimodal data fusion according to claim 6, characterized in that: The step of quantitatively fusing the changing trends of dynamic characteristics to generate a comprehensive adjudication score includes: Obtain the initial weights for the time series stability score and the overall confidence score; Obtain the historical averages of the time series stability score and the overall confidence score, and record them as the stability baseline value and the confidence baseline value, respectively; Calculate the deviation ratio between the current time series stability score and the stability benchmark value, and record it as the first adjustment factor; Calculate the percentage deviation between the current overall confidence score and the confidence baseline, and record it as the second adjustment factor; The initial weights are dynamically adjusted based on the first and second adjustment factors to obtain the updated time series stability score weights and confidence score weights. The updated stability score weights and confidence score weights are multiplied by the current time-series stability score and overall confidence score, respectively, and the result is output as the overall adjudication score.
8. The guidance and assessment method based on multimodal data fusion according to claim 1, characterized in that: The steps of identifying abnormal behaviors that do not comply with the training process based on comprehensive adjudication scores, determining the occurrence path and nodes of abnormal behaviors by tracing back dynamic characteristics, and cross-validating them with spatiotemporally aligned multimodal data to generate a traceable chain of adjudication evidence include: Target behaviors with a comprehensive adjudication score below a preset evaluation threshold are identified as abnormal behaviors. The stability benchmark value is shifted downward to form a risk-sensitive threshold. Starting from the moment when abnormal behavior is triggered, the time-series stability score sequence of dynamic characteristics is traced back along the time axis to identify the moment when the time-series stability score first falls below the risk sensitivity threshold as the node where the abnormality occurs. Extract the multimodal data segment of the current moment of the anomaly occurrence node, analyze the change in target thermal radiation intensity gradient from image feedback data, extract the motion direction offset from video feedback data, and verify the consistency of target contour through geometric similarity. The changes in thermal radiation intensity gradient, the offset of motion direction, and the contour consistency results are spatiotemporally aligned and fused to construct an abnormal behavior feature vector. The abnormal behavior feature vector is compared with a preset library of typical violation patterns to match and output the abnormal type. Then, based on the timestamp, the original data source is associated to generate a complete chain of adjudication evidence that includes node location, abnormal behavior feature evidence, and pattern matching results.
9. A guidance and assessment system based on multimodal data fusion, characterized in that: The guidance and assessment method based on multimodal data fusion according to any one of claims 1 to 8 includes: The data acquisition module is used to collect multimodal data in real time during the training process, including image feedback data and video feedback data. The data processing module is used to preprocess the image feedback data and combine it with the time series of the video feedback data to perform spatiotemporal alignment on the preprocessed image feedback data, eliminate the time offset and spatial distortion between different modalities, and obtain the data to be verified. The parameter extraction module is used to match and compare the data to be verified with the preset standard template to extract the associated feature parameters used as the basis for adjudication. The associated feature parameters include dynamic features and static features. The quantitative analysis module is used to perform time-series analysis on dynamic features based on the sliding window mechanism, determine the changing trend of dynamic features, and evaluate the confidence of static features. Then, it combines the changing trend of dynamic features for quantitative fusion to generate a comprehensive adjudication score. The anomaly identification module is used to identify abnormal behaviors that do not comply with the behavior compliance during the training and exercise process based on the comprehensive adjudication score. It determines the occurrence path and node of the abnormal behavior by tracing back the dynamic characteristics, and performs cross-validation by combining spatiotemporally aligned multimodal data to generate a traceable adjudication evidence chain.
10. An electronic device, characterized in that: The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the guidance and assessment method based on multimodal data fusion as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Data processing method and device for guidance evaluation
CN119989302A
Dynamic multi-modal data fusion and real-time analysis method
CN120046119A
Behavior analysis method based on video recognition, processor and storage medium
CN120182766A
Immersive VR psychological detection system and method based on multi-modal AI
CN120436642A
Multi-modal data alignment method and system
CN120492947A