Dynamic multi-modal data fusion and real-time analysis method

By standardizing multimodal data and time axis alignment, evaluating stability scores and calculating the synergistic offset strength, the problems of synergistic offset and cascade errors in multimodal data fusion are solved, and more stable and accurate data fusion effects are achieved, and the robustness and intelligent decision-making capabilities of the system are enhanced.

CN120046119APending Publication Date: 2025-05-27中铁水利信息科技有限公司
View PDF 0 Cites 16 Cited by

Patent Information

Application Number
CN202510518727.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When processing multimodal data, the prior art tends to cause catastrophic output due to the intermodal synergistic shift and cascading amplification effects, and it is difficult to predict and handle abnormal situations in all combinations.

Method used

The stability scores for each modality are evaluated by standardizing the multimodal data and timeline alignment, and features are extracted using deep neural networks for fusing. During the fusion process, the characteristic change trend and response behavior are analyzed in real time, the coordinated offset intensity is calculated, and the suspected abnormal mode is soft-blocked based on the joint feedback of stability score and coordinated offset intensity, so as to predict future fusion risks and determine whether the model execution process is interrupted.

Benefits of technology

It effectively avoids misjudgment and system instability caused by intermodal coordinated shifts and cascading errors, significantly improves the stability and accuracy of the multimodal data fusion system, and enhances the robustness and intelligent decision-making capabilities of the system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046119A_ABST
    Figure CN120046119A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic multi-modal data fusion and real-time analysis method, and particularly relates to the technical field of data analysis. Carrying out standardization and time alignment on the multi-modal data, evaluating the stability of each modal in the current scene, and calculating a stability score; extracting features by using a deep neural network, fusing the features, detecting abnormal synchronization offset between modes in real time, and calculating collaborative offset intensity; based on the combined feedback of the stability score and the offset intensity, soft shielding processing is executed on the suspected abnormal mode; the future fusion risk is further evaluated through a time sequence prediction algorithm, and the model execution process is interrupted and an exception handling mechanism is triggered when necessary, so that the reliability and the safety of the multi-modal system in a complex environment are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and particularly to a method for dynamic multimodal data fusion and real-time analysis. Background Art

[0002] Dynamic multimodal data fusion and real-time analysis refers to the ability of a system to dynamically adjust the fusion strategy according to the evolution of time and perform immediate analysis and decision-making when processing multimodal data generated from different sensors or data sources (such as images, audio, text, radar, GPS, etc.). The core lies in effectively integrating heterogeneous data in the time dimension through intelligent algorithms to achieve a more comprehensive and accurate understanding and prediction. However, when multiple modalities trigger similar bias responses under the same abnormal state (for example, being allergic to a certain kind of noise simultaneously), the model will produce a cascading amplification effect, resulting in catastrophic outputs. In addition, due to the difficulty in covering all combination scenarios during model training, this problem is extremely difficult to predict. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for dynamic multimodal data fusion and real-time analysis to solve the deficiencies in the background art.

[0004] To achieve the above purpose, the present invention provides the following technical solution: A method for dynamic multimodal data fusion and real-time analysis, including: Performing normalization processing on the obtained multimodal data and aligning it on the time axis; Based on the normalized multimodal data, evaluating the volatility of each modality in the current scenario to obtain stability scores at different modality levels; Using a deep neural network to extract and fuse different modality features. During the fusion process, real-time analysis is carried out on whether the feature change trends and response behaviors of different modalities show abnormal synchronous offsets, and the collaborative offset intensity is calculated; Based on the joint feedback of the stability scores and collaborative offset intensities of different modalities, performing soft masking processing on suspected abnormal modalities; Predicting the fusion risk of suspected abnormal modality data in subsequent time periods, and judging whether to interrupt the model execution process according to the prediction result and performing abnormal handling.

[0005] Preferably, the multimodal data includes: visual modality data, auditory modality data, language modality data, and environmental modality data.

[0006] Preferably, when processing the auditory modality, it includes unifying the audio sampling rate, performing noise reduction processing, and dividing the audio signal into frames according to a fixed window; the processing of the language modality includes word segmentation, encoding, temporal annotation, and using a pre-trained language model to convert the text into context embedding vectors; the processing of the environmental modality includes unit unification, sampling frequency synchronization, interpolation filling, and removing outliers using the Z-score method.

[0007] Preferably, the input is a multi-modal feature sequence after standardization , where is the feature sequence of the z-th modality in the current time period. For each modality , a time window is selected to extract the volatility-sensitive features in the multi-modal feature sequence: including short-term variance and signal-to-noise ratio; after normalizing the short-term variance and signal-to-noise ratio, a standardized index set for each modality is constructed, and the stability score is calculated using weighted sum averaging for the normalized short-term variance and signal-to-noise ratio.

[0008] Preferably, analyze whether the feature change trends and response behaviors of different modalities show abnormal synchronous offsets, and calculate the co-offset intensity. The calculation method is as follows: Set the feature time series of the image modality and the audio modality, calculate the feature distance between each pair of time points, generate an n×m distance matrix, use a recursive formula to calculate the cumulative distance matrix of the shortest path, initialize the distance, obtain the total distance of the optimal alignment path, and calculate the ratio of the total distance to the length of the optimal path to obtain the co-offset intensity.

[0009] Preferably, for each modality, normalize the stability score and the co-offset intensity so that they are both in the range of [0,1], and fuse its stability score and the offset participation degree in the co-offset intensity to generate a comprehensive anomaly score.

[0010] Preferably, compare the obtained anomaly score with the set threshold. If the anomaly score is greater than the set threshold, it is regarded as a suspected abnormal modality and shielding processing is performed; if the anomaly score is less than or equal to the set threshold, no additional processing is performed.

[0011] Preferably, for each modality marked as suspected abnormal, extract its anomaly score, stability score, and co-offset intensity in the nearest K time windows; form a multi-dimensional time anomaly sequence as the input of the Transformer model; The model predicts the fusion risk score trend in the future W time windows; Output the risk score curve; According to the prediction result, judge whether to trigger the model interruption and anomaly handling mechanism: if the risk score at any future time point > 0.8, then trigger the model interruption; If the score is > 0.7 at multiple time points continuously, an early warning is given and the degraded operation mode is executed. If an interruption is triggered, immediately stop the current model fusion and inference process.

[0012] Preferably, the degraded operation mode includes: temporarily removing high-risk modalities, only using the remaining modalities for fusion and inference, and marking the risk-limited inference status in the inference result.

[0013] In the above technical solution, the technical effects and advantages provided by the present invention are as follows: 1. The dynamic multi-modal data fusion and real-time analysis method provided by the present invention can dynamically identify and process potential risks caused by collaborative offset between modalities during the fusion process. By standardizing multi-modal data and aligning it on the time axis, combining modal volatility analysis and collaborative offset intensity calculation, a comprehensive anomaly score at the modal level is constructed. When a suspected abnormal modality is detected, a soft shielding strategy is executed, effectively avoiding misjudgment or cascading errors caused by the instability of weak modalities, and significantly improving the stability and accuracy of the fusion system.

[0014] 2. The present invention introduces a risk trend prediction mechanism based on Transformer, which can pre-evaluate the fusion risk of abnormal modalities in future time periods, and dynamically decide whether to interrupt the model process or execute the degraded operation according to the prediction result, effectively ensuring the safe operation and real-time response ability of the system in the face of complex environments, sudden interferences or collaborative biases between modalities. Overall, the present invention enhances the robustness, reliability and intelligent decision-making level of the multi-modal system in real scenarios, and has broad application value. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0016] Figure 1 It is the method flow chart of the present invention. Detailed Embodiments

[0017] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] For the embodiments, please refer to Figure 1 As shown, the dynamic multimodal data fusion and real-time analysis method described in this embodiment includes: Perform normalization processing on the acquired multimodal data and align it along the time axis; Based on the normalized multimodal data, evaluate the volatility of each modality in the current scenario to obtain stability scores at different modality levels; Use a deep neural network to extract and fuse features of different modalities. During the fusion process, real-time analyze whether the feature change trends and response behaviors of different modalities show abnormal synchronous offsets, and calculate the collaborative offset intensity; Based on the joint feedback of the stability scores and collaborative offset intensities of different modalities, perform soft masking on suspected abnormal modalities; Predict the fusion risk of suspected abnormal modality data in subsequent time periods, and determine whether to interrupt the model execution process according to the prediction result and perform anomaly handling.

[0019] Determine the types of input modalities: such as visual modality (image / video), auditory modality (audio / speech), language modality (text), environmental modality (sensor data, such as temperature, speed, position, etc.). Configure corresponding processing modules for each modality to ensure that subsequent processes can run in parallel.

[0020] Scale all images to a unified size (such as 224×224 or 512×512). Normalize the image color values to [0,1] or standardize them to zero mean unit variance. If it is a video input, extract key frames at a fixed frame rate (such as 10fps).

[0021] Audio modality processing includes: unified sampling rate: such as unifying to 16kHz. Noise reduction processing: remove background noise (spectral subtraction, Wiener filtering, etc. can be used). Signal segmentation: segment into frames according to the window size (such as 25ms, with an overlap of 10ms).

[0022] Text modality processing includes: word segmentation and cleaning: remove punctuation and stop words, and perform word segmentation (such as based on the BERT tokenizer). Encoding processing: use a pre-trained language model (such as BERT, RoBERTa) to convert the text into an embedding vector. Temporal annotation: if the text has a timestamp (such as speech transcription), its temporal information needs to be retained.

[0023] Sensor modality processing includes: unified unit: such as converting all GPS units to meters and unifying the speed to m / s. Sampling synchronization: unify the sampling frequency (such as once per second), and interpolate and fill in missing frames if there are any. Outlier removal: use a sliding window median or Z-score method to detect abnormal readings and replace them.

[0024] Extract high-dimensional semantic features in the image using convolutional neural networks (such as ResNet, EfficientNet). The output is a fixed-length vector or feature map (such as a 2048-d vector or a 7×7×512 feature map).

[0025] Extract frequency-domain features such as Mel-Spectrogram and MFCC. Feed them into a CNN or 1D-CNN for deeper feature extraction. After inputting the text tokens, use a pre-trained model (such as BERT) to obtain context embeddings. Obtain the CLS vector or token-level embeddings (which can be used for the attention mechanism).

[0026] Extract statistical features (such as mean, volatility, rate of change, etc.). If it is time-series data, feed it into an LSTM / GRU or a temporal CNN to extract dynamic features.

[0027] Unify all modal data based on a unified reference time axis (such as UTC time or the system start time). All data frames of all modalities need to carry corresponding timestamps.

[0028] For low-frequency modalities (such as sensors), supplement synchronization with high-frequency modalities through linear interpolation. Aggregate the feature vectors of all modalities with a fixed window (such as 1 second) to form synchronous segments. At each time point, select the features of each modality that are closest to each other to construct an aligned feature set.

[0029] Detect whether there are missing modalities in the aligned segments (such as missing frames in images or abnormal interruptions in audio). Record the missing modalities for subsequent processing (such as soft masking, weight adjustment). Package the aligned features of all modalities into a unified format according to time segments (such as a multi-modal vector sequence).

[0030] Input the standardized multi-modal feature sequence , where is the feature sequence of the z-th modality (such as image, audio, text, sensor, etc.) within the current time period. For each modality , select a time window (such as the past 5 seconds or N time slices) and extract the following volatility-sensitive features: Including short-term variance and signal-to-noise ratio. The short-term variance is used to measure the local stability of the modal features (such as the change of image feature vectors between consecutive frames). The signal-to-noise ratio is used to evaluate the ratio of useful information to noise in the signal, which is particularly applicable to audio and sensor modalities.

[0031] The calculation method of the short-term variance is: ; is the short-term variance, is the multi-modal feature of the th frame, is the average feature, is the total number of frames; signal-to-noise ratio The calculation method is as follows: ; where is the variance of the effective signal, and the calculation method is: if the signal can be directly measured, take the variance within the time window of the original input data; if it is a "clean signal" after filtering, then this value is determined by the stable eigenvalue output by the filter; in deep learning, the variance of the activation value of the feature layer can also be used to represent the signal strength. is the variance of the noise part, which is obtained through residual estimation is a positive number to prevent the denominator from being 0.

[0032] Since the short-term variance and the signal-to-noise ratio have different dimensions, each index needs to be normalized to the interval [0,1]. For example, use min-max normalization or Z-score standardization to obtain the set of standardized indexes for each modality, and use weighted summation average to calculate the stability score MSI. The higher the stability score, the more stable it represents.

[0033] The smaller the stability score, the greater the feature volatility of this modality in the current time period, the more unstable the signal, and it may be affected by interference, noise, or data anomalies. At this time, the credibility of the feature output of this modality is relatively low, and it is easy to introduce biases when participating in fusion, thereby increasing the abnormal risk of the overall system. On the contrary, the larger the stability score, the more stable the signal of this modality, the smaller the fluctuation, and it has higher reliability, which is conducive to the collaboration and accurate decision-making between modalities.

[0034] Use a deep neural network to extract and fuse features of different modalities, specifically including: For heterogeneous data from different sources and forms (modalities), such as images, audio, text, sensor data, etc., use a deep neural network to perform high-level feature extraction respectively, and then achieve effective fusion in a unified representation space, so as to obtain comprehensive features containing cross-modal semantic associations.

[0035] Image modality feature extraction: The input image data usually comes from cameras, video streams, etc. After being standardized (size scaling, pixel normalization), it is input into the image feature extraction network. Common network structures include: ResNet, EfficientNet, VGG, Vision Transformer (ViT), etc. The output is a high-dimensional semantic feature vector or feature map. If it is a video frame sequence, a temporal network (such as 3D CNN or T-CNN) can be further used to extract temporal dynamic features.

[0036] Audio modality feature extraction: After the original audio data is preprocessed (noise reduction, unified sampling rate), it is usually converted into Mel-Spectrogram or MFCC features. 1D-CNN, 2D-CNN (processing spectrograms), or audio Transformer models are used for feature extraction. The output is an audio semantic feature vector.

[0037] Text modality feature extraction: Text data usually comes from speech transcription, user input, or sensor event descriptions, and is first tokenized and encoded. Pretrained language models (such as BERT, RoBERTa, GPT) are used to extract semantic features. The output can be the [CLS] vector representing the whole sentence, or sentence-level embedding representation.

[0038] Sensor modality feature extraction: Including GPS, IMU, temperature and humidity sensors, etc. After the raw data is standardized, statistical features (mean, rate of change, volatility) are extracted or fed into a time series model (such as LSTM, TCN) for modeling. The output is a physical state semantic feature.

[0039] After feature extraction is completed, there are still differences in dimension, structure, and information density among the high-dimensional semantic features of each modality. Therefore, a reasonable fusion mechanism needs to be designed. The fusion methods are divided into types such as early fusion, alignment fusion, and deep fusion.

[0040] Align the dimensionality of the feature vectors of different modalities (using methods such as fully connected layers, projection matrices, etc.), and map all modalities to a unified shared representation space.

[0041] Concatenate all modality features and input them into the subsequent deep network; Introduce a cross-modal attention mechanism to calculate the attention between modalities (such as the focus of text on images), and achieve dynamic weighted fusion. For example, use the Multi-head Attention structure.

[0042] Design a gating factor for each modality feature to determine its importance in the fusion; Construct a multi-modal graph structure, regard each modality as a graph node, model the relationship between modalities through edge weights, and use GAT or GCN for fusion learning.

[0043] The finally fused unified feature vector is used for subsequent analysis, such as in classification, regression, detection, or control decision-making systems.

[0044] Analyze whether the feature change trends and response behaviors of different modalities show abnormal synchronous offsets, and calculate the co-offset intensity. The calculation method is: Set the feature time series of two modalities as: Image modality: ; Audio modality: ; where: and represent high-dimensional feature vectors at time n or m; calculate the feature distance (such as Euclidean distance) between each pair of time points: ; generate a distance matrix D, calculate the cumulative distance matrix C(i,j) (dynamic programming), and use the recursive formula to calculate the cumulative distance matrix of the shortest path: ; Initialization: C(0,0)=D(0,0), and the remaining boundaries can be set to positive infinity. The final C(n,m) is the total distance of the optimal alignment path.

[0045] Calculate the co-offset intensity , and the expression is: ; where: L is the length of the optimal path (i.e., the number of paired time points). The smaller this value is, the more synchronous the feature trends of the two modalities are, and the higher the degree of cooperation. However, if this synchronization occurs in an abnormal change section, it is a co-offset.

[0046] The greater the co-offset intensity, the more consistent the feature change trends of multiple modalities are within the same time period, and they occur concentrated in the abnormal or drastic change intervals. This may be an indication that the modalities are affected by a common abnormal source (such as noise resonance, environmental mutation). Such abnormal linkages often have an amplification effect and can easily lead to the system fusion result deviating seriously from the expectation. On the contrary, the smaller the co-offset intensity, the more independent the modalities are in terms of feature changes, the more robust the system fusion is, and the lower the abnormal risk.

[0047] For each modality, normalize the stability score and the co-offset intensity so that they are both within [0,1], fuse the stability score and the offset participation degree in the co-offset intensity, and generate a comprehensive abnormal score. For example, calculate the comprehensive abnormal score by weighted average summation of the stability score and the co-offset intensity.

[0048] Compare the obtained abnormal score with the set threshold. If the abnormal score is greater than the set threshold, it is regarded as a suspected abnormal modality and prepare to perform shielding processing; if the abnormal score is less than or equal to the set threshold, no additional processing is performed.

[0049] For each modality determined to be abnormal, perform one or a combination of the following operations during the fusion stage: Reduce the attention or weight parameter of this modality in the fusion module (such as shrinking the attention score). Stop the feature of this modality from participating in the update of gradient propagation to prevent it from continuing to affect model learning. Only retain the high-confidence part of the features of this modality, and mask the remaining features. Temporarily skip the input of this modality in some subtasks or decision-making links.

[0050] For each modality marked as suspected of being abnormal , extract the anomaly scores, stability scores, and co - shift intensities for its nearest K time windows; form an anomaly sequence in multi - dimensional time as the input to the Transformer model.

[0051] The model predicts the trend of the fusion risk score for the next W time windows; Output the risk score curve, e.g., 0.61, 0.69, 0.74, 0.79, 0.82; this represents the risk trend of this modality within the next 5 windows.

[0052] Based on the prediction results, determine whether to trigger the model interruption and anomaly handling mechanism, which can be based on the following strategies: If the risk score at any future time point > 0.8, trigger model interruption; If the scores at multiple consecutive time points > 0.7 (e.g., 3 time points), issue an early warning and execute the degraded operation mode.

[0053] If interruption is triggered, immediately stop the current model fusion and inference process.

[0054] Write the anomaly event to the system log, return the interpretable anomaly information (e.g., there is a continuous anomaly risk in the future for the image modality, and the process has been terminated), and notify the manual intervention module, anomaly review service, or backup decision system to take over.

[0055] If the degraded mode is triggered: temporarily exclude the high - risk modality, and only use the remaining modalities for fusion and inference. Mark the risk in the output result to limit the inference and improve the interpretability of the system.

[0056] In this embodiment, for multi-source heterogeneous modal data such as images, audio, text, and sensors, first, standardize the data and align it along the time axis to construct a multi-modal aligned feature sequence with a unified reference time. Subsequently, extract the volatility-sensitive indicators of each modality in the current scenario through a sliding window, such as short-term variance and signal-to-noise ratio, and calculate the stability score (MSI) after normalization. On this basis, use deep neural networks (such as CNN, Transformer, LSTM, etc.) to extract the high-level semantic features of each modality respectively, and perform cross-modal fusion through methods such as feature concatenation, attention mechanism, gated fusion, or graph neural network to generate a unified comprehensive feature representation. During the fusion process, monitor in real-time whether the change trends of the features of each modality show abnormal synchronous offsets by calculating the co-offset intensity between modality pairs. Based on the weighted fusion of the stability score and the co-offset participation degree, generate an anomaly score at the modality level and compare it with a set threshold. If the score is too high, trigger a soft masking mechanism and execute strategies such as weight decay, gradient freezing, feature masking, or fusion skipping on the abnormal modality to suppress its negative impact on the overall fusion effect. In addition, for the identified abnormal modality, input its historical score sequence into the Transformer model to predict the fusion risk trend in multiple future time windows. If the predicted risk is higher than the set threshold or remains high, interrupt the current inference process, record the abnormal information, and trigger the abnormal handling mechanism or the degraded operation mode, only retaining the low-risk modalities to participate in the decision-making, so as to achieve the forward-looking identification and adaptive response of the system to co-bias and cascading errors, and significantly enhance the robustness, security, and intelligence of the multi-modal system in complex environments.

[0057] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application.

Claims

1. Dynamic multimodal data fusion and real-time analysis method, characterized by: include: Standardize the acquired multimodal data and align the time axis; Based on standardized multimodal data, the volatility of each mode in the current scenario is evaluated to obtain stability scores at different modal levels; Use deep neural networks to extract and fuse features of different modalities. During the fusion process, analyze in real time whether the feature change trends and response behaviors of different modalities show abnormal synchronization deviation, and calculate the strength of the coordinated deviation. Based on the joint feedback of stability scores and co-shift strength of different modes, the suspected abnormal modes are soft-shielded; Predict the fusion risk of suspected abnormal modal data in the subsequent time period, determine whether to interrupt the model execution process based on the prediction results, and handle the exception.

2. The dynamic multimodal data fusion and real-time analysis method according to claim 1, characterized in that: The multimodal data includes: visual modality data, auditory modality data, language modality data and environmental modality data.

3. The dynamic multimodal data fusion and real-time analysis method according to claim 1, characterized in that: When processing the auditory modality, it includes unifying the audio sampling rate, performing noise reduction, and dividing the audio signal into frames according to fixed windows; the processing of the language modality includes word segmentation, encoding, time series labeling, and using a pre-trained language model to convert the text into a context embedding vector; the processing of the environmental modality includes unit unification, sampling frequency synchronization, interpolation and padding, and the use of the Z-score method to eliminate outliers.

4. The dynamic multimodal data fusion and real-time analysis method according to claim 1, characterized in that: Input standardized multimodal feature sequence ,in is the feature sequence of the zth mode in the current time period. For each mode , select a time window to extract volatility-sensitive features in the multimodal feature sequence: Including short-term variance and signal-to-noise ratio; after normalizing the short-term variance and signal-to-noise ratio, a standardized indicator set for each mode is constructed, and the stability score is calculated using weighted sum and average of the normalized short-term variance and signal-to-noise ratio.

5. The dynamic multimodal data fusion and real-time analysis method according to claim 4, characterized in that: Analyze whether the characteristic change trends and response behaviors of different modes show abnormal synchronous offset, and calculate the intensity of the coordinated offset. The calculation method is: Set the characteristic time series of the image modality and the audio modality, calculate the characteristic distance between each pair of time points, generate an n×m distance matrix, use the recursive formula to calculate the cumulative distance matrix of the shortest path, initialize the distance, obtain the total distance of the optimal alignment path, calculate the ratio of the total distance to the optimal path length, and obtain the collaborative offset strength.

6. The dynamic multimodal data fusion and real-time analysis method according to claim 5, characterized in that: For each mode, the stability score and the co-shift strength are normalized to be between [0, 1], and its stability score and the shift participation in the co-shift strength are fused to generate a comprehensive anomaly score.

7. The dynamic multimodal data fusion and real-time analysis method according to claim 6, characterized in that: The obtained anomaly score is compared with the set threshold. If the anomaly score is greater than the set threshold, it is regarded as a suspected abnormal mode and shielding is performed; if the anomaly score is less than or equal to the set threshold, no additional processing is performed.

8. The dynamic multimodal data fusion and real-time analysis method according to claim 7, characterized in that: For each modality marked as suspected anomaly, extract the anomaly score, stability score, and co-shift strength of the most recent K time windows; form a multi-dimensional time anomaly sequence as the input of the Transformer model; The model predicts the trend of fusion risk scores in the future W time windows; Output risk score curve; Based on the prediction results, determine whether to trigger the model interruption and abnormal handling mechanism: if the risk score at any time point in the future is greater than 0.8, the model interruption will be triggered; If the score is > 0.7 for multiple time points, an early warning will be issued and the downgraded operation mode will be implemented; If an interrupt is triggered, the current model fusion and reasoning process will be stopped immediately.

9. The dynamic multimodal data fusion and real-time analysis method according to claim 8, characterized in that: The degraded operation mode includes: temporarily eliminating high-risk modes, using only the remaining modes for fusion and reasoning, and marking the risk-limited reasoning status in the reasoning results.

Citation Information

Cited By

  • Multi-mode perception fused man-machine interaction intelligent identification system

    CN120354100A

  • Multimodal perception fusion human-computer interaction intelligent recognition system

    CN120354100B

  • Standardized detection result calibration method based on multi-modal fusion

    CN120597221A

  • Method for preventing leakage of sovereign data in combination with multi-mode deception feature perception

    CN120597326A

  • A sovereign data leak prevention method combined with multi-modal deception feature perception

    CN120597326B