Multi-mode Internet of Things sensor data fusion processing method and system

By employing a multimodal IoT sensor data fusion processing method, the problem of misjudgment caused by cross-modal noise coupling is solved, enabling high-reliability monitoring in complex environments and improving the system's accuracy and robustness.

CN121302228AInactive Publication Date: 2026-01-09KUNSHAN XIAOCHENG INTERNET OF THINGS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511314983.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-01-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing IoT multimodal sensor fusion technology suffers from noise amplification and misjudgment of target events in complex environments due to neglecting cross-modal noise coupling, thus reducing the reliability and availability of the system.

Method used

By acquiring multimodal sensor data, interference source localization analysis is performed, interference source distribution maps are generated, interference waveform features are separated, cross-modal interference propagation networks are constructed, noise correlation is weakened, dynamic fusion analysis is performed, and target events or states are detected.

Benefits of technology

It effectively eliminates cross-modal noise coupling, improves the accuracy and robustness of multimodal data fusion, reduces false alarm rate, and ensures high-reliability monitoring in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121302228A_ABST
    Figure CN121302228A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of Internet of Things data processing, and discloses a multi-mode Internet of Things sensor data fusion processing method and system. Comprising the following steps: acquiring multi-modal sensor data in a monitoring environment; performing interference source positioning analysis on the data to generate an interference source distribution map; separating interference waveform characteristics of each modal data based on the distribution map; constructing an interference propagation network by using the features so as to quantify the coupling relationship of interference among different modes; performing noise correlation weakening on original data based on the network to obtain decoupled sensor data; performing dynamic fusion analysis on the decoupling data, and constructing a multi-modal fusion feature vector; and finally, accurately detecting a target event or state based on the feature vector. According to the method, the negative influence of noise correlation on fusion is eliminated, and the accuracy and robustness of target detection and the reliability of the system in a complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet of Things (IoT) data processing technology, and more specifically, to a method and system for multimodal IoT sensor data fusion processing. Background Technology

[0002] With the rapid development of Internet of Things (IoT) technology and the continuous reduction in sensor costs, deploying multiple different types of sensors (such as cameras, microphones, accelerometers, and infrared sensors) in the same monitoring environment has become a mainstream trend. By fusing data from these multimodal sensors, the system can overcome the limitations of a single sensor and obtain a more comprehensive, accurate, and robust environmental perception capability than any independent source. This has enormous application value in fields such as smart manufacturing, structural health monitoring, smart homes, and security monitoring.

[0003] Currently, research on multimodal data fusion technology mainly focuses on three levels: data-level fusion, feature-level fusion, and decision-level fusion. Among them, feature-level fusion is widely used due to its good balance between information preservation and computational complexity. Its typical process includes: first, independently preprocessing the raw data collected by each sensor, such as filtering and noise reduction, and normalization; then, extracting the features from the preprocessed data; finally, combining these features from different modalities through methods such as weighted averaging, Bayesian inference, neural networks, or DS evidence theory to form a unified multimodal fusion feature vector, and performing final pattern recognition or state determination based on this vector.

[0004] However, existing IoT multimodal sensor fusion technologies face a bottleneck: most are based on a fragile and unrealistic "independent noise assumption." Traditional methods typically treat the noise from different sensors as uncorrelated random variables and independently denoise each data stream, but this severely ignores the ubiquitous cross-modal noise coupling phenomenon in the real physical world. In complex industrial, urban, or indoor environments, a single, powerful source of interference (such as a starting air conditioner compressor, a running elevator, or stamping equipment on a large production line) does not only produce a single-mode disturbance but also creates a pervasive interference characteristic that diffuses throughout the monitoring space. For example, when a central air conditioning unit is running, its low-frequency mechanical vibrations are captured by a high-precision accelerometer, its fan and airflow noise are recorded by a microphone array, its electromagnetic radiation from its motor interferes with nearby magnetometers, and even the slight structural vibrations it causes can appear as an imperceptible "jelly effect" in video surveillance footage. These seemingly disparate "noises"—vibrations, sounds, electromagnetic waves, and image jitter—are not independent; they originate from the same physical process, thus exhibiting high temporal synchronization and an inherent structural correlation in their spectral characteristics. Traditional data processing workflows, lacking a cognitive framework for recognizing these shared interference sources, render their independent noise reduction methods ineffective. When this data, containing strongly correlated noise, enters the fusion stage, it produces a distorted effect of noise amplification rather than suppression: the true target signal, which should be highlighted through multi-source information complementarity, is instead submerged by this structured interference feature, present in all modalities and mistakenly considered a valid feature. The fusion algorithm incorrectly identifies this cross-modal consistent noise as a high-confidence event, contaminating the final fused feature vector. Consequently, the system misjudges a normal air conditioner start-up and shutdown as an equipment malfunction or security intrusion event, generating numerous false alarms. This severely undermines the original purpose of data fusion and significantly limits its reliability and usability in real-world, complex scenarios.

[0005] In view of this, the present invention proposes a multimodal IoT sensor data fusion processing method and system to solve the above problems. Summary of the Invention

[0006] To overcome the aforementioned shortcomings of the prior art and to achieve the above objectives, the present invention provides a multimodal Internet of Things (IoT) sensor data fusion processing method, comprising: Acquire multimodal sensor data from multiple IoT sensors in the monitoring environment at each sampling time, wherein the multimodal sensor data includes sensor data from at least two different modes; Interference source localization analysis is performed on the multimodal sensor data at each sampling time to generate an interference source distribution map; Based on the interference source distribution map, the interference waveform characteristics of each mode sensor data are separated; A cross-modal interference propagation network is constructed using the aforementioned interference waveform characteristics; Based on the cross-modal interference propagation network, noise correlation is reduced on the multimodal sensor data to obtain decoupled sensor data; Dynamic fusion analysis is performed on the decoupled sensor data at each sampling time to construct a multimodal fusion feature vector; Based on the multimodal fusion feature vector, target events or target states in the monitoring environment are detected.

[0007] On the other hand, the present invention provides a multimodal Internet of Things (IoT) sensor data fusion processing system, comprising: The data acquisition module is used to acquire multimodal sensor data from multiple IoT sensors in the monitoring environment at each sampling time. The multimodal sensor data includes sensor data from at least two different modes. The interference source localization module is used to perform interference source localization analysis on the multimodal sensor data at each sampling time and generate an interference source distribution map; The interference feature separation module is used to separate the interference waveform features of each mode sensor data based on the interference source distribution map. An interference propagation network construction module is used to construct a cross-modal interference propagation network using the characteristics of the interference waveform; The data decoupling processing module is used to reduce the noise correlation of the multimodal sensor data based on the cross-modal interference propagation network to obtain decoupled sensor data. The dynamic fusion analysis module is used to perform dynamic fusion analysis on the decoupled sensor data at each sampling time and construct a multimodal fusion feature vector; The target detection module is used to detect target events or target states in the monitoring environment based on the multimodal fusion feature vector. The modules are connected via wired and / or wireless means to enable data transmission between them.

[0008] The technical effects and advantages of the multimodal IoT sensor data fusion processing method and system of this invention are as follows: This invention effectively analyzes and deconstructs the complex coupling relationships that pervade different sensor data streams, caused by common environmental interference sources. By deeply identifying and quantifying this cross-modal systematic disturbance, this invention eliminates spurious feature correlations caused by noise correlation, thereby preventing interference signals from being erroneously amplified during the fusion process. This ensures that multi-source information fusion truly achieves its intended function of enhancing target features and suppressing irrelevant disturbances. This process improves the signal integrity of each sensor data stream, providing a clean data foundation for subsequent analysis. Furthermore, the adaptive fusion mechanism introduced in this invention enables the system to dynamically evaluate the instantaneous quality of data from each modality, allowing for intelligent information integration based on reliability and information contribution, enhancing the system's robustness in the face of local sensor failures or transient strong interference. This invention can construct a high-fidelity representation of target events or states with high temporal coherence and logical consistency from complex raw sensing data, significantly reducing false alarms and false negatives, and enabling monitoring and detection results to reach the high reliability level necessary for long-term stable operation in complex real-world environments. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the multimodal IoT sensor data fusion processing method of the present invention; Figure 2 This is a schematic diagram of the multimodal IoT sensor data fusion processing system of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] This application provides a method and system for multimodal IoT sensor data fusion processing. The system's execution entities include, but are not limited to, data processing equipment, IoT control systems, monitoring gateways, and edge computing units, which can be considered general computing nodes in this application. The IoT control system includes, but is not limited to, at least one of the following: an IoT PLC controller, a distributed IoT monitoring system, and a programmable IoT controller.

[0012] This invention provides a multimodal IoT sensor data fusion processing method. By acquiring multimodal data from multiple IoT sensors in a monitoring environment in real time, it analyzes the distribution characteristics of interference sources, separates interference waveform features, constructs a cross-modal interference propagation network, and weakens the noise correlation of sensor data. This achieves dynamic fusion analysis of multimodal data, ultimately accurately detecting target events or states in the monitoring environment. It is highly adaptable, capable of accurately identifying and eliminating noise interference in complex monitoring environments, significantly improving the accuracy and reliability of multimodal data fusion.

[0013] In this embodiment of the invention, the detailed implementation steps of the multimodal IoT sensor data fusion processing method include: First, multimodal sensor data from multiple IoT sensors in the monitoring environment is acquired at each sampling time. This multimodal sensor data includes sensor data from at least two different modes, such as sensing data of different physical quantities like temperature, humidity, sound, image, vibration, and gas concentration. This data is collected in real-time by various IoT sensors distributed throughout the monitoring environment, forming a multi-dimensional sensing dataset at each sampling time. The acquisition of multimodal data provides a rich source of information for subsequent analysis, enabling the description of changes in the monitoring environment's state from different perspectives, ensuring the comprehensiveness and accuracy of the monitoring.

[0014] Interference source localization analysis is performed on multimodal sensor data at each sampling time to generate an interference source distribution map. This map is a visual representation of the spatial distribution of various interference factors in the monitored environment, reflecting the location, intensity, and range of influence of the interference sources. The localization analysis process involves in-depth mining of the time-frequency characteristics of the multimodal data to identify energy anomaly regions, track their temporal variation characteristics, and determine the location coordinates of the interference sources. The interference source distribution map provides a spatial reference for subsequent interference feature separation and is fundamental to achieving accurate noise cancellation.

[0015] Based on the interference source distribution map, interference waveform characteristics of each modal sensor data are separated. Interference waveform characteristics refer to the typical waveform patterns exhibited by the interference source in each modal sensor data, including frequency characteristics, amplitude characteristics, and time-varying patterns. The separation process first determines the propagation range of the interference source, extracts the corresponding local data segments, separates the interference components through spectral analysis, reconstructs the interference waveform, and extracts key feature parameters. These interference waveform characteristics accurately describe how interference affects different modal data, providing a basis for subsequent construction of the interference propagation network.

[0016] A cross-modal interference propagation network is constructed using interference waveform characteristics. This network is a mathematical model describing the propagation relationship of interference between different modal sensor data, revealing the influence mechanism of interference sources on each modal data. The construction process is based on interference waveform characteristics, analyzing the interference propagation relationship between any two modal data points, calculating propagation delay, constructing an initial propagation map, performing propagation intensity analysis, and finally generating a stable network structure through dynamic optimization. This network model accurately characterizes the noise coupling relationship between multimodal data, providing a theoretical basis for subsequent data decoupling.

[0017] Based on cross-modal interference propagation networks, noise correlation attenuation is performed on multimodal sensor data to obtain decoupled sensor data. Noise correlation attenuation is a processing technique to reduce noise interference between multimodal data, aiming to eliminate interrelated noise components in different modal data. The process first identifies high-impact propagation paths, analyzes the influence of each modal data on interference propagation, constructs an interference suppression matrix, performs a linear transformation on the original data, and finally removes residual interference components through residual analysis. Decoupling the sensor data preserves the effective information of each modality while significantly reducing the impact of interference factors, laying the foundation for subsequent fusion analysis.

[0018] Dynamic fusion analysis is performed on decoupled sensor data at each sampling time to construct a multimodal fusion feature vector. This multimodal fusion feature vector is a comprehensive representation of information from multiple sensor modalities, providing a more complete description of the state characteristics of the monitored environment. The fusion analysis process includes time-series segmentation, analysis of temporal consistency, distribution analysis, construction of dynamic fusion weights, and finally, generation of the fusion feature vector. This dynamic fusion strategy adaptively adjusts the fusion weights based on the reliability and information value of different modal data, ensuring the optimality and stability of the fusion results.

[0019] Based on multimodal fusion feature vectors, this method detects target events or target states in the monitoring environment. This is the ultimate goal of the entire approach: applying the results of the preliminary processing to actual monitoring tasks. The detection process first acquires a pre-defined reference feature library, calculates the dynamic matching degree between the fused feature vector and the reference features, identifies candidate target events or states, and finally generates the final detection result through temporal smoothing processing. This multimodal fusion-based detection method fully utilizes the complementary advantages of different sensors, significantly improving the accuracy, reliability, and robustness of the detection.

[0020] In this embodiment of the invention, the detailed implementation steps for performing interference source localization analysis on the multimodal sensor data at each sampling time and generating an interference source distribution map include: At each sampling time, time-frequency decomposition is performed on the sensor data for each modality to generate a time-frequency distribution map. The time-frequency distribution map represents the energy distribution of sensor data in both time and frequency dimensions, visually displaying the frequency characteristics of the data and their changes over time. The decomposition process employs time-frequency analysis techniques such as short-time Fourier transform or wavelet transform to convert one-dimensional time-series data into a two-dimensional time-frequency representation. Appropriate window functions and window lengths are selected during processing to balance time and frequency resolution, ensuring accurate capture of short-term interference characteristics. Each point in the time-frequency distribution map represents the energy value at a specific time and frequency; color coding visually displays the energy intensity, providing a foundation for anomaly region identification.

[0021] In the time-frequency distribution map, high-energy anomalous regions are identified. These regions are defined as areas with energy values ​​exceeding a preset energy threshold. High-energy anomalous regions typically correspond to abnormal signals generated by interference sources and are crucial clues for locating these sources. The identification process employs adaptive thresholding technology, determining an appropriate energy threshold based on the overall energy distribution characteristics of the time-frequency distribution map. The preset energy threshold is usually set to 3-5 times the average energy of the time-frequency distribution map, or statistically determined using the 95th percentile of the energy distribution. The identification results are marked on the time-frequency distribution map as connected regions. Each connected region corresponds to a potential interference event, containing information such as time range, frequency range, and energy distribution.

[0022] A time-series tracking method is used to analyze high-energy anomaly regions, calculating their movement trajectories over multiple consecutive sampling times. These trajectories reflect the evolution of interference characteristics in the time-frequency domain and are crucial for determining the location of the interference source. The tracking process employs a target tracking algorithm to establish the correspondence between anomaly regions across consecutive sampling times, forming a complete time-series chain. The algorithm considers regional overlap, energy variation trends, and frequency drift characteristics, effectively handling complex situations such as region splitting and merging. The movement trajectories are represented as time-frequency coordinate sequences, recording the change in the centroid position of the anomaly region over time, providing dynamic features for subsequent analysis.

[0023] Based on the movement trajectory, candidate locations of interference sources are determined. These candidate locations are the convergence points of the movement trajectories. Convergence points typically correspond to the actual location of the interference source, reflecting the source region of energy propagation. The determination process analyzes the directionality and convergence characteristics of the movement trajectories, identifying regions where the trajectory direction change rate is small and multiple trajectories tend to converge. The algorithm calculates the direction vector and curvature features of the trajectories, and uses threshold filtering to identify potential convergence regions. Candidate locations are represented in spatial coordinates, with each location assigned a confidence score reflecting the likelihood that the location is a true interference source.

[0024] Cluster analysis is performed on candidate locations, merging those within a preset distance threshold to generate an interference source distribution map. This map includes the location and intensity of the interference sources. Cluster analysis aims to eliminate redundant candidate locations and improve the accuracy of interference source localization. The analysis employs density-based clustering algorithms, such as DBSCAN or a modified K-means, automatically determining the number of clusters based on the spatial distribution characteristics of the candidate locations. The preset distance threshold is determined based on the spatial scale of the monitoring environment, typically 5-10% of the monitoring area diameter. Each cluster center is considered an interference source location, and the number of candidate locations within each cluster and the average confidence score are used to calculate the interference source intensity. The final interference source distribution map is visualized as a heatmap, intuitively showing the spatial distribution and relative intensity of interference sources in the monitoring environment, providing a precise spatial reference for subsequent interference feature analysis.

[0025] In this embodiment of the invention, the detailed implementation steps for separating the interference waveform characteristics of each mode sensor data based on the interference source distribution map include: Based on the interference source distribution map, the propagation range of each interference source is determined. The propagation range is a circular area centered on the location of the interference source with a radius equal to the interference intensity. The propagation range defines the spatial extent of the interference source's influence and is a key area for extracting interference features. The determination process considers the relationship between interference intensity and spatial attenuation, and uses a radiation model to calculate the effective influence radius. Interference intensity is directly mapped to the propagation radius; the greater the intensity, the wider the influence range, ensuring the capture of complete interference features. For complex environments, the blocking effect of spatial obstacles is also considered, and the geometry of the propagation range is corrected using ray tracing technology to better reflect actual propagation characteristics. The propagation range of each interference source is described in analytical geometry to facilitate subsequent region extraction and analysis.

[0026] In the time-series data of each mode of sensor data, local data segments corresponding to the propagation range are extracted. These local data segments are subsets of sensor data directly affected by the interference source, containing complete characteristics of the interference waveform. The extraction process first identifies sensor nodes that spatially intersect with the propagation range, and then extracts time windows corresponding to the interference events from the time-series data of these nodes. The determination of the time windows considers the interference propagation speed and duration to ensure complete capture of the interference events. For each mode of sensor, different spatial mapping and temporal alignment strategies are adopted based on its characteristics to ensure that the extracted data segments accurately correspond to the interference events. Finally, a local data set organized by interference source and sensor mode is formed, providing a focused research object for subsequent analysis.

[0027] Spectral decomposition is performed on local data segments to separate high-frequency interference components and low-frequency trend components. Spectral decomposition aims to separate the interference signal from the background signal in the frequency domain, extracting pure interference features. The decomposition process employs multi-resolution analysis techniques, such as wavelet decomposition or empirical mode decomposition, to decompose the signal into sub-components in different frequency bands. The decomposition level is dynamically adjusted according to the signal complexity, typically 3-5 levels, to ensure sufficient separation accuracy. High-frequency interference components usually correspond to transient signals generated by the interference source, characterized by concentrated energy and distinct frequency characteristics; low-frequency trend components correspond to the environmental background and sensor baseline, exhibiting gradual changes. The separation thresholds for both are adaptively determined based on the signal energy distribution, ensuring the accuracy and stability of the separation.

[0028] Waveform reconstruction is performed on the high-frequency interference components to generate the interference waveform. Waveform reconstruction transforms the frequency domain analysis results back into a time domain representation to obtain the temporal characteristics of the interference signal. The reconstruction process employs inverse transform technology to resynthesize the separated high-frequency components into a time-domain signal. Before reconstruction, the high-frequency components undergo appropriate filtering and denoising to eliminate artifacts that may be introduced during the decomposition process. The reconstruction algorithm considers the phase-preserving principle to ensure that the temporal characteristics of the reconstructed waveform remain consistent with the original interference signal. For multimodal data, different reconstruction parameters are used based on the characteristics of different modes to ensure the accuracy of the reconstruction results. The reconstructed interference waveform is represented as a time-series curve, intuitively demonstrating the complete temporal performance of the interference and providing a foundation for feature extraction.

[0029] Feature extraction is performed on the interference waveform to generate interference waveform features corresponding to each modal sensor data. Feature extraction includes calculating the peak frequency, peak amplitude, and periodicity exponent of the waveform. Feature extraction simplifies complex waveforms into a few key parameters, facilitating subsequent analysis and comparison. The peak frequency is obtained through power spectrum analysis of the reconstructed waveform, reflecting the main frequency components of the interference signal; the peak amplitude is determined by statistically analyzing the extreme value distribution of the waveform, reflecting the interference intensity; and the periodicity exponent is calculated through autocorrelation analysis, quantifying the repetitive pattern of the interference signal. In addition, auxiliary features such as rise time, duration, and spectral entropy are extracted to comprehensively describe the characteristics of the interference waveform. The feature extraction employs a sliding window technique to ensure the capture of the time-varying characteristics of the waveform. Finally, a multi-dimensional feature vector is formed, accurately characterizing the interference waveform features in each modal sensor data, providing a data foundation for the subsequent construction of the interference propagation network.

[0030] In this embodiment of the invention, the detailed implementation steps for constructing a cross-modal interference propagation network using interference waveform characteristics include: At each sampling time, based on the interference waveform characteristics of the data from each modal sensor, the interference propagation relationship between any two modal sensor data points is analyzed. The interference propagation relationship describes how interference propagates from one modality to another and is the foundation for constructing the propagation network. The analysis process employs correlation analysis and causal inference techniques to assess the similarity and temporal dependence of interference features between different modalities. The algorithm calculates the correlation coefficient matrix between feature vectors, applies mutual information analysis to identify nonlinear dependencies, and determines the propagation direction through Granger causality tests. The analysis considers the physical characteristics and spatial distribution of each modal sensor, identifying physically reasonable propagation relationships and avoiding misjudgments caused by spurious correlations. The propagation relationship is represented in the form of directed connections, recording the source mode, target mode, and relationship strength, providing basic elements for subsequent propagation network construction.

[0031] The propagation delay of interference propagation relationships is calculated, which involves comparing the peak time differences of interference waveform characteristics. Propagation delay reflects the time required for interference to propagate from one mode to another and is a crucial indicator of the dynamic characteristics of interference propagation. The calculation process first identifies the characteristic peak times of each mode of interference waveform, and then calculates the time difference between the peak occurrences of different modes. To improve accuracy, the algorithm employs cross-correlation analysis to determine the optimal time alignment, taking into account the impact of sampling rate differences. For complex waveforms, a multi-peak matching strategy is used, and the correspondence is determined through pattern recognition. Propagation delay is expressed in time units; positive values ​​indicate forward propagation from the source mode to the target mode, while negative values ​​indicate reverse influence, and the magnitude of the value reflects the time scale of the propagation process.

[0032] Based on propagation delay, an initial propagation graph is constructed. The nodes of the initial propagation graph represent the data from each modality of the sensor, and the edges represent the propagation delay. The initial propagation graph is a network representation of the interference propagation relationships, intuitively displaying the topology of interference propagation among multimodal data. The construction process employs a graph theory model, treating each modality as a node and establishing directed edges between modal pairs with propagation relationships. Edge attributes include propagation delay and initial relationship strength. The graph construction considers a balance between completeness and simplicity, retaining significant propagation relationships through threshold filtering to avoid overly complex networks. For bidirectional propagation, the decision to retain bidirectional edges or only the primary direction is based on the propagation strength. The initial propagation graph is represented in both adjacency matrix and visual graph formats, providing a basic data structure for subsequent analysis.

[0033] Propagation intensity analysis was performed on the initial propagation graph, calculating the propagation intensity of each edge. The calculation included analyzing the peak-to-amplitude ratio (P / A ratio) and periodicity exponential difference of the interference waveform characteristics. Propagation intensity quantifies the impact of interference from the source mode to the target mode and is a key indicator for evaluating the importance of propagation. The calculation process comprehensively considered multiple factors: the P / A ratio reflects the energy change of the interference signal during propagation, and the periodicity exponential difference reflects the degree of waveform structure preservation. In addition, auxiliary indicators such as waveform similarity and spectral consistency were considered, and a comprehensive evaluation was formed through weighted fusion. The propagation intensity was normalized, with a value range of [0,1], where a larger value indicates a more significant propagation impact. The intensity analysis results were added as new attributes of the edges to the initial propagation graph, providing an important basis for network optimization.

[0034] The initial propagation graph is dynamically optimized by removing edges with propagation strength below a preset strength threshold, generating a cross-modal interference propagation network. Dynamic optimization aims to extract the main propagation paths of interference, eliminate noise and redundant connections, and make the network structure clearer and more stable. The optimization process first determines an appropriate strength threshold, typically using data-driven methods such as adaptive thresholds based on network connectivity or statistical thresholds based on strength distribution. The preset strength threshold is usually set to the 75th percentile of the strength distribution, ensuring that significant propagation relationships are preserved while effectively reducing network complexity. After removing low-strength edges, connectivity analysis is performed on the network, and the threshold is adjusted as necessary to ensure overall network connectivity. The final generated cross-modal interference propagation network retains the critical paths and main structure of interference propagation, providing accurate model support for subsequent noise correlation reduction. The optimized network is stored in the form of an adjacency list, and a visual representation is generated for intuitive understanding of the interference propagation patterns.

[0035] In this embodiment of the invention, the detailed implementation steps for reducing noise correlation in multimodal sensor data and obtaining decoupled sensor data based on a cross-modal interference propagation network include: In cross-modal interference propagation networks, high-impact propagation paths are identified, defined as those whose sum of propagation intensities exceeds a preset path threshold. These high-impact paths are the primary channels through which interference propagates across multimodal data and have a decisive impact on noise correlation. The identification process employs a path analysis algorithm to enumerate all possible propagation paths in the network and calculate the cumulative propagation intensity of each path. The cumulative intensity calculation considers the influence of path length and uses an attenuation coefficient model to ensure the rationality of the assessment. The preset path threshold is typically determined based on the path intensity distribution, using the 80-90% quantile to ensure the selection of the most influential path set. For complex networks, a heuristic search strategy is used, prioritizing the analysis of paths composed of high-intensity edges to improve computational efficiency. The identification results are represented as a set of paths, with each path containing a node sequence and cumulative intensity, providing an object for subsequent impact analysis.

[0036] Impact analysis was conducted on high-impact propagation paths to analyze the influence of sensor data from each modality on interference propagation within each path. The aim of the impact analysis was to determine the role and importance of each modality's data in interference propagation, providing a basis for constructing interference suppression strategies. The analysis process first calculated the frequency and location distribution of each modality across all high-impact paths, assessing its proportion as a source, intermediate node, and endpoint. Then, it analyzed the connectivity and centrality indices of modal nodes, such as degree centrality, betweenness centrality, and eigenvector centrality, quantifying their structural importance in the network. Finally, it assessed the interference generation capability and sensitivity of each modality by combining its own interference waveform characteristics. The analysis results generated an influence score for each modality, reflecting the magnitude and characteristics of different modalities' roles in interference propagation, providing a quantitative basis for constructing the interference suppression matrix.

[0037] Based on impact analysis, an interference suppression matrix is ​​constructed, which involves weighted adjustment of the adjacency matrix of the cross-modal interference propagation network. The interference suppression matrix is ​​a special linear transformation matrix designed to reduce noise coupling between multimodal data. The construction process first obtains the adjacency matrix of the propagation network, with matrix elements representing the propagation strength of corresponding edges. Then, based on the impact analysis results, the matrix elements are weighted to enhance the suppression effect on high-impact paths. The adjustment uses nonlinear mapping functions, such as exponential decay or sigmoid transformation, to amplify the suppression weights of high-strength connections. Diagonal elements are set to unit values ​​to ensure preservation of the intrinsic information of each mode. Finally, the matrix is ​​normalized and regularized to ensure numerical stability and transformation effectiveness. The constructed suppression matrix is ​​a square matrix with a dimension equal to the number of modes, and each element represents the interference suppression coefficient from the row-indexed mode to the column-indexed mode.

[0038] A linear transformation is performed on multimodal sensor data using an interference suppression matrix to generate intermediate decoupled data. The linear transformation is the core step in applying the interference suppression strategy, achieving initial decoupling of the multimodal data through matrix operations. The transformation process first organizes the sensor data for each modality into a feature matrix, where each row represents a sampling time and each column represents a mode. Then, matrix multiplication is performed between the feature matrix and the interference suppression matrix to obtain the transformed intermediate decoupled data. The transformation algorithm considers differences in data dimensionality and scale, employing appropriate preprocessing and normalization strategies to ensure numerical stability. For time-series data, a sliding window technique is used to maintain temporal continuity and smooth transitions. The intermediate decoupled data retains the basic structure of the original data but significantly reduces noise coupling between modes, providing a foundation for subsequent residual analysis.

[0039] Residual analysis is performed on the intermediate decoupled data to remove components highly correlated with interference waveform features, generating decoupled sensor data. Residual analysis is a refined step in the decoupling process, aiming to eliminate interference components remaining after linear transformation. The analysis process first calculates the difference between the intermediate decoupled data and the original data to obtain the initial residuals. Then, feature analysis is performed on the residuals to extract their time-frequency characteristics and statistical distribution. Correlation analysis is performed between the residual features and pre-extracted interference waveform features to identify residual components highly correlated with interference. Matched filtering or projection techniques are used to accurately remove these components from the intermediate decoupled data. A soft thresholding strategy is employed during the removal process, adjusting the suppression strength according to the correlation to avoid over-elimination and loss of useful information. The final generated decoupled sensor data retains the effective information of each mode to the maximum extent while significantly reducing the influence of interference factors, providing a high-quality data foundation for subsequent fusion analysis. The decoupling effect is quantitatively evaluated by the degree of improvement in signal-to-noise ratio and the degree of enhancement in inter-modal independence to ensure the effectiveness of the processing.

[0040] In this embodiment of the invention, the detailed implementation steps for dynamically fusing and analyzing the decoupled sensor data at each sampling time to construct a multimodal fusion feature vector include: Decoupled sensor data at each sampling time point is segmented into multiple local time-series segments. Time-series segmentation is a fundamental preprocessing step for analyzing time-varying data, dividing a continuous data stream into meaningful analytical units. The segmentation process employs adaptive windowing, dynamically adjusting the window size and overlap based on data characteristics. The segmentation algorithm comprehensively considers signal stationarity and event integrity, setting segment boundaries at points of significant data change to ensure relatively consistent characteristics within each segment. For event-driven data, an event detection algorithm is used to assist segmentation, ensuring that important events are not truncated by window boundaries. The segmentation results form a set of time-series segments, each containing a time index and multimodal data values, providing basic units for subsequent consistency analysis. This segmentation strategy improves the relevance and computational efficiency of the analysis while maintaining temporal continuity.

[0041] Within each local time segment, the temporal consistency of sensor data from different modalities is analyzed. This analysis includes calculating the curvature of the dynamic time warping path for decoupled sensor data. Temporal consistency reflects the degree of coordination among different modal data over time and is a crucial indicator of fusion reliability. The analysis employs Dynamic Time Warping (DTW) to calculate the optimal alignment path between any two modal data points. The curvature of the alignment path reflects the difficulty of matching temporal changes; a smaller curvature indicates greater consistency in temporal changes. The calculation process first constructs a distance matrix, then applies a dynamic programming algorithm to solve for the optimal path, and finally calculates the geometric properties of the path. To improve computational efficiency, acceleration algorithms such as FastDTW are used, while window constraints are applied to limit the search range. For multimodal scenarios, a hierarchical alignment strategy is adopted, first calculating the consistency of modal pairs, and then comprehensively evaluating the overall consistency. Temporal consistency is represented as a consistency coefficient matrix, where each element quantifies the degree of temporal coordination of the corresponding modal pair.

[0042] Distribution analysis was performed on decoupled sensor data at all sampling times to analyze the distribution stability of each modality. The stability analysis included kernel density estimation of the decoupled sensor data to generate probability density distribution curves. These curves describe the distribution characteristics of the data values ​​and are fundamental to assessing data stability. The estimation process employed kernel density estimation, using a Gaussian kernel function to smooth the data points, and the bandwidth parameter was determined through cross-validation. The estimation algorithm considered both the data volume and distribution characteristics, employing an adaptive bandwidth strategy for sparse regions to improve estimation accuracy. For multidimensional data, dimensionality reduction techniques were used to extract the main directions of change, or multidimensional kernel density estimation was applied directly. The generated density curves visually demonstrate the central tendency and dispersion of the data, providing a foundation for subsequent distribution feature extraction.

[0043] The probability density distribution curve is segmented for analysis to extract local distribution features within each time period, including peak positions and peak widths. These local features characterize the key properties of the data distribution within a specific time period, reflecting the central tendency and fluctuation range of the data. The extraction process first performs peak detection on the density curve to identify local maximum points. Then, the position coordinates and half-width at half-maximum (WHM) of each peak are calculated to quantify its central location and dispersion. For multi-peak distributions, the number of peaks and their relative heights are recorded to reflect the complexity of the distribution. Feature extraction considers the effects of noise and random fluctuations, employing smoothing and statistical significance testing to ensure the stability and representativeness of the extracted features. Local distribution features are represented as feature vectors, containing complete distribution information and providing a data foundation for subsequent stability analysis.

[0044] Peak drift analysis examines the peak shift of local distribution characteristics over continuous time periods. This analysis includes calculating the distance between peak positions in adjacent time periods. Peak drift reflects the degree of change in the data distribution center over time and is a direct indicator of distribution stability. The analysis first establishes the correspondence between peaks in continuous time periods, identifying peak pairs through proximity principles or feature matching. Then, the Euclidean or Mahalanobis distance between corresponding peaks is calculated to quantify positional changes. Peaks that appear or disappear are recorded as special events and given additional weight. The drift analysis employs a time-weighted strategy, where recent changes have a greater impact than long-term changes, more accurately reflecting current stability. The analysis results form a peak drift sequence, describing the trajectory of the distribution center over time; smaller amplitude changes indicate a more stable distribution.

[0045] The shape variation of the probability density distribution curve is analyzed, including the calculation of its normalized value. Shape variation reflects the stability of the data distribution pattern and is an important supplement to the distribution characteristics. The analysis uses distributional distance measures, such as Kullback-Leibler divergence or Wasserstein distance, to calculate the degree of difference between density curves in adjacent time periods. To eliminate the influence of absolute values, the density curve is normalized, focusing only on the distribution shape. Shape analysis also considers higher-order statistics, such as changes in skewness and kurtosis, to comprehensively assess distribution characteristics. For multimodal data, the shape variation of the joint distribution is calculated to assess the stability of relationships between modes. Shape variation is represented as a rate of change series, reflecting the speed of distribution pattern evolution over time; a lower rate of change indicates a more stable distribution shape. The comprehensive results of the distribution stability analysis include both peak shift and shape variation, comprehensively assessing the stability of the statistical characteristics of each modality of data.

[0046] Based on temporal consistency and distribution stability, dynamic fusion weights are constructed. The construction of dynamic fusion weights involves analyzing the relationship between temporal consistency and distribution stability. Dynamic fusion weights determine the contribution ratio of each modality in the fusion process and are a key parameter for achieving adaptive fusion. The construction process first normalizes the temporal consistency and distribution stability indices to ensure comparability. Then, the correlation between the two types of indices is analyzed to determine a suitable combination strategy. For highly correlated cases, principal component analysis is used to extract common factors; for low-correlation cases, a weighted average strategy is used to retain individual information. Weight construction uses nonlinear mapping functions, such as sigmoid or exponential functions, to convert the normalized indices into weight values ​​in the [0,1] interval. The construction algorithm considers the differences in the importance of the indices and data characteristics, assigning higher basic weights to key modalities while maintaining sufficient dynamic adjustment space. The final generated dynamic fusion weights change in real time with the data characteristics, adaptively adjusting the fusion ratio of each modality to ensure the optimality of the fusion result.

[0047] By utilizing dynamic fusion weights, decoupled sensor data is weighted and fused to generate a multimodal fusion feature vector. Weighted fusion is the process of integrating multimodal data information into a unified representation, a key step in achieving complementary advantages of multi-source information. The fusion process first extracts features from each modality of data, converting them into a unified-dimensional feature representation. Then, dynamic fusion weights are applied to calculate a weighted average or weighted combination, generating a preliminary fusion result. For heterogeneous features, nonlinear fusion strategies, such as deep feature fusion or kernel method fusion, are employed to handle complex feature relationships. The fusion algorithm considers the complementarity and redundancy between features, adjusting the fusion strategy through correlation analysis to maximize information utilization efficiency. The final generated multimodal fusion feature vector is a comprehensive representation of the monitored environmental state, integrating the advantageous information from each modality, providing high-quality feature input for subsequent target detection. The fusion effect is evaluated using metrics such as information entropy gain and classification accuracy improvement to ensure the effectiveness and stability of the fusion process.

[0048] In this embodiment of the invention, the detailed implementation steps for detecting target events or target states in a monitoring environment based on multimodal fusion feature vectors include: The system acquires a pre-defined reference feature library for target events or states. This library forms the knowledge base for target detection, containing standard feature representations of various target events or states. The library is constructed using a combination of prior knowledge and data-driven methods, extracting feature templates from expert-annotated sample data to form a standard reference. The library content is organized by target type, with each target corresponding to a set of feature vectors describing its typical characteristics and variation range. Feature representations use a format compatible with fused feature vectors, ensuring direct comparison. For complex targets, multi-template representations are used to capture their different state characteristics. The feature library supports incremental updates, dynamically expanding and optimizing based on new samples to improve adaptability to target changes. The reference feature library is stored in structured data format, supporting efficient querying and similarity calculation, providing benchmark data for subsequent matching analysis.

[0049] At each sampling time, the dynamic matching degree between the multimodal fusion feature vector and each reference feature in the reference feature library is calculated. The calculation of the dynamic matching degree includes analyzing the similarity between the multimodal fusion feature vector and the reference features. The dynamic matching degree quantifies the degree of conformity between the current observation and the standard target, and is the core basis for target detection. The calculation process uses multidimensional similarity measures, such as cosine similarity, Mahalanobis distance, or DTW distance, to evaluate the matching degree between feature vectors. To improve sensitivity, the algorithm considers the differences in feature importance and assigns higher weights to key dimensions. The matching calculation employs a parallel processing strategy, comparing multiple reference features simultaneously to improve computational efficiency. For temporal features, dynamic warping techniques are used to handle changes in time scale, enhancing the ability to recognize temporal patterns. The matching degree calculation also considers the influence of environmental conditions and observation quality, adjusting the judgment criteria through adaptive thresholds. The calculation results form a matching degree matrix, where each element represents the degree of matching between the current feature and a specific reference target, providing a quantitative basis for target recognition.

[0050] Based on the maximum value of the dynamic matching degree, candidate target events or target states at the corresponding sampling time are determined. Candidate targets are the preliminary identification results, reflecting the most likely target type or state at present. The determination process first finds the maximum value in the matching degree matrix and identifies the reference feature for the best match. Then, the target type corresponding to this reference feature is taken as a candidate target, and the matching degree is recorded as a confidence score. For multiple candidates with similar matching degrees, the top N are retained as a candidate set, reflecting the uncertainty of identification. The determination algorithm considers the absolute value of the matching degree and filters by the minimum matching threshold to avoid misjudgment when there is no matching target. For novel events, when all matching degrees are below a preset threshold, they are marked as "unknown targets," triggering the anomaly detection process. Candidate target information includes target type, confidence score, and key feature deviations, providing input for subsequent time-series smoothing.

[0051] Temporal smoothing is performed on candidate target events or target states across multiple consecutive sampling times. This smoothing process involves constructing a state transition model, eliminating detection results with state transition probabilities below a preset probability threshold, and generating the final target event or target state. Temporal smoothing aims to eliminate instantaneous fluctuations and misjudgments, improving the stability and consistency of detection results. The process first constructs a transition probability model of the target state based on historical data and prior knowledge, describing the reasonable transition patterns between different states. The model employs a Markov chain or conditional random field structure to capture the temporal dependence of the state sequence. The smoothing algorithm calculates the probability of each transition in the candidate target sequence, identifying low-probability transitions as potential misjudgments. The preset probability threshold is typically set to 10-20% of the normal transition probability to ensure that only obviously abnormal transitions are eliminated. For identified anomalies, corrections are made based on the context, or interpolation techniques are used to fill in missing values. The smoothed result forms a coherent sequence of target events or states, eliminating random fluctuations and isolated misjudgments, and truly reflecting the dynamic changes of the monitoring environment. The final detection results are output in time series format, including timestamps, target types, state parameters, and confidence scores, providing a reliable basis for subsequent analysis and decision-making.

[0052] In this embodiment of the invention, the detailed implementation steps for time-series tracking of high-energy anomaly regions and calculation of the movement trajectory of high-energy anomaly regions within multiple consecutive sampling times include: In the time-frequency distribution map of multiple consecutive sampling times, the centroid positions of high-energy anomaly regions are extracted. The centroid position is a key parameter characterizing the spatial properties of the anomaly region, providing the basic coordinate points for trajectory construction. The extraction process first performs connected component analysis on the high-energy anomaly regions at each time point to identify independent anomaly blocks. Then, the weighted centroid of each block is calculated, using energy values ​​as weights to ensure that the centroid position is closer to the energy concentration region. The calculation formula is: ; ; Here, the x-coordinate and y-coordinate of the centroid position are defined, x_i and y_i are the coordinates of the i-th point within the region, and w_i is the energy value corresponding to the i-th point within the region. For regions with complex shapes, principal component analysis is used to assist in determining the principal axis direction and feature point positions, improving the accuracy of centroid extraction. The centroid position is represented by coordinates on the time-frequency plane, forming a discrete position sequence to provide raw data points for subsequent trajectory construction.

[0053] Temporal interpolation is performed on the centroid positions to generate continuous centroid trajectories. Temporal interpolation aims to connect discrete centroid positions into a continuous trajectory curve, eliminating discontinuities caused by sampling intervals. The interpolation process employs spline interpolation techniques, such as cubic splines or B-splines, to generate smooth intermediate transitions while preserving the original data points. The interpolation algorithm considers the uncertainty of centroid positions, reducing constraint strength for points with low confidence to minimize the impact of outliers. For moments with occlusion or detection failures, a combination of forward prediction and backward correction is used to estimate missing locations. The interpolation resolution is dynamically adjusted according to analysis needs, typically 3-5 times the original sampling rate, ensuring the continuity and smoothness of the trajectory. The generated continuous centroid trajectory is represented as a high-density point sequence or a parametric curve, accurately describing the movement path of the anomaly region in the time-frequency plane.

[0054] The centroid trajectory is smoothed to remove noise and jitter. This smoothing process includes calculating the energy intensity of high-energy anomaly regions. The smoothing aims to eliminate trajectory jitter caused by centroid position measurement errors and random fluctuations, resulting in a more physically consistent movement path. The process first calculates the energy intensity of the anomaly region at each time step, serving as the basis for the smoothing weights. The energy intensity is determined by the sum or maximum value of the energy values ​​within the region, reflecting the significance of the anomaly signal.

[0055] Weighting coefficients are constructed based on energy intensity. These weighting coefficients determine the influence of each point in the smoothing process and are key parameters for achieving adaptive smoothing. The construction process employs normalization, mapping energy intensity to weight values ​​in the [0,1] interval. The mapping function uses a non-linear design, such as a sigmoid or logarithmic function, emphasizing the importance of high-energy points. The weighting coefficient calculation considers temporal continuity and treats abrupt energy changes specially to prevent the smoothing process from eliminating true trajectory changes. The calculation formula is as follows: ; Where E_i is the energy intensity at the i-th point, f() is the nonlinear mapping function, and w_i is the corresponding normalized weighting coefficient. The weighting coefficient directly affects the local characteristics of the smoothing process; regions with high energy retain more original features, while regions with low energy undergo stronger smoothing, achieving adaptive processing.

[0056] A weighted moving average is applied to the centroid trajectory using weighting coefficients. Weighted moving average is a commonly used signal smoothing technique that effectively eliminates random noise while preserving the main trend. The averaging process employs a sliding window strategy, weighting the points within the window to generate smoothed points. The window size is dynamically adjusted based on the trajectory characteristics, typically ranging from 5 to 9 points, balancing smoothing effect with detail preservation. The calculation formula is: ; Here, P_{i+j} represents the position point in the original centroid trajectory, where i is the index of the current processing point and j is the relative index within the window. For example, j=0 represents the current point itself, j=1 represents the next point, and j=-1 represents the previous point; w_{i+j} represents the weighting coefficient of the corresponding point P_{i+j}, which is a weight value constructed based on the energy intensity of the high-energy anomaly region, with points having higher energy having higher weights; j is the relative index within the window, P_i is the original point position, and P'_i is the smoothed position. For trajectory endpoints, boundary processing techniques, such as mirror extension or truncated windows, are used to ensure the smoothness and consistency of the entire trajectory. The smoothed trajectory retains the original movement trend while significantly reducing random jitter, providing a stable foundation for subsequent directional analysis.

[0057] Directional analysis is performed on the smoothed centroid trajectory to analyze its directional changes. This analysis includes extracting adjacent points from the centroid trajectory. The aim of directional analysis is to identify turning points and trend changes in the trajectory, providing a basis for convergence point identification. The analysis process first extracts adjacent point pairs from the smoothed trajectory as the basic unit for directional calculation. Point pair selection considers the curvature characteristics of the trajectory, increasing sampling density in high-curvature regions to improve analysis accuracy.

[0058] Calculate the angle between adjacent points, i.e., the change in direction. Angle calculation is a direct method to quantify changes in trajectory direction, reflecting the stability of the movement trend. The calculation uses the vector angle formula: ; Where v1 and v2 are the direction vectors of adjacent segments, and θ is the included angle. To eliminate the discontinuity caused by the 2π periodicity, cumulative angle change is used to represent and record the overall turning trend of the trajectory. The included angle calculation takes into account the influence of vector length, and the shorter segments are weighted to reduce the interference of local noise. The calculation results form a sequence of direction changes, reflecting the turning characteristics of each point on the trajectory; the smaller the direction change, the more stable the trajectory.

[0059] Based on directional changes, convergence points of the trajectory are identified. Convergence points are regions where the directional change is below a preset threshold, generating the movement trajectory. Convergence point identification is the core step in trajectory analysis, aiming to pinpoint the potential locations of interference sources. The identification process first determines an appropriate change threshold, typically set at 20-30% of the average directional change, to filter out directionally stable regions. Then, local energy maxima are searched within these regions as candidate convergence points. For complex trajectories, cluster analysis is used to identify multiple convergence regions, reflecting potential multi-source interference. The identification algorithm considers both the overall trend and local characteristics of the trajectory, enhancing sensitivity to convergence regions through directional gradient analysis. The finally identified convergence points, together with the original trajectory, constitute a complete representation of the movement trajectory, containing key information such as the starting point, path, and convergence points, providing direct evidence for subsequent interference source localization. The movement trajectory is stored in structured data format, containing point sequences, directional characteristics, and convergence information, comprehensively describing the evolution of the abnormal region in the time-frequency plane.

[0060] In this embodiment of the invention, the detailed implementation steps for constructing dynamic fusion weights based on temporal consistency and distribution stability include: Time series consistency and distribution stability are normalized to generate normalized consistency and normalized stability. Normalization aims to eliminate scale differences between different indicators, making them comparable and combinable. The process uses the Min-Max normalization method to map the original indicator values ​​to the [0,1] interval. The calculation formula is as follows: Normalized value = (original value - minimum value) / (maximum value - minimum value); For timing consistency, considering its inverse property (the smaller the value, the higher the consistency), a reverse mapping is used: Normalization consistency = 1 - (original value - minimum value) / (maximum value - minimum value); The normalization process takes into account the skewness of the index distribution. For non-normally distributed data, logarithmic or power transformation preprocessing is used to improve the normalization effect. The processing result forms a normalized index matrix, where each element corresponds to the normalized score of a mode on a specific index. The closer the value is to 1, the better the performance, providing a standardized basis for subsequent analysis.

[0061] The correlation between normalized consistency and normalized stability is analyzed to generate correlation coefficients. Correlation analysis aims to understand the strength and pattern of the relationship between the two types of indicators, providing a basis for weighting strategies. Pearson correlation coefficient or Spearman rank correlation coefficient are used to quantify the linear or monotonic relationship between indicators. The calculation process considers the characteristic differences of each mode, calculating correlations at both the mode-level and global levels. For nonlinear relationships, techniques such as mutual information analysis or Maximum Information Coefficient are used to capture complex statistical dependencies. The correlation analysis results are represented as a coefficient matrix, where each element reflects the strength of the relationship between consistency and stability for a specific mode. Coefficients close to 1 indicate a high positive correlation, close to -1 indicate a high negative correlation, and close to 0 indicate a weak relationship. These coefficients directly influence subsequent weighting strategies, ensuring that weight allocation fully considers the interaction effects between indicators.

[0062] Based on the correlation coefficient, the weight ratios of normalized consistency and normalized stability are adjusted. Weight adjustment is a crucial step in achieving optimal indicator combination, determining the contribution ratio of each indicator in the final score. The adjustment process employs an adaptive strategy, dynamically determining weight allocation based on correlation strength. For highly correlated cases (|correlation coefficient|>0.7), the total weight is reduced to avoid information redundancy; for low or negative correlation cases, a higher total weight is maintained to preserve complementary information. The adjustment algorithm considers the differences in indicator importance; generally, temporal consistency has a higher weight in dynamic scenarios, while distribution stability has a higher weight in static scenarios. The calculation formula is: Temporal consistency weight = base weight × (1 - α × |correlation coefficient|); Distribution stability weight = basic weight × (1 - α × |correlation coefficient|); Here, α is the adjustment coefficient, typically ranging from 0.3 to 0.5, while the base weights are preset based on the application scenario. Weight adjustment ensures that the final fusion fully utilizes the indicator information while avoiding redundant amplification, thus improving the rationality and adaptability of weight allocation.

[0063] Based on weight ratios, normalization consistency and normalization stability are weighted and combined to generate dynamic fusion weights. This weighted combination is the core step in generating the final score, integrating multiple indicators into a comprehensive metric. The combination employs either a weighted average or geometric average method, selecting an appropriate fusion function based on the indicator characteristics. The calculation formula is as follows: Dynamic fusion weight = w1 × normalized consistency + w2 × normalized stability; or Dynamic fusion weight = (normalized consistency) w1 )×(Normalized stability) w2 ); Where w1 and w2 are the adjusted weight ratios, and w1 + w2 = 1. The combination process considers the complementarity and redundancy between indicators, and adopts a more balanced weight allocation for highly complementary indicator pairs. The final generated dynamic fusion weight is a score value in the range [0,1], reflecting the comprehensive reliability of each modality data at the current moment, and is directly used in the subsequent data fusion process. This dynamic weight construction method based on multiple indicators can adjust the fusion strategy in real time according to data characteristics, significantly improving the adaptability and robustness of the fusion results.

[0064] This invention achieves high-precision detection of target events or states in a monitoring environment through multimodal sensor data acquisition, interference source localization and analysis, interference waveform feature separation, cross-modal interference propagation network construction, noise correlation reduction, and multimodal dynamic fusion analysis. The multimodal IoT sensor data fusion processing method of this invention effectively solves the performance degradation problem of traditional fusion methods in complex noisy environments.

[0065] The above describes the multimodal IoT sensor data fusion processing method in the embodiments of this application. The following describes the multimodal IoT sensor data fusion processing system in the embodiments of this application. Please refer to [link / reference]. Figure 2 One embodiment of the multimodal IoT sensor data fusion processing system in this application includes: The data acquisition module is used to acquire multimodal sensor data from multiple IoT sensors in the monitoring environment at each sampling time. The multimodal sensor data includes sensor data from at least two different modes. The interference source localization module is used to perform interference source localization analysis on the multimodal sensor data at each sampling time and generate an interference source distribution map; The interference feature separation module is used to separate the interference waveform features of each mode sensor data based on the interference source distribution map; The interference propagation network construction module is used to construct a cross-modal interference propagation network by utilizing the characteristics of interference waveforms. The data decoupling processing module is used to reduce the noise correlation of multimodal sensor data based on the cross-modal interference propagation network and obtain decoupled sensor data. The dynamic fusion analysis module is used to perform dynamic fusion analysis on the decoupled sensor data at each sampling time and construct a multimodal fusion feature vector; The target detection module is used to detect target events or target states in the monitoring environment based on multimodal fusion feature vectors. The modules are connected via wired and / or wireless means to enable data transmission between them.

[0066] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0067] It should be noted that all formulas in this manual are calculated by removing dimensions and taking their numerical values. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0068] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for multimodal IoT sensor data fusion processing, characterized in that, include: Acquire multimodal sensor data from multiple IoT sensors in the monitoring environment at each sampling time, wherein the multimodal sensor data includes sensor data from at least two different modes; Interference source localization analysis is performed on the multimodal sensor data at each sampling time to generate an interference source distribution map; Based on the interference source distribution map, the interference waveform characteristics of each mode sensor data are separated; A cross-modal interference propagation network is constructed using the aforementioned interference waveform characteristics; Based on the cross-modal interference propagation network, noise correlation is reduced on the multimodal sensor data to obtain decoupled sensor data; Dynamic fusion analysis is performed on the decoupled sensor data at each sampling time to construct a multimodal fusion feature vector; Based on the multimodal fusion feature vector, target events or target states in the monitoring environment are detected.

2. The multimodal IoT sensor data fusion processing method according to claim 1, characterized in that, The step of performing interference source localization analysis on the multimodal sensor data at each sampling time to generate an interference source distribution map includes: At each sampling time, time-frequency decomposition is performed on the data from each modal sensor to generate a time-frequency distribution map; In the time-frequency distribution map, high-energy anomaly regions are identified, which are regions where the energy value is greater than a preset energy threshold. The high-energy anomaly region is time-series tracked to calculate the movement trajectory of the high-energy anomaly region within multiple consecutive sampling times; Based on the movement trajectory, candidate locations of the interference source are determined, where the candidate locations are the convergence points of the movement trajectories. Cluster analysis is performed on the candidate locations, and candidate locations with a distance less than a preset distance threshold are merged to generate the interference source distribution map, which includes the location and intensity of the interference sources.

3. The multimodal IoT sensor data fusion processing method according to claim 1, characterized in that, The step of separating the interference waveform characteristics of each mode sensor data based on the interference source distribution map includes: Based on the interference source distribution map, the propagation range of each interference source is determined. The propagation range is a circular area centered on the location of the interference source and with a radius equal to the interference intensity. In the time-series data of each modal sensor, extract the local data segment corresponding to the propagation range; The local data segment is subjected to spectral decomposition to separate high-frequency interference components and low-frequency trend components; The high-frequency interference components are reconstructed to generate an interference waveform; The interference waveform is subjected to feature extraction to generate interference waveform features corresponding to the sensor data of each mode. The feature extraction includes calculating the peak frequency, peak amplitude and periodicity index of the waveform.

4. The multimodal IoT sensor data fusion processing method according to claim 1, characterized in that, The method of constructing a cross-modal interference propagation network using the interference waveform characteristics includes: At each sampling time, based on the interference waveform characteristics of each modal sensor data, the interference propagation relationship between any two modal sensor data is analyzed; Calculate the propagation delay of the interference propagation relationship, wherein the calculation of the propagation delay includes comparing the peak time difference of the interference waveform characteristics; Based on the propagation delay, an initial propagation graph is constructed, where the nodes of the initial propagation graph are the data from each modal sensor, and the edges are the propagation delay; A propagation intensity analysis is performed on the initial propagation graph to calculate the propagation intensity of each edge. The calculation of the propagation intensity includes analyzing the peak amplitude ratio and periodicity exponential difference of the interference waveform characteristics. The initial propagation graph is dynamically optimized by removing edges whose propagation intensity is lower than a preset intensity threshold, thereby generating the cross-modal interference propagation network.

5. The multimodal IoT sensor data fusion processing method according to claim 1, characterized in that, The step of reducing noise correlation in the multimodal sensor data based on the cross-modal interference propagation network to obtain decoupled sensor data includes: In the cross-modal interference propagation network, high-impact propagation paths are identified, which are paths whose sum of propagation intensities is greater than a preset path threshold. An impact analysis was performed on the high-impact propagation paths to analyze the impact of sensor data of each mode on interference propagation in each path; Based on the aforementioned impact analysis, an interference suppression matrix is ​​constructed, the construction of which includes a weighted adjustment of the adjacency matrix of the cross-modal interference propagation network. Using the interference suppression matrix, the multimodal sensor data is linearly transformed to generate intermediate decoupled data; Residual analysis is performed on the intermediate decoupled data to remove components in the residuals that are highly correlated with the characteristics of the interference waveform, thereby generating the decoupled sensor data.

6. The multimodal IoT sensor data fusion processing method according to claim 1, characterized in that, The dynamic fusion analysis of the decoupled sensor data at each sampling time to construct a multimodal fusion feature vector includes: The decoupled sensor data at each sampling time is time-series segmented to generate multiple local time segments; Within each local time segment, the temporal consistency of the sensor data of each modality is analyzed. The temporal consistency analysis includes calculating the curvature of the dynamic time warping path of the decoupled sensor data. Distribution analysis is performed on the decoupled sensor data at all sampling times to analyze the distribution stability of the sensor data for each mode. The analysis of distribution stability includes the following steps: Kernel density estimation is performed on the decoupled sensor data to generate a probability density distribution curve; The probability density distribution curve is segmented and analyzed to extract local distribution features within each time period. The local distribution features include peak position and peak width. Analyze the peak shift of the local distribution characteristics within a continuous time period, wherein the peak shift analysis includes calculating the distance between peak positions in adjacent time periods; The shape change of the probability density distribution curve is analyzed, and the analysis of the shape change includes calculating the normalized value of the probability density distribution curve; Based on the aforementioned temporal consistency and distribution stability, a dynamic fusion weight is constructed; The decoupled sensor data is weighted and fused using the dynamic fusion weights to generate the multimodal fusion feature vector.

7. The multimodal IoT sensor data fusion processing method according to claim 1, characterized in that, The detection of target events or target states in the monitoring environment based on the multimodal fusion feature vector includes: Obtain a preset reference feature library of target events or target states; At each sampling time, the dynamic matching degree between the multimodal fusion feature vector and each reference feature in the reference feature library is calculated; Based on the maximum value of the dynamic matching degree, the candidate target event or target state at the corresponding sampling time is determined; The candidate target event or target state at multiple consecutive sampling times is subjected to temporal smoothing processing, which includes constructing a state transition model, eliminating detection results with state transition probabilities lower than a preset probability threshold, and generating the final target event or target state.

8. The multimodal IoT sensor data fusion processing method according to claim 2, characterized in that, The step of time-series tracking of the high-energy anomaly region and calculating the movement trajectory of the high-energy anomaly region over multiple consecutive sampling times includes: In the time-frequency distribution map at multiple consecutive sampling times, the centroid position of the high-energy anomaly region is extracted; Temporal interpolation is performed on the centroid position to generate a continuous centroid trajectory; The centroid trajectory is smoothed to remove noise and jitter. The smoothing process includes the following steps: Calculate the energy intensity of the high-energy anomaly region; Based on the energy intensity, a weighting coefficient is constructed; Using the weighting coefficients, a weighted moving average is performed on the centroid trajectory; The smoothed centroid trajectory is subjected to directional analysis to analyze the changes in trajectory direction. The directional analysis includes the following steps: Extract the adjacent points of the centroid trajectory; Calculate the angle between the adjacent points, i.e., the change in direction; Based on the change in direction, the convergence point of the trajectory is identified. The convergence point is the area where the change in direction is less than a preset change threshold, and the movement trajectory is generated.

9. The multimodal IoT sensor data fusion processing method according to claim 6, characterized in that, The construction of dynamic fusion weights based on the temporal consistency and distribution stability includes: The temporal consistency and the distribution stability are normalized to generate normalized consistency and normalized stability. Analyze the correlation between the normalization consistency and the normalization stability, and generate a correlation coefficient; Based on the correlation coefficient, adjust the weight ratio of the normalization consistency and the normalization stability. Based on the weight ratio, the normalized consistency and the normalized stability are weighted and combined to generate the dynamic fusion weight.

10. A multimodal IoT sensor data fusion processing system, used to implement the multimodal IoT sensor data fusion processing method according to any one of claims 1 to 9, characterized in that, include: The data acquisition module is used to acquire multimodal sensor data from multiple IoT sensors in the monitoring environment at each sampling time. The multimodal sensor data includes sensor data from at least two different modes. The interference source localization module is used to perform interference source localization analysis on the multimodal sensor data at each sampling time and generate an interference source distribution map; The interference feature separation module is used to separate the interference waveform features of each mode sensor data based on the interference source distribution map. An interference propagation network construction module is used to construct a cross-modal interference propagation network using the characteristics of the interference waveform; The data decoupling processing module is used to reduce the noise correlation of the multimodal sensor data based on the cross-modal interference propagation network to obtain decoupled sensor data. The dynamic fusion analysis module is used to perform dynamic fusion analysis on the decoupled sensor data at each sampling time and construct a multimodal fusion feature vector; The target detection module is used to detect target events or target states in the monitoring environment based on the multimodal fusion feature vector. The modules are connected via wired and / or wireless means to enable data transmission between them.

Citation Information

Cited By

  • Multi-source data coupling system and method

    CN121997280A