Intelligent hidden danger diagnosis method based on multi-modal feature fusion
Through the spatiotemporal alignment and feature fusion processing of multimodal data, the problems of multimodal data synchronization and noise processing are solved, and high-precision hidden danger diagnosis and real-time monitoring are achieved, which is suitable for industrial equipment and public facilities.
Patent Information
- Application Number
- CN202511213069.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-09-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The multimodal data processing methods in the existing technology have deficiencies in data synchronization and feature correlation, resulting in low accuracy and efficiency of hidden danger diagnosis, and imperfect noise processing, which affects the reliability of the diagnosis results.
Visual data, voiceprint data and vibration data are collected synchronously through multimodal sensing devices, and spatiotemporal alignment preprocessing, sub-modal feature extraction and cross-modal association encoding are performed. The LSTM network is used for time series synchronization calibration and noise filtering to generate a multimodal association feature set, which is then combined with a deep learning model to identify hidden dangers.
It achieves comprehensive capture of multi-dimensional information, improves the accuracy and efficiency of diagnosis, reduces missed diagnoses and misdiagnoses, and is suitable for real-time hidden danger monitoring and diagnosis of various industrial equipment and public facilities.
Smart Images

Figure CN120727263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent diagnosis of hidden dangers, and specifically to an intelligent diagnosis method of hidden dangers based on multimodal feature fusion. Background Art
[0002] In many fields, including industrial production, equipment operation and maintenance, and public safety, the timely detection and accurate diagnosis of hidden dangers are crucial for ensuring stable system operation and reducing accident risks. Traditional hidden danger diagnosis methods rely heavily on manual inspections or single-sensor monitoring, which presents significant limitations. Manual inspections are not only labor-intensive and resource-intensive, but also limited by the inspectors' experience and responsibility, making them prone to missed and false detections. This makes it difficult to meet the needs of large-scale, high-precision hidden danger diagnosis.
[0003] Monitoring with a single sensor device can only capture data on a single physical characteristic of the target object, such as collecting only image information through a visual device or relying solely on a vibration sensor to obtain vibration signals. However, many hidden dangers often manifest themselves in multiple dimensions, making it difficult to fully reflect the true state of the hazard with a single modal data source. For example, a mechanical equipment failure may be accompanied by abnormal vibration, distinctive sounds, and subtle changes in appearance. Relying on only one type of data for diagnosis is likely to result in inaccurate results due to incomplete information.
[0004] With the development of technologies such as the Internet of Things and artificial intelligence, multimodal data fusion technology has gradually been applied to the field of hidden danger diagnosis. However, current multimodal data processing methods still have shortcomings in terms of data synchronization and feature correlation. Data acquisition devices of different modalities may have time deviations, making it difficult to align the collected multimodal data in the time dimension, affecting the effectiveness of subsequent feature extraction and fusion. At the same time, during the feature processing process, there is often a lack of effective exploration of the inherent correlation between different modal features, making it impossible for the fused features to fully reflect the synergistic advantages of multimodal data, thereby affecting the accuracy and efficiency of hidden danger diagnosis.
[0005] Existing technologies for multimodal data noise processing are inadequate. The presence of noise can interfere with the extraction of effective features and reduce data quality. Furthermore, after feature fusion, the models used for hidden danger identification often lack generalization capabilities and are difficult to adapt to the needs of hidden danger diagnosis in different scenarios, resulting in the reliability of diagnostic results requiring improvement. Therefore, developing an intelligent diagnostic method that can achieve precise alignment and effective fusion of multimodal data, while improving the accuracy and efficiency of hidden danger diagnosis, has become a challenge that needs to be addressed in related fields. Summary of the Invention
[0006] The purpose of the present invention is to provide an intelligent hidden danger diagnosis method based on multimodal feature fusion to solve the problems raised in the above background technology.
[0007] To achieve the above objectives, the present invention provides a method for intelligent diagnosis of hidden dangers based on multimodal feature fusion, the method comprising: Using a multimodal sensing device to synchronously collect visual data, voiceprint data, and vibration data of a target object to form a multimodal raw data set, and transmitting the multimodal raw data set to a diagnostic terminal; At the diagnosis terminal, performing spatiotemporal alignment preprocessing on the multimodal original data set to obtain a multimodal aligned data set; At the diagnostic terminal, performing sub-modal feature extraction on the multimodal aligned data set to obtain a visual feature sequence, a voiceprint feature sequence, and a vibration feature sequence, and performing cross-modal association encoding on the visual feature sequence, the voiceprint feature sequence, and the vibration feature sequence to obtain a multimodal association feature set; At the diagnosis terminal, a hidden danger diagnosis result is outputted through a hidden danger discrimination model according to the multimodal association feature set, and corresponding disposal suggestions are generated based on the hidden danger diagnosis result.
[0008] Preferably, performing spatiotemporal alignment preprocessing on the multimodal original dataset to obtain a multimodal aligned dataset comprises: Extracting a timestamp of each modality data from the multimodal original dataset, and dividing the multimodal original dataset into time periods to obtain a forward time period multimodal dataset and a backward time period multimodal dataset; According to the timestamp mark, respectively performing time synchronization calibration on the forward period multimodal dataset and the backward period multimodal dataset to obtain a forward synchronization multimodal subset and a backward synchronization multimodal subset; Noise components are filtered on the forward synchronized multimodal subset and the backward synchronized multimodal subset to obtain multimodal alignment data subsets, and the multimodal alignment data subsets are combined to form the multimodal alignment data set.
[0009] Preferably, extracting the timestamp of each modality data from the multimodal original data set and dividing the multimodal original data set into time periods to obtain a forward time period multimodal data set and a backward time period multimodal data set includes: Preliminarily identifying the synchronization time node of each modality data in the multimodal original data set to obtain the timestamp mark; The multimodal original data set is divided into time periods using the timestamp as a dividing point to obtain the forward time period multimodal data set and the backward time period multimodal data set.
[0010] Preferably, according to the timestamp mark, performing time synchronization calibration on the forward time period multimodal dataset and the backward time period multimodal dataset respectively to obtain a forward synchronous multimodal subset and a backward synchronous multimodal subset, including: Using the timestamp mark as a reference synchronization point, calculating the timing offset between the reference synchronization point and each data point in the forward time period multimodal dataset to obtain a forward timing offset sequence, and calculating the timing offset between the reference synchronization point and each data point in the backward time period multimodal dataset to obtain a backward timing offset sequence; The forward timing offset sequence and the backward timing offset sequence are respectively subjected to synchronization error compensation encoding to obtain a forward synchronization compensation feature vector as a calibration basis for the forward synchronization multimodal subset and a backward synchronization compensation feature vector as a calibration basis for the backward synchronization multimodal subset.
[0011] Preferably, performing synchronization error compensation encoding on the forward timing offset sequence and the backward timing offset sequence respectively to obtain a forward synchronization compensation feature vector and a backward synchronization compensation feature vector comprises: Inputting the forward timing offset sequence into a forward synchronization compensation encoder based on a forward LSTM network to obtain the forward synchronization compensation feature vector; The backward timing offset sequence is input into a backward synchronization compensation encoder based on a backward LSTM network to obtain the backward synchronization compensation feature vector.
[0012] Preferably, filtering the forward synchronized multimodal subset and the backward synchronized multimodal subset for noise components to obtain multimodal alignment data subsets, and combining the multimodal alignment data subsets to form the multimodal alignment dataset comprises: Performing noise feature detection on the forward synchronous multimodal subset and the backward synchronous multimodal subset respectively to obtain a forward noise feature set and a backward noise feature set; performing noise removal processing on the forward synchronous multimodal subset according to the forward noise feature set to obtain a forward aligned data subset, and performing noise removal processing on the backward synchronous multimodal subset according to the backward noise feature set to obtain a backward aligned data subset; The forward aligned data subset and the backward aligned data subset are spliced in time sequence to form the multimodal aligned data set.
[0013] Preferably, performing noise feature detection on the forward synchronous multimodal subset and the backward synchronous multimodal subset respectively to obtain a forward noise feature set and a backward noise feature set comprises: performing outlier detection on the forward synchronous multimodal subset to extract forward anomaly features, and performing noise pattern matching on the forward anomaly features to obtain the forward noise feature set; Outlier detection is performed on the backward synchronous multimodal subset to extract backward anomaly features, and noise pattern matching is performed on the backward anomaly features to obtain the backward noise feature set.
[0014] Preferably, performing noise removal processing on the forward synchronized multimodal subset according to the forward noise feature set to obtain a forward aligned data subset comprises: Performing a feature difference operation on the forward synchronous multimodal subset and the forward noise feature set to obtain a forward difference feature sequence; An effective signal preservation process is performed on the forward differential feature sequence to obtain the forward alignment data subset.
[0015] Preferably, the forward aligned data subset and the backward aligned data subset are temporally spliced to form the multimodal aligned data set, comprising: Extracting the end time mark of the forward aligned data subset and the start time mark of the backward aligned data subset, and calculating the time interval between the two; Performing a time shift adjustment on the backward aligned data subset according to the time interval so that a start time stamp of the backward aligned data subset is continuous with an end time stamp of the forward aligned data subset; The adjusted backward aligned data subset is appended to the forward aligned data subset to form the multimodal aligned data set.
[0016] Preferably, performing sub-modal feature extraction on the multimodal aligned dataset to obtain a visual feature sequence, a voiceprint feature sequence, and a vibration feature sequence includes: Performing image edge detection and texture analysis on the visual data in the multimodal alignment dataset to extract local visual features, and performing sequence integration of the local visual features in a time dimension to obtain the visual feature sequence; Performing frequency domain conversion and energy spectrum analysis on the voiceprint data in the multimodal alignment dataset to extract voiceprint spectrum features, and performing sliding aggregation of the voiceprint spectrum features in a time window to obtain the voiceprint feature sequence; Acceleration signal decomposition and amplitude statistics are performed on the vibration data in the multimodal alignment data set to extract vibration modal features, and time sequence continuity verification is performed on the vibration modal features to obtain the vibration feature sequence.
[0017] Compared with the prior art, the present invention has the following beneficial effects: This method uses multimodal sensing equipment to synchronously collect the visual data, voiceprint data and vibration data of the target object, which can comprehensively capture the multi-dimensional information of the target object. Compared with the traditional single-modal data diagnosis method, it greatly enriches the diagnostic basis, avoids missed diagnosis and misdiagnosis caused by missing information, and lays a solid data foundation for subsequent accurate diagnosis.
[0018] During the data preprocessing phase, a spatiotemporal alignment preprocessing step was implemented. This effectively addressed the temporal asynchrony of multimodal data by extracting timestamps, dividing the data into time periods, performing time series synchronization calibration, and filtering noise. Timestamp extraction and time period division enabled data to be processed segmented based on synchronized time nodes, improving the specificity of data processing. Time series synchronization calibration ensured temporal consistency of multimodal data within different time periods by calculating time series offsets and performing error compensation coding. Noise filtering further purified the data, reducing noise interference with subsequent feature extraction and significantly improving the quality of the multimodal alignment dataset.
[0019] In the feature extraction and fusion stages, submodal feature extraction uses corresponding extraction methods based on the characteristics of visual, voiceprint, and vibration data, accurately capturing key features of each modality, such as image edge and texture information in visual feature sequences, frequency and energy spectrum features in voiceprint feature sequences, and acceleration and amplitude features in vibration feature sequences. Cross-modal correlation coding further explores the inherent connections between different modal features, integrating dispersed single-modal features into a highly correlated multimodal correlation feature set. This fully leverages the synergistic advantages of multimodal data, enabling the fused features to more comprehensively reflect the state of the target object and provide more effective feature support for hidden danger identification.
[0020] By using a hidden danger identification model to output diagnostic results based on a multimodal correlation feature set and generate corresponding treatment recommendations, this not only improves the automation and intelligence level of hidden danger diagnosis and reduces reliance on manual experience, but also quickly provides targeted treatment plans, helping relevant personnel take timely measures to eliminate hidden dangers, reduce the risk of accidents, and improve the safety and stability of system operations. Furthermore, the entire method completes data processing and diagnosis at the diagnostic terminal, achieving high processing efficiency and meeting the needs of real-time diagnosis. It is suitable for hidden danger monitoring and diagnosis in a variety of industrial equipment, public facilities, and other scenarios, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a diagram showing the working principle of the intelligent hidden danger diagnosis method based on multimodal feature fusion according to the present invention; Figure 2 Flowchart of preprocessing for spatiotemporal alignment; Figure 3Flowchart for timing synchronization calibration; Figure 4 Flowchart for noise removal processing; Figure 5 This is a flowchart of time sequence splicing. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0023] See also Figure 1-Figure 5 The present invention provides a hidden danger intelligent diagnosis method based on multimodal feature fusion, and the specific implementation steps are as follows: Multimodal sensing devices are used to synchronously collect visual data, voiceprint data, and vibration data of the target object to form a multimodal raw data set, which is then transmitted to the diagnostic terminal. Multimodal sensing devices include but are not limited to high-definition cameras, microphone arrays, and vibration sensors. Each device uses a unified clock synchronization mechanism to ensure the time consistency of data collection.
[0024] At the diagnosis terminal, the multimodal original dataset is preprocessed for spatiotemporal alignment to obtain a multimodal aligned dataset.
[0025] At the diagnostic terminal, sub-modal feature extraction is performed on the multimodal aligned dataset to obtain visual feature sequences, voiceprint feature sequences and vibration feature sequences, and the visual feature sequences, voiceprint feature sequences and vibration feature sequences are cross-modally associated encoded to obtain a multimodal associated feature set.
[0026] At the diagnostic terminal, a hidden danger identification model outputs a hidden danger diagnosis result based on the multimodal correlation feature set, and generates corresponding treatment recommendations based on the hidden danger diagnosis results. The hidden danger identification model uses a deep learning architecture trained on historically annotated data. It analyzes and determines the multimodal correlation feature set, outputs a specific diagnosis result, including the hidden danger type and severity, and generates corresponding treatment recommendations based on pre-set treatment rules.
[0027] Example 1: In the process of performing spatiotemporal alignment preprocessing on the multimodal original data set to obtain a multimodal aligned data set, it is necessary to complete the following key operations in sequence: extracting timestamps and time period division, timing synchronization calibration, and noise component filtering.
[0028] The timestamps of each modal data set are extracted from the multimodal raw dataset, and the dataset is divided into time periods to produce a forward-period multimodal dataset and a backward-period multimodal dataset. The multimodal raw dataset contains visual data, voiceprint data, and vibration data, all of which are timestamped at the time of acquisition. The synchronization time nodes of each modal data set are initially identified, for example, by detecting the start time of data acquisition or a specific synchronization signal to determine the timestamp. The synchronization time node here refers to the key time point at which the modal data can be aligned in time, which can be a specific moment or a shorter time range. Using this timestamp as the demarcation point, the multimodal raw dataset is divided into a forward-period multimodal dataset and a backward-period multimodal dataset. The forward-period multimodal dataset contains all data before the timestamp, while the backward-period multimodal dataset contains all data after the timestamp. This division establishes a clear temporal boundary between the two time periods, providing a temporal benchmark for subsequent processing.
[0029] According to the timestamp mark, the forward period multimodal dataset and the backward period multimodal dataset are respectively subjected to time synchronization calibration to obtain the forward synchronized multimodal subset and the backward synchronized multimodal subset. The timestamp mark is used as the reference synchronization point. For each data point in the forward period multimodal dataset, the difference between its timestamp and the timestamp of the reference synchronization point is calculated to obtain the forward time series offset sequence; similarly, for each data point in the backward period multimodal dataset, the difference between its timestamp and the timestamp of the reference synchronization point is calculated to obtain the backward time series offset sequence. These time series offset sequences reflect the time offset of each data point relative to the reference synchronization point. Since there may be time deviations in the acquisition of different modal data, these offsets need to be compensated.
[0030] Synchronization error compensation encoding is performed on the forward and backward timing offset sequences, respectively. The forward timing offset sequence is input into a forward synchronization compensation encoder based on a forward LSTM network. The forward LSTM network is a neural network model capable of processing time series data. It has a memory function and can learn the temporal dependencies in the timing offset sequence. The forward LSTM network calculates and processes the forward timing offset sequence, outputting a forward synchronization compensation feature vector. This vector contains the information required for synchronization calibration of the forward period data and serves as the basis for calibrating the forward synchronization multimodal subset. Similarly, the backward timing offset sequence is input into a backward synchronization compensation encoder based on a backward LSTM network. The backward LSTM network processes the timing offset sequence from back to front, capturing the reverse temporal dependencies in the sequence. This processing outputs a backward synchronization compensation feature vector, which serves as the basis for calibrating the backward synchronization multimodal subset.
[0031] After completing the timing synchronization calibration, the noise components of the forward synchronization multimodal subset and the backward synchronization multimodal subset are filtered to obtain a multimodal alignment data subset, and the multimodal alignment data subsets are combined to form a multimodal alignment data set. First, noise feature detection is performed on the forward synchronization multimodal subset and the backward synchronization multimodal subset respectively. Outlier detection is performed on the forward synchronization multimodal subset. A variety of methods can be used for outlier detection, such as statistical methods or machine learning algorithms. These methods identify data points in the forward synchronization multimodal subset that deviate from the normal range, and extract the features of these outliers as forward anomaly features. Then, the forward anomaly features are compared with a preset noise pattern library. The noise pattern library stores the characteristic patterns of various common noises. The noise type corresponding to the forward anomaly features is determined by matching, thereby obtaining a forward noise feature set. Similarly, outlier detection is performed on the backward synchronization multimodal subset, the backward anomaly features are extracted, and noise pattern matching is performed to obtain a backward noise feature set.
[0032] The forward-synchronized multimodal subset is subjected to noise removal processing based on the forward noise feature set. A feature differential operation is performed on the forward-synchronized multimodal subset and the forward noise feature set, and the feature difference between the two is calculated to obtain a forward differential feature sequence. In this differential feature sequence, the noise features are weakened, while the features of the valid signal are retained. The forward differential feature sequence is subjected to valid signal preservation processing, such as using an appropriate filtering method to remove the noise component, to obtain a forward-aligned data subset. Similarly, the backward-synchronized multimodal subset is subjected to noise removal processing based on the backward noise feature set to obtain a backward-aligned data subset.
[0033] The forward-aligned data subset and the backward-aligned data subset are spliced in time sequence to form a multimodal alignment dataset. The end time stamp of the forward-aligned data subset and the start time stamp of the backward-aligned data subset are extracted, and the time interval between them is calculated. If a time interval exists, the backward-aligned data subset is time-shifted according to the time interval so that the start time stamp of the backward-aligned data subset is temporally continuous with the end time stamp of the forward-aligned data subset. After the time shift adjustment is completed, the adjusted backward-aligned data subset is attached to the forward-aligned data subset to form a complete multimodal alignment dataset.
[0034] Example 2: In the process of spatiotemporal alignment preprocessing of multimodal raw data sets, extracting the timestamp mark of each modal data from the multimodal raw data set and dividing the multimodal raw data set into time periods to obtain a forward time period multimodal data set and a backward time period multimodal data set are important foundations for subsequent processing. When the visual data, voiceprint data, and vibration data in the multimodal raw data set are collected, they are all assigned corresponding timestamps by the multimodal sensing device. These timestamps record the specific moment of data collection. By analyzing the timestamps of each modal data, the synchronization time nodes of each modal data are preliminarily identified. The synchronization time nodes here can be the same or similar moments in the timestamps of each modal data, such as finding the moment within the same second that the timestamps of each modal data first appear, or detecting the common starting point of the collection of each modal data through a specific algorithm to obtain a timestamp mark.
[0035] The multimodal original data set is divided into two parts using the timestamp as the dividing point. Among them, all data with a timestamp earlier than the marker constitute the forward period multimodal data set, and all data with a timestamp later than the marker constitute the backward period multimodal data set. This division method makes the data sets of the forward period and the backward period have clear boundaries in time, and can provide a clear time range definition for subsequent processing such as timing synchronization calibration and noise filtering. For example, assuming that the timestamp is marked as time T0, then the forward period multimodal data set contains all data collected before time T0, and the backward period multimodal data set contains all data collected after time T0. This division ensures that the data of the two periods are non-overlapping and continuous on the time axis.
[0036] When performing timing synchronization calibration on the forward period multimodal dataset and the backward period multimodal dataset based on the timestamp mark, the timestamp mark is used as the reference synchronization point. For each data point in the forward period multimodal dataset, the difference between its timestamp and the timestamp of the reference synchronization point is calculated to obtain a forward timing offset sequence. Specifically, the timestamp of each data point is Ti, and the timestamp of the reference synchronization point is T0, then the timing offset of the data point is Ti-T0, and the timing offsets of all the forward period data points are arranged in sequence to form a forward timing offset sequence. Similarly, for each data point in the backward period multimodal dataset, the difference between its timestamp and the timestamp of the reference synchronization point is calculated to obtain a backward timing offset sequence, that is, the timing offset of each backward data point is Ti-T0, and they are arranged in sequence to form a backward timing offset sequence.
[0037] Due to clock skew between the devices collecting data from different modalities or delays in data transmission, the timestamps of each modal data may differ from the actual acquisition time. These timing offset sequences reflect the time offsets of each data point relative to the reference synchronization point. To compensate for these offsets, the forward timing offset sequence is input into a forward synchronization compensation encoder based on a forward LSTM network. The forward LSTM network is a recurrent neural network designed to effectively process time series data. It uses memory cells to store historical information, thereby capturing the temporal dependencies in the timing offset sequence. The forward LSTM network processes each element of the forward timing offset sequence in chronological order, calculating the offset using the network's internal weight parameters and activation function. It ultimately outputs a forward synchronization compensation feature vector. This vector contains the information necessary for synchronization calibration of the forward period data and serves as the basis for calibrating the forward synchronized multimodal subset.
[0038] Similarly, the backward timing offset sequence is input into a backward synchronization compensation encoder based on a backward LSTM network. The backward LSTM network processes the timing offset sequence from backward to forward, in the opposite direction of the forward LSTM network. This allows it to capture reverse temporal dependencies within the sequence. Starting from the last timing offset, the backward LSTM network processes each element forward in sequence. Through network calculation and processing, it outputs a backward synchronization compensation feature vector, which serves as the basis for calibrating the backward synchronization multimodal subset.
[0039] When filtering noise components in the forward-synchronized multimodal subset and the backward-synchronized multimodal subset, outlier detection is first performed on the forward-synchronized multimodal subset. Various methods can be used for outlier detection, such as the Z-score method, which identifies outliers based on the mean and standard deviation of the data, treating data points that deviate from the mean by a certain standard deviation as outliers; or the IQR method, which identifies outliers by calculating quartiles and interquartile ranges. These methods identify data points in the forward-synchronized multimodal subset that deviate from the normal range, and extract features of these outliers, such as amplitude and frequency, as forward anomaly features.
[0040] The forward anomaly signature is then matched against patterns in a pre-established noise pattern library. This library stores characteristic patterns of various common noises, such as electromagnetic interference noise signatures during equipment operation and environmental background noise, derived through analysis and summary of extensive historical noise data. By comparing the forward anomaly signature with the patterns in the noise pattern library, the noise type corresponding to the forward anomaly signature is determined, thereby generating a forward noise feature set. Similarly, outlier detection is performed on the backward synchronized multimodal subset, and backward anomaly signatures are extracted. Noise pattern matching is then performed to generate a backward noise feature set.
[0041] The forward synchronous multimodal subset is subjected to noise removal processing according to the forward noise feature set. A feature difference operation is performed on the forward synchronous multimodal subset and the forward noise feature set, that is, the difference between the feature of each data point in the forward synchronous multimodal subset and the corresponding noise feature of the forward noise feature set is calculated to obtain a forward differential feature sequence. In this differential feature sequence, the features of the noise are weakened, while the features of the effective signal are highlighted. The forward differential feature sequence is subjected to effective signal retention processing, such as using a wavelet denoising method. By selecting an appropriate wavelet basis and number of decomposition layers, the differential feature sequence is decomposed and reconstructed to remove the noise component and obtain a forward aligned data subset. Similarly, the backward synchronous multimodal subset is processed according to the backward noise feature set to obtain a backward aligned data subset.
[0042] When combining the forward-aligned data subset and the backward-aligned data subset into a multimodal alignment dataset, the end time stamp of the forward-aligned data subset and the start time stamp of the backward-aligned data subset are extracted, and the time interval between the two is calculated. Assuming that the end time stamp of the forward-aligned data subset is T1 and the start time stamp of the backward-aligned data subset is T2, then the time interval is T2-T1. If the time interval is not zero, the backward-aligned data subset is time-shifted according to the interval, and the timestamps of all data points in the backward-aligned data subset are subtracted by the time interval so that the start time stamp of the backward-aligned data subset and the end time stamp of the forward-aligned data subset are temporally continuous, that is, the start time stamp of the adjusted backward-aligned data subset becomes T1. After the time shift adjustment is completed, the adjusted backward-aligned data subset is placed after the forward-aligned data subset and arranged in chronological order to form a complete multimodal alignment dataset.
[0043] Example 3: In the process of extracting sub-modal features from the multimodal alignment dataset to obtain visual feature sequences, voiceprint feature sequences, and vibration feature sequences, it is necessary to adopt corresponding processing methods based on the characteristics of different modal data, as follows: To process visual data, we first perform image edge detection on the visual data in the multimodal alignment dataset. Image edge detection uses a specific algorithm to identify areas within an image where pixel grayscale changes significantly. The commonly used Canny operator for edge detection can be represented as follows: Gaussian filtering is performed on the image to remove noise, gradient magnitude and direction are calculated, non-maximum suppression is performed, and finally, a dual-threshold algorithm is used to determine and connect edges. Edge detection can highlight the outlines and boundaries of objects in an image, providing a foundation for subsequent analysis.
[0044] After edge detection, texture analysis is performed on the visual data. Texture analysis aims to describe the spatial distribution pattern of pixel grayscale in an image. This can be achieved by calculating texture feature parameters such as contrast, correlation, energy, and entropy. Contrast reflects the degree of difference in grayscale between adjacent pixels in an image. The calculation formula is: ; in, represents the number of gray levels, is the position in the gray-level co-occurrence matrix The formula measures the contrast of the image by calculating the relationship between the elements and positions in the gray level co-occurrence matrix.
[0045] Through edge detection and texture analysis, local visual features can be extracted. In order to integrate these local features in the time dimension and form a visual feature sequence with time series characteristics, it is necessary to collect and organize the local visual features at each time point. Specifically, for each time point , combine the edge features and texture features of the image at that moment to form a feature vector , and then arrange these feature vectors into sequences in chronological order , thus obtaining a visual feature sequence, which reflects the changes in the visual features of the target object at different time points.
[0046] For voiceprint data processing, the voiceprint data in the multimodal alignment dataset is first converted to the frequency domain. The commonly used frequency domain conversion method is Fourier transform, which is based on the principle of decomposing the time domain signal into a superposition of sine waves of different frequencies. For discrete voiceprint signals, , its discrete Fourier transform can be expressed as: ; in, is the signal length, represents the frequency index, is an imaginary unit. Through Fourier transform, the voiceprint signal in the time domain can be converted to the frequency domain to obtain the spectrum distribution of the voiceprint, making the frequency components of the signal more clearly visible.
[0047] Based on the frequency domain conversion, the energy spectrum analysis is performed on the voiceprint data. Energy spectrum analysis is to calculate the energy distribution of different frequency components. , its energy spectrum It can be expressed as , which reflects the energy of the corresponding frequency component. Through energy spectrum analysis, the voiceprint spectrum characteristics can be extracted, such as the frequency band where energy is concentrated and the main frequency components.
[0048] In order to aggregate these spectral features in time and form a voiceprint feature sequence, a sliding aggregation method of a time window is used. Set a suitable time window size and sliding step length , slide the window on the time axis. For each voiceprint spectrum feature in the window, calculate its mean, variance and other statistics, for example, for the first The corresponding spectral feature mean can be expressed as: ; in, For the In the window The energy value of each frequency component. The aggregation result of each window is used as the voiceprint feature at that time point, and these features are arranged in chronological order to obtain the voiceprint feature sequence. , this sequence reflects the temporal variation characteristics of the voiceprint signal.
[0049] For vibration data processing, the acceleration signal of the vibration data in the multi-modal alignment dataset is first decomposed. Acceleration signal decomposition can use spectrum analysis or modal analysis to decompose the complex vibration acceleration signal into different frequency components or modes. For example, the acceleration signal in the time domain is transformed into Convert to frequency domain and get its spectrum , thereby determining the main frequency components of the vibration signal and the corresponding amplitude .
[0050] After completing the acceleration signal decomposition, the vibration data is subjected to amplitude statistics. Amplitude statistics is to calculate the amplitude of the vibration signal at different times and perform statistical analysis, such as calculating the mean, maximum, minimum, etc. of the amplitude. , in the time period The mean within can be expressed as: ; Vibration modal characteristics can be extracted through amplitude statistics, such as the amplitude of each frequency component and the range of amplitude variation.
[0051] In order to ensure the temporal continuity of the vibration feature sequence, the temporal continuity check is performed on the vibration modal features. The temporal continuity check is to check whether there are sudden changes or unreasonable jumps in the vibration features of adjacent time points. Specifically, for adjacent time points and , calculate its vibration eigenvector and Difference , can be calculated using Euclidean distance: ; in, is the dimension of the feature vector, and Time points and No. The value of the feature. Set a threshold ,when When the feature at this time point is considered to have abnormal jumps, it needs to be corrected or smoothed, such as using the weighted average of adjacent features to replace the abnormal feature, thereby obtaining the vibration feature sequence ,This sequence can accurately reflect the change of the vibration state of the target object over time.
[0052] Through the above submodal feature extraction and processing of visual, voiceprint and vibration data, corresponding feature sequences are obtained. These sequences describe the characteristics of the target object from different modalities, providing rich feature information for subsequent cross-modal association coding and hidden danger diagnosis.
[0053] Example 4: Extracting timestamps from each modality and dividing them into time periods are crucial foundational steps when preprocessing multimodal raw datasets for spatiotemporal alignment. When collecting visual, voiceprint, and vibration data from a target object, multimodal sensing devices add precise timestamps to each piece of data. These timestamps can be generated based on the device's internal clock or an external synchronized clock. For example, in an industrial equipment monitoring scenario, high-definition cameras, microphone arrays, and vibration sensors collect data simultaneously. Each device records data with a time stamp, such as 10:15:23 AM on July 1, 2025.
[0054] By traversing the multimodal raw dataset, extracting the timestamp information for each modal data, and then finding a common reference point among the timestamps of each modality. For example, assuming the first timestamp of the visual data is 2025-07-01 10:15:23.001, the first timestamp of the voiceprint data is 2025-07-01 10:15:23.003, and the first timestamp of the vibration data is 2025-07-01 10:15:23.002, then a close moment among these three timestamps, such as 10:15:23 on July 1, 2025, can be used as a timestamp marker. Using this marker as the boundary, the dataset is divided into a forward period and a backward period multimodal dataset. The forward period multimodal dataset contains all data collected before 10:15:23 on July 1, 2025, while the backward period multimodal dataset contains all data collected after that moment. This division ensures that the two period datasets have a clear temporal order, facilitating subsequent processing.
[0055] When performing timing synchronization calibration, the timestamp mark is used as the benchmark to calculate the timing offset of each data point in the forward time period multimodal dataset from the benchmark. For example, the timestamp of a visual data point in the forward time period is 2025-07-0110:15:22.500, and the benchmark timestamp is 2025-07-0110:15:23.000. Then the timing offset of this data point is -0.5 seconds. The timing offsets of all data points in the forward time period are arranged in order to form a forward timing offset sequence. Similarly, the timing offsets of the data points in the backward time period are calculated. For example, if the timestamp of a backward voiceprint data point is 2025-07-0110:15:23.500, its timing offset is +0.5 seconds, and the backward timing offset sequence is obtained.
[0056] To compensate for these offsets, a forward LSTM network and a backward LSTM network are used for encoding. The forward LSTM network processes the forward time series offset sequence in chronological order, such as processing the offsets of -0.5 seconds, -0.3 seconds, and -0.1 seconds at each time point. It uses memory cells to store historical information, capture the temporal dependencies in the sequence, and output a forward synchronization compensation feature vector, which contains the correction information required to synchronize the forward time period data. The backward LSTM network processes the backward time series offset sequence in reverse chronological order, such as starting from +0.5 seconds, +0.3 seconds, and +0.1 seconds. It obtains clues from future information and outputs a backward synchronization compensation feature vector for synchronization calibration of the backward time period data.
[0057] During the noise filtering phase, outlier detection is performed on both the forward and backward synchronized multimodal subsets. For example, in the forward synchronized multimodal subset, if the amplitude of a vibration data point is significantly higher than that of other data points, a threshold is set, such as a point exceeding three standard deviations from the mean, to be considered an outlier. This outlier is identified and its features extracted. For example, if the amplitude is 20g, while the normal range is between 0-5g, this feature is used as the forward anomaly feature. The forward anomaly feature is then matched with patterns in a noise pattern library, which stores noise features generated by loose equipment during operation and noise features from environmental electromagnetic interference. Assuming that this anomaly feature matches the noise feature generated by loose equipment, the forward noise feature set is derived as the equipment looseness noise feature. Similarly, outlier detection is performed on the backward synchronized multimodal subset, and the backward anomaly feature is extracted. Noise pattern matching is then performed to obtain the backward noise feature set.
[0058] The forward-synchronized multimodal subset is subjected to noise removal based on the forward noise feature set. Feature differential operations are performed on the forward-synchronized multimodal subset and the forward noise feature set, for example, by calculating the difference between the amplitude of each vibration data point and the amplitude of the equipment loose noise feature, to obtain a forward differential feature sequence. This forward differential feature sequence is then subjected to effective signal-retention processing, such as median filtering, to remove noise components and obtain a forward-aligned data subset. Similarly, the backward-synchronized multimodal subset is processed based on the backward noise feature set to obtain a backward-aligned data subset.
[0059] Finally, the forward-aligned data subset and the backward-aligned data subset are spliced in time sequence. The end time stamp of the forward-aligned data subset is extracted, such as 2025-07-0110:15:23.000, and the start time stamp of the backward-aligned data subset is 2025-07-0110:15:23.100, with a calculation time interval of 0.1 seconds. According to this interval, the backward-aligned data subset is time-shifted and adjusted, and the timestamps of all data in the backward-aligned data subset are subtracted by 0.1 seconds, so that its start time stamp becomes 2025-07-0110:15:23.000, which is continuous with the end time stamp of the forward-aligned data subset. The adjusted backward-aligned data subset is connected to the forward-aligned data subset to form a complete multimodal alignment data set. After this processing, the modal data in the multimodal alignment data set are aligned in time, and noise interference is removed, providing reliable data for subsequent operations such as submodal feature extraction.
[0060] Example 5: When extracting sub-modal features from a multi-modal aligned dataset, differentiated processing flows are required based on the characteristics of different modal data. Taking visual, voiceprint, and vibration data as an example, the specific implementation methods are as follows: To process visual data, image edge detection is first required. Taking monitoring images of industrial equipment as an example, the image is processed using the Canny or Sobel operators to identify areas within the image where pixel grayscale changes significantly, such as the outlines of equipment components and the edges of seams. Edge detection can separate the structural features of the equipment from the background, forming clear contour lines. Texture analysis is then performed to describe the spatial distribution pattern of pixels in the image by calculating texture feature parameters. For example, for texture features such as rust and wear on the surface of equipment, parameters such as contrast and energy are calculated. Contrast measures the grayscale difference between adjacent pixels, while energy reflects the uniformity of the texture. Through edge detection and texture analysis, local visual features such as edge contour coordinates and texture parameter values can be extracted.
[0061] To form a visual feature sequence, the local features at each time point must be integrated in chronological order. Suppose the coordinate sequence of the equipment bolt edge detected at time t1 is (x1, y1), (x2, y2), and so on. The contrast of the texture in this area is calculated as C1 and the energy is E1. The edge coordinates and texture parameters at time t2 are updated to the new values. By combining the edge and texture features at each time point into a feature vector and arranging them in the order t1, t2, t3, and so on, we can obtain a sequence that reflects the changes in the equipment's visual features over time, such as the shift in edge position when a bolt loosens or the evolution of surface texture wear.
[0062] For voiceprint data, frequency domain conversion is first performed. Taking the sound of a running motor as an example, Fourier transform is used to convert the time-domain voiceprint signal into a frequency-domain signal, revealing the energy distribution of the sound at different frequency components. For example, during normal operation, the motor's voiceprint spectrum has strong energy around 50Hz, while anomalies may cause new energy peaks at 100Hz or higher. Energy spectrum analysis is performed in the frequency domain to calculate the energy value of each frequency component and extract spectral features such as energy peak frequency and bandwidth.
[0063] The voiceprint feature sequence is generated by sliding aggregation of time windows. The time window size is set to 100ms, the sliding step is 50ms, and the spectrum features in each window are counted. For example, in the first window, the energy mean of the 50Hz frequency is calculated to be , the average energy at 100Hz is ; After the next window slides, we get 、 The statistical results of each window are arranged in chronological order to form a soundprint feature sequence, which can reflect the dynamic changes in the sound frequency characteristics when the motor is running, such as the gradual increase in high-frequency energy when the bearing wears.
[0064] For vibration data, the acceleration signal is first decomposed. Taking vibration monitoring of rotating machinery as an example, spectrum analysis is used to decompose the time-domain acceleration signal into vibration components of different frequencies. For example, under normal conditions, vibration energy is primarily concentrated at the rotor rotation frequency (e.g., 20 Hz). However, when an imbalance fault occurs, significant energy may appear at twice the rotation frequency (40 Hz). Amplitude statistics are then performed on the decomposed signal, calculating the amplitude, mean, and maximum values of each frequency component.
[0065] In order to ensure the temporal continuity of the vibration feature sequence, the features need to be verified. For example, the vibration feature vectors at adjacent time points t and t+1 are (Including 20Hz amplitude , 40Hz amplitude )and . Calculate the difference between the two, if Comparison If the value suddenly increases by 50% and exceeds the preset threshold, it is judged that the feature may have a mutation due to noise or abnormal interference, and it needs to be smoothed by weighted averaging of adjacent features to ensure that the changes in the vibration characteristics in the sequence conform to the physical laws of equipment operation, such as the gradual increase in vibration amplitude when an unbalanced fault develops.
[0066] After obtaining the visual, voiceprint, and vibration feature sequences, cross-modal correlation coding is required. For example, when a bolt position offset is detected in the visual feature sequence, the voiceprint feature sequence is simultaneously checked to see if there is an increase in high-frequency energy and an increase in the 40Hz amplitude in the vibration feature sequence. The modal features are aligned by timestamps to establish correlation between the features. Bolt loosening may be accompanied by an increase in vibration amplitude and the appearance of abnormal noise. Through this correlation coding, a feature set containing multi-modal feature correlation relationships is formed, providing comprehensive input information for subsequent hidden danger diagnosis models, such as comprehensively judging the degree of bolt loosening and whether it requires treatment.
[0067] During the entire sub-modal feature extraction process, each link is based on the physical meaning of the data and the principles of signal processing. Through step-by-step processing of visual, voiceprint, and vibration data, the conversion from raw data to feature sequences is achieved, ultimately providing multi-dimensional feature support for hidden danger diagnosis, ensuring that each modal feature can not only independently reflect the equipment status, but also reveal potential hidden danger association patterns through cross-modal associations.
[0068] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0069] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for intelligent diagnosis of hidden dangers based on multimodal feature fusion, characterized in that: The method comprises: Using a multimodal sensing device to synchronously collect visual data, voiceprint data, and vibration data of a target object to form a multimodal raw data set, and transmitting the multimodal raw data set to a diagnostic terminal; At the diagnosis terminal, performing spatiotemporal alignment preprocessing on the multimodal original data set to obtain a multimodal aligned data set; At the diagnostic terminal, performing sub-modal feature extraction on the multimodal aligned data set to obtain a visual feature sequence, a voiceprint feature sequence, and a vibration feature sequence, and performing cross-modal association encoding on the visual feature sequence, the voiceprint feature sequence, and the vibration feature sequence to obtain a multimodal association feature set; At the diagnosis terminal, a hidden danger diagnosis result is outputted through a hidden danger discrimination model according to the multimodal association feature set, and corresponding disposal suggestions are generated based on the hidden danger diagnosis result.
2. The method for intelligent diagnosis of hidden dangers based on multimodal feature fusion according to claim 1 is characterized in that: Performing spatiotemporal alignment preprocessing on the multimodal original dataset to obtain a multimodal aligned dataset includes: Extracting a timestamp of each modality data from the multimodal original dataset, and dividing the multimodal original dataset into time periods to obtain a forward time period multimodal dataset and a backward time period multimodal dataset; According to the timestamp mark, respectively performing time synchronization calibration on the forward period multimodal dataset and the backward period multimodal dataset to obtain a forward synchronization multimodal subset and a backward synchronization multimodal subset; Noise components are filtered on the forward synchronized multimodal subset and the backward synchronized multimodal subset to obtain multimodal alignment data subsets, and the multimodal alignment data subsets are combined to form the multimodal alignment data set.
3. The method for intelligent diagnosis of hidden dangers based on multimodal feature fusion according to claim 2 is characterized in that: Extracting a timestamp of each modality data from the multimodal original data set, and dividing the multimodal original data set into time periods to obtain a forward time period multimodal data set and a backward time period multimodal data set, including: Preliminarily identifying the synchronization time node of each modality data in the multimodal original data set to obtain the timestamp mark; The multimodal original data set is divided into time periods using the timestamp as a dividing point to obtain the forward time period multimodal data set and the backward time period multimodal data set.
4. The method for intelligent diagnosis of hidden dangers based on multimodal feature fusion according to claim 3 is characterized in that: According to the timestamp mark, respectively performing time synchronization calibration on the forward period multimodal dataset and the backward period multimodal dataset to obtain a forward synchronous multimodal subset and a backward synchronous multimodal subset, comprising: Using the timestamp mark as a reference synchronization point, calculating the timing offset between the reference synchronization point and each data point in the forward time period multimodal dataset to obtain a forward timing offset sequence, and calculating the timing offset between the reference synchronization point and each data point in the backward time period multimodal dataset to obtain a backward timing offset sequence; The forward timing offset sequence and the backward timing offset sequence are respectively subjected to synchronization error compensation encoding to obtain a forward synchronization compensation feature vector as a calibration basis for the forward synchronization multimodal subset and a backward synchronization compensation feature vector as a calibration basis for the backward synchronization multimodal subset.
5. The method for intelligent diagnosis of hidden dangers based on multimodal feature fusion according to claim 4 is characterized in that: Performing synchronization error compensation encoding on the forward timing offset sequence and the backward timing offset sequence to obtain a forward synchronization compensation feature vector and a backward synchronization compensation feature vector, respectively, includes: Inputting the forward timing offset sequence into a forward synchronization compensation encoder based on a forward LSTM network to obtain the forward synchronization compensation feature vector; The backward timing offset sequence is input into a backward synchronization compensation encoder based on a backward LSTM network to obtain the backward synchronization compensation feature vector.
6. The method for intelligent diagnosis of hidden dangers based on multimodal feature fusion according to claim 5 is characterized in that: Filtering the forward synchronized multimodal subset and the backward synchronized multimodal subset for noise components to obtain multimodal alignment data subsets, and combining the multimodal alignment data subsets to form the multimodal alignment dataset, comprising: Performing noise feature detection on the forward synchronous multimodal subset and the backward synchronous multimodal subset respectively to obtain a forward noise feature set and a backward noise feature set; performing noise removal processing on the forward synchronous multimodal subset according to the forward noise feature set to obtain a forward aligned data subset, and performing noise removal processing on the backward synchronous multimodal subset according to the backward noise feature set to obtain a backward aligned data subset; The forward aligned data subset and the backward aligned data subset are spliced in time sequence to form the multimodal aligned data set.
7. The method for intelligent diagnosis of hidden dangers based on multimodal feature fusion according to claim 6 is characterized in that: Performing noise feature detection on the forward synchronous multimodal subset and the backward synchronous multimodal subset respectively to obtain a forward noise feature set and a backward noise feature set, comprising: performing outlier detection on the forward synchronous multimodal subset to extract forward anomaly features, and performing noise pattern matching on the forward anomaly features to obtain the forward noise feature set; Outlier detection is performed on the backward synchronous multimodal subset to extract backward anomaly features, and noise pattern matching is performed on the backward anomaly features to obtain the backward noise feature set.
8. The method for intelligent diagnosis of hidden dangers based on multimodal feature fusion according to claim 7 is characterized in that: Performing noise removal processing on the forward synchronized multimodal subset according to the forward noise feature set to obtain a forward aligned data subset, comprising: Performing a feature difference operation on the forward synchronous multimodal subset and the forward noise feature set to obtain a forward difference feature sequence; An effective signal preservation process is performed on the forward differential feature sequence to obtain the forward alignment data subset.
9. The method for intelligent diagnosis of hidden dangers based on multimodal feature fusion according to claim 8, characterized in that: Temporally splicing the forward aligned data subset and the backward aligned data subset to form the multimodal aligned data set, comprising: Extracting the end time mark of the forward aligned data subset and the start time mark of the backward aligned data subset, and calculating the time interval between the two; Performing a time shift adjustment on the backward aligned data subset according to the time interval so that a start time stamp of the backward aligned data subset is continuous with an end time stamp of the forward aligned data subset; The adjusted backward aligned data subset is appended to the forward aligned data subset to form the multimodal aligned data set.
10. The method for intelligent diagnosis of hidden dangers based on multimodal feature fusion according to claim 1, characterized in that: Performing sub-modal feature extraction on the multimodal aligned dataset to obtain a visual feature sequence, a voiceprint feature sequence, and a vibration feature sequence, including: Performing image edge detection and texture analysis on the visual data in the multimodal alignment dataset to extract local visual features, and performing sequence integration of the local visual features in a time dimension to obtain the visual feature sequence; Performing frequency domain conversion and energy spectrum analysis on the voiceprint data in the multimodal alignment dataset to extract voiceprint spectrum features, and performing sliding aggregation of the voiceprint spectrum features in a time window to obtain the voiceprint feature sequence; Acceleration signal decomposition and amplitude statistics are performed on the vibration data in the multimodal alignment data set to extract vibration modal features, and time sequence continuity verification is performed on the vibration modal features to obtain the vibration feature sequence.
Citation Information
Patent Citations
Electromechanical equipment multi-modal data fusion fault diagnosis method and device
CN118395231A
Industrial intelligent detection method and system based on multi-modal large model
CN118503832A
Remote online monitoring method and system based on machine vision and artificial intelligence
CN120105312A
Cited By
Hydroelectric generating set rotor acoustic diagnosis method and device based on space-time joint feature map and medium
CN121051656A
Hydroelectric generator rotor acoustic diagnosis method and device based on spatio-temporal joint feature map and medium
CN121051656B
Quality evaluation method and system based on senile syndrome data
CN121565501A
Equipment defect detection method and device, equipment and medium
CN121682463A
Automatic production and processing system for degradable meal boxes
CN122058475A