A Fusion Method and System for Automatic Processing of Multi-Source Heterogeneous Data Based on Convolutional Neural Network
Through the method based on convolutional neural network, spatial mapping relationship is established, high-order features of multimodal data are extracted and fusion, which solves the problems of time-consuming feature extraction, difficulty in data alignment and difficulty in deep correlation capture in traditional methods, and achieves efficient multimodal data fusion and prediction accuracy improvement.
Patent Information
- Application Number
- CN202510240927.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The traditional multimodal data fusion method has problems such as manual feature extraction, difficulty in data alignment, and difficulty in capturing multimodal deep correlation, resulting in poor fusion effect.
Using a method based on convolutional neural network, a spatial mapping relationship is established by acquiring sensor data and image data, the sensor data is organized into two-dimensional data, the resonance region is determined using the LSTM model, higher-order features are extracted and convolutional operations are fused, and fusion features are generated by weighting processing of the resonance region.
It realizes efficient extraction and integration of high-order features of multimodal data, solves the problems of data alignment and deep correlation capture, and improves prediction accuracy and robustness.
Smart Images

Figure CN119720109B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing, and more particularly relates to a fusion method and system for automatic processing of multi-source heterogeneous data based on a convolutional neural network. Background Art
[0002] In fields such as complex environmental monitoring, intelligent transportation, and industrial equipment fault diagnosis, it is usually necessary to process multi-source heterogeneous data from multiple sensors and vision devices. These data come from diverse sources, including sensors for temperature, humidity, vibration, etc. installed at different locations, as well as image or video data collected by cameras. With the increase in the number and diversity of data sources, how to efficiently integrate different modal data and make full use of their respective advantages has become the key to improving prediction accuracy and system robustness.
[0003] Traditional multi-modal data fusion methods usually rely on manual feature extraction and rule formulation, and such methods have the following problems:
[0004] Traditional manual feature extraction methods rely on the knowledge of domain experts, which is not only time-consuming and laborious, but also prone to missing potential key features, resulting in low data utilization rate and affecting model performance.
[0005] Sensor data and image data usually have inconsistencies in space. How to accurately align different modal data to achieve effective fusion is a major challenge faced by traditional methods.
[0006] Due to significant differences in structure, scale, and dimension between sensor and image data, traditional fusion methods are difficult to capture the deep associations of multi-modal data, resulting in poor fusion effects.
[0007] Traditional methods often ignore the dynamic interaction relationships between multi-modal features during the feature fusion process of multi-modal data, resulting in the model being unable to fully explore the potential patterns and associations in the data. Summary of the Invention
[0008] To solve the problems in the prior art, the present invention provides a fusion method for automatic processing of multi-source heterogeneous data based on a convolutional neural network. The method includes the following steps:
[0009] Obtain sensor data at different positions and obtain image data including the sensor installation positions; establish a spatial mapping relationship according to the spatial correspondence between the sensor installation positions and the image feature regions; organize the sensor data into sensor two-dimensional data according to the mapping relationship; use an LSTM model to determine the resonance region between the sensor two-dimensional data and the image data; use a first two-dimensional convolutional neural network to extract the first high-order features of the image data, and use a second two-dimensional convolutional neural network to extract the second high-order features of the sensor two-dimensional data; fuse the first high-order features and the second high-order features using a convolutional operation to obtain a first fusion feature; determine the fusion feature based on the resonance region and the first fusion feature.
[0010] Further, the establishing of the spatial mapping relationship includes:
[0011] Obtain the physical installation positions of each sensor, including geographic coordinates, relative height, and direction angle information, to form three-dimensional coordinates;
[0012] Collect image data and obtain the pixel coordinates of the image feature regions;
[0013] Specifically, the following formula is used for projective transformation:
[0014]
[0015] where are the horizontal and vertical pixel coordinates of the image, are the coordinates in the three directions of the sensor three-dimensional space; is the internal parameter matrix of the camera, is the rotation matrix, representing the rotation attitude of the camera in three-dimensional space;
[0016] Determine the corresponding pixel regions of each sensor in the image through projective transformation, realize the spatial alignment of the sensor data and the image feature regions, and generate a spatial mapping relationship table.
[0017] Further, the organizing of the sensor data into sensor two-dimensional data according to the mapping relationship includes:
[0018] According to the spatial mapping relationship, map the time series data of each sensor to the corresponding image plane coordinates;
[0019] Construct a two-dimensional matrix, the size of which is consistent with the spatial distribution of the image feature regions, and each element corresponds to the sensor data at a position;
[0020] Fill the measurement data at each sensor position into the corresponding pixel positions to form a two-dimensional representation;
[0021] The interpolation method is used to complete the pixel positions without filled data.
[0022] Further, using the LSTM model to determine the resonance region between the two-dimensional data of the sensor and the image data includes:
[0023] The resonance region is specifically the region where the sensor data and the image data change significantly simultaneously in time and space;
[0024] The two-dimensional data of the sensor is regarded as time-series features, where the two-dimensional matrix at each time point represents the spatial distribution of the sensor at the corresponding time;
[0025] The image data is extracted in the form of a time series;
[0026] A dual-input LSTM model is constructed to process the two-dimensional data of the sensor and the time-series features of the image in parallel. Through the LSTM, time-related spatial features are extracted, and through another LSTM, time-related image features are extracted;
[0027] The sensor features and the image features are subjected to time synchronization alignment and similarity analysis;
[0028] The similarity between the sensor features and the image features is calculated in the time series. When the similarity exceeds a preset threshold, the corresponding time point is determined as the time resonance point;
[0029] For the time resonance point, further identify the co-activation region of the sensor data and the image features in the two-dimensional space;
[0030] The two-dimensional features of the sensor and the image features are spatially matched, and the Hadamard product is used for calculation to obtain the resonance intensity,
[0031]
[0032] where represents the horizontal pixel position, represents the vertical pixel position, represents the time parameter, represents the Hadamard product, represents the resonance intensity, represents the sensor features, represents the image features; by setting the resonance threshold, the positions in the resonance intensity greater than the resonance threshold are marked as spatial resonance points, thereby determining the resonance region in the two-dimensional space.
[0033] Further, based on the resonance region and the first fusion feature, determining the fusion feature includes:
[0034] The high-order features of the image data are extracted by the first two-dimensional convolutional neural network to obtain the high-order image features;
[0035] The high-order features of the two-dimensional sensor data are extracted by the second two-dimensional convolutional neural network to obtain the high-order sensor features;
[0036] Align the high-order image features and the high-order sensor features in the spatial dimension;
[0037] Concatenate the high-order image features and the high-order sensor features in the channel dimension to form a new joint feature matrix, the concatenated feature;
[0038] Use a new two-dimensional convolutional layer to perform a convolution operation on the concatenated feature, so as to extract the deep interaction pattern between the multi-modal features;
[0039] After the convolution operation, apply a non-linear activation function to enhance the expression ability of the features and retain the fused feature information with positive activation;
[0040] The result obtained after the convolution operation is the first fused feature;
[0041] Generate a binary mask matrix, the mask, of the same size as the first fused feature to mark the position of the resonance region;
[0042] Multiply the mask matrix pixel by pixel with the first fused feature to obtain a weighted feature matrix, the weighted fused feature;
[0043] Perform non-linear activation processing on the weighted fused feature and perform normalization processing on the weighted feature, and finally output the fused feature.
[0044] On the other hand, the present invention also provides a fusion system for automatic processing of multi-source heterogeneous data based on a convolutional neural network. The system includes the following modules:
[0045] A data acquisition module for acquiring sensor data at different positions and acquiring image data including the sensor installation position;
[0046] A mapping module for establishing a spatial mapping relationship according to the spatial correspondence between the sensor installation position and the image feature region;
[0047] A processing module for organizing the sensor data into two-dimensional sensor data according to the mapping relationship;
[0048] A detection module for using an LSTM model to determine the resonance region between the two-dimensional sensor data and the image data;
[0049] An extraction module, configured to extract first high-order features of the image data by using a first two-dimensional convolutional neural network, and extract second high-order features of the two-dimensional sensor data by using a second two-dimensional convolutional neural network;
[0050] A fusion module, configured to fuse the first high-order features and the second high-order features through a convolution operation to obtain first fused features; and determine fused features based on the resonance region and the first fused features.
[0051] Further, the establishing of the spatial mapping relationship includes:
[0052] Obtaining the physical installation positions of each sensor, including geographical coordinates, relative height, and direction angle information, to form three-dimensional coordinates;
[0053] Collecting image data and obtaining pixel coordinates of the image feature region;
[0054] Specifically, the following formula is used for projective transformation:
[0055] ,
[0056] where are the horizontal and vertical pixel coordinates of the image, are the coordinates in three directions of the three-dimensional space of the sensor; is the internal parameter matrix of the camera, is the rotation matrix, indicating the rotation attitude of the camera in the three-dimensional space;
[0057] Determining the corresponding pixel region of each sensor in the image through projective transformation, realizing the spatial alignment of the sensor data and the image feature region, and generating a spatial mapping relationship table.
[0058] Further, the organizing of the sensor data into two-dimensional sensor data according to the mapping relationship includes:
[0059] Mapping the time series data of each sensor to the corresponding image plane coordinates according to the spatial mapping relationship;
[0060] Constructing a two-dimensional matrix, the size of which is consistent with the spatial distribution of the image feature region, and each element corresponds to the sensor data at a position;
[0061] Filling the measurement data at each sensor position into the corresponding pixel position to form a two-dimensional representation;
[0062] Using an interpolation method to fill in the pixel positions without filled data.
[0063] Further, using an LSTM model to determine the resonance region between the two-dimensional sensor data and the image data includes:
[0064] The resonance region specifically refers to the region where significant changes occur simultaneously in the sensor data and the image data in terms of time and space;
[0065] The two-dimensional data of the sensor is regarded as time-series features, where the two-dimensional matrix at each time point represents the spatial distribution of the sensor at the corresponding time;
[0066] The image data is extracted in the form of time series;
[0067] Construct a dual-input LSTM model for parallel processing of the sensor two-dimensional data and the image time-series features. Extract time-related spatial features through one LSTM and time-related image features through another LSTM;
[0068] Perform time synchronization alignment and similarity analysis on the sensor features and the image features;
[0069] Calculate the similarity between the sensor features and the image features in the time series. When the similarity exceeds a preset threshold, determine the corresponding time point as the time resonance point;
[0070] For the time resonance point, further identify the co-activation region of the sensor data and the image features in the two-dimensional space;
[0071] After spatial matching of the sensor two-dimensional features and the image features, calculate the resonance intensity using the Hadamard product,
[0072] ,
[0073] where represents the horizontal pixel position, represents the vertical pixel position, represents the time parameter, represents the Hadamard product, represents the resonance intensity, represents the sensor features, represents the image features; By setting a resonance threshold, mark the positions in the resonance intensity that are greater than the resonance threshold as spatial resonance points, thereby determining the resonance region in the two-dimensional space.
[0074] Furthermore, based on the resonance region and the first fusion feature, it is determined that the fusion feature includes:
[0075] The high-order features of the image data extracted by the first two-dimensional convolutional neural network, obtaining the image high-order features;
[0076] The high-order features of the sensor two-dimensional data extracted by the second two-dimensional convolutional neural network, obtaining the sensor high-order features;
[0077] Align the high - order image features and the high - order sensor features in the spatial dimension;
[0078] Concatenate the high - order image features and the high - order sensor features in the channel dimension to form a new joint feature matrix, i.e., the concatenated feature;
[0079] Use a new two - dimensional convolutional layer to perform a convolution operation on the concatenated feature, so as to extract the deep interaction patterns between multi - modal features;
[0080] After the convolution operation, apply a non - linear activation function to enhance the feature representation ability and retain the fused feature information with positive activation;
[0081] The result obtained after the convolution operation is the first fused feature;
[0082] Generate a binary mask matrix, i.e., the mask, with the same size as the first fused feature to mark the position of the resonance region;
[0083] Multiply the mask matrix with the first fused feature pixel - by - pixel to obtain a weighted feature matrix, i.e., the weighted fused feature;
[0084] Perform non - linear activation processing on the weighted fused feature and normalize the weighted feature, and finally output the fused feature.
[0085] The fusion method for automatic processing of multi - source heterogeneous data based on convolutional neural network proposed by the present invention provides the following beneficial effects by making full use of the characteristics of multi - modal data:
[0086] By using the automatic feature extraction capabilities of convolutional neural network (CNN) and long short - term memory network (LSTM), the present invention can efficiently extract and integrate high - order features in sensor data and image data, avoid the low - efficiency problem of manual feature extraction in traditional methods, and greatly improve the fusion efficiency.
[0087] By establishing the spatial correspondence relationship between the sensor installation position and the image feature region, the present invention can accurately align different - modal data in space, enabling multi - modal data to perform feature fusion under a unified spatial coordinate system, and effectively solving the problem of spatial inconsistency of multi - modal data in traditional methods.
[0088] Use the LSTM model to identify the "resonance region" of sensor data and image features, i.e., the co - activation region of multi - modal data in time and space, so as to better capture the key patterns in the data and improve the mining effect of the correlation of multi - modal data.
[0089] Through convolution operations, the high-order features of sensors and images are deeply fused, and combined with weighted processing of the resonance region, enabling the model to more accurately identify and predict change trends and abnormal conditions in complex environments, improving the prediction accuracy and robustness of the system.
[0090] The method of the present invention has a high degree of automation, can adaptively adjust the fusion strategy according to the characteristics of different modality data without human intervention, and is applicable to a variety of complex practical application scenarios, such as disaster monitoring, intelligent transportation, industrial equipment fault diagnosis, etc.
[0091] In summary, through multi-modal deep fusion, the present invention not only solves the deficiencies of traditional methods in feature extraction, spatial alignment, heterogeneous data integration, etc., but also significantly improves the accuracy and real-time performance of data processing, having broad application value and innovative significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0093] Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0094] Next, with reference to the drawings and specific embodiments, a preferred description of the invention will be given.
[0095] The present embodiment solves the above problems through the following steps:
[0096] In one embodiment, referring to Figure 1 , the present invention provides a fusion method for automatic processing of multi-source heterogeneous data based on a convolutional neural network. This method uses a convolutional neural network to perform feature extraction and fusion processing on sensor data and image data from different positions, realizing automatic alignment and deep interaction of multi-source heterogeneous data, and specifically includes the following steps:
[0097] Obtain sensor data at different positions, and obtain image data including the installation positions of the sensors.
[0098] Obtain sensor data collected by sensors installed at multiple different locations. The sensors include temperature sensors, humidity sensors, vibration sensors, or gas sensors, and are arranged at different physical locations, such as different parts of industrial equipment, different floors or areas of building structures, different geographical locations in the outdoor environment, etc. The sensor data includes numerical information of physical quantities such as temperature values, humidity values, vibration intensities, or gas concentrations at the corresponding locations, and these data are collected in the form of time series or at regular sampling intervals.
[0099] Obtain image data corresponding to the sensor installation locations. The image data is collected by cameras or monitoring devices installed near these locations, and the image data includes the coverage area of the sensor installation locations. For example, when a sensor is installed at the bearing of an industrial equipment, the image data is a device monitoring image including the bearing and its surrounding environment; when a sensor is installed on a certain floor of a building, the image data is a video frame or image including the internal environment of that floor; when a sensor is arranged at a certain geographical location, the image data is a ground monitoring image or an aerial drone image covering that geographical location and its surrounding areas.
[0100] The image data not only contains spatial location information corresponding to the sensor installation locations, but also includes environmental feature information of these locations, such as visual features like the operating state of equipment, the appearance of building structures, the activities of people in the environment, the occurrence of flames or smoke, etc. By obtaining the above-mentioned sensor data and image data, the acquisition of multi-modal data at different locations is realized, providing data support for subsequent multi-modal data alignment, feature extraction, and fusion.
[0101] Establish a spatial mapping relationship according to the spatial correspondence between the sensor installation locations and the image feature regions.
[0102] In this step, according to the spatial correspondence between the installation locations of the sensors and the image feature regions, a spatial mapping relationship of multi-modal data is constructed. The spatial mapping relationship is used to align and match the sensor data and the image data in the same coordinate system. Specifically, by measuring physical parameters such as the geographical location, installation height, and orientation of the sensors, the absolute position of the sensors in three-dimensional space is determined and converted into pixel coordinates in the image data, generating a spatial mapping model between the sensor data and the image feature regions.
[0103] The method for establishing the mapping relationship is as follows:
[0104] Obtain the physical installation positions of each sensor, including information such as geographical coordinates, relative height, and direction angle, to form three-dimensional coordinates.
[0105] Collect image data through cameras or other imaging devices to obtain the pixel coordinates of the image feature regions.
[0106] The projection transformation is specifically carried out using the following formula:
[0107] ,
[0108] where are the horizontal and vertical pixel coordinates of the image, are the coordinates in the three directions of the sensor's three-dimensional space; is the internal parameter matrix of the camera, is the rotation matrix, representing the rotation posture of the camera in three-dimensional space.
[0109] By means of projection transformation, the corresponding pixel area of each sensor in the image is determined, realizing the spatial alignment of sensor data and the image feature area, and generating a spatial mapping relationship table.
[0110] For example, when a temperature sensor is installed at the bearing of an industrial device, by measuring its specific physical coordinates, the corresponding pixel area of this position in the device monitoring image is determined, thereby establishing a spatial mapping relationship between the sensor data and the image features of this area; when a humidity sensor is installed on the ceiling of a building floor, by measuring its floor position and relative height, the feature area of this position in the indoor monitoring image is determined, and the value of the humidity sensor is aligned with the image features through spatial mapping; when a vibration sensor is arranged at a certain pier of an outdoor bridge, its position in the bridge structure monitoring image is obtained through GPS positioning information, and a spatial mapping relationship between the sensor data and the pier image features is established.
[0111] Through the above spatial mapping relationship, the physical coordinates of the sensor data are aligned with the pixel coordinates of the image feature area, so as to maintain spatial consistency in the process of multi-modal data fusion, ensuring the feature interaction and integration of different modal data within the same spatial range.
[0112] Organize the sensor data into sensor two-dimensional data according to the mapping relationship.
[0113] In this step, according to the spatial correspondence relationship between the installation position of the sensor and the image feature area, a spatial mapping relationship of multi-modal data is constructed, and the spatial mapping relationship is used to realize the alignment and matching of sensor data and image data in the same coordinate system. By measuring the physical parameters such as the geographical location, installation height, and orientation of the sensor, the absolute position of the sensor in three-dimensional space is determined, and it is converted into pixel coordinates in the image data, generating a spatial mapping model between the sensor data and the image feature area.
[0114] Collect physical quantity data from sensors installed at multiple different locations, including temperature, humidity, vibration, or other measurement data, and the data of each sensor is represented in the form of a time series.
[0115] According to the previously established spatial mapping relationship, map the time series data of each sensor to the corresponding image plane coordinates.
[0116] Construct a two-dimensional matrix, the size of which is consistent with the spatial distribution of the image feature region, and each element corresponds to the sensor data at a specific position.
[0117] Fill the measurement data at each sensor position into the corresponding pixel position to form a two-dimensional representation.
[0118] If there is no corresponding sensor data at some image pixel positions, interpolation methods (such as bilinear interpolation or Gaussian filling, etc.) can be used to complete the two-dimensional matrix.
[0119] Use the LSTM model to determine the resonance region between the two-dimensional sensor data and the image data.
[0120] The resonance region refers to the region where the sensor data and the image data change significantly simultaneously in time and space, that is, the multimodal data shows similar responses to a certain event or change within a specific time period. Such a region often reflects the high correlation between the data and is the most critical part in multimodal fusion, which can reveal the synchronous responses of different modalities to the same event.
[0121] The present invention determines resonance from two dimensions of time and space.
[0122] Resonance in the time dimension: The sensor and image data have similar fluctuation patterns in the time series. For example, when the measured value of the sensor changes significantly within a certain time period, corresponding changes are also detected in the corresponding image frame.
[0123] Resonance in the space dimension: The two-dimensional data of the sensor and the image features have overlapping activation regions in space, indicating that the two modalities have similar spatial responses in this region.
[0124] Regard the two-dimensional data of the sensor as time series features, where the two-dimensional matrix at each time point represents the spatial distribution of the sensor at the corresponding time.
[0125] Extract the image data in the form of a time series, which can be the time series of image frames or specific image features (such as high-order features extracted by a convolutional neural network), to form a corresponding feature matrix.
[0126] Construct a dual-input LSTM model for parallel processing of two-dimensional sensor data and image time-series features. Extract time-related spatial features through LSTM and time-related image features through another LSTM.
[0127] Perform time synchronization alignment and similarity analysis on sensor features and image features to determine the common activation time period between the two in time.
[0128] Calculate the similarity between sensor features and image features on the time series. This can be done through measurement methods such as cosine similarity or correlation coefficient. When the similarity exceeds a preset threshold, determine that time point as a time resonance point.
[0129] For time resonance points, further identify the common activation regions of sensor data and image features in two-dimensional space.
[0130] The two-dimensional sensor features and image features are spatially matched, and the resonance intensity is calculated using the Hadamard product.
[0131] ,
[0132] where represents the horizontal pixel position, represents the vertical pixel position, represents the time parameter, represents the Hadamard product, represents the resonance intensity, represents the sensor features, represents the image features. By setting a resonance threshold, mark the positions in the resonance intensity that are greater than the resonance threshold as spatial resonance points, thereby determining the resonance region in two-dimensional space.
[0133] Combine the resonance results in time and space to form the final resonance region. This resonance region can reflect the common response pattern of sensors and images in time and space and is an important basis for subsequent multi-modal feature fusion and decision-making.
[0134] Use a first two-dimensional convolutional neural network to extract the first high-order features of the image data, and use a second two-dimensional convolutional neural network to extract the second high-order features of the two-dimensional sensor data.
[0135] The image data is an input two-dimensional matrix representing the visual information around the sensor installation position. It may be a static image, a single frame of a video frame, or a preprocessed feature matrix. The input image data contains spatial distribution information, has a certain width and height, and usually has multiple color channels, such as red, green, and blue (RGB).
[0136] The first two-dimensional convolutional neural network (CNN1) is used to extract high-order features of image data. This network consists of multiple convolutional layers, pooling layers, and activation functions, extracting the spatial features of the image layer by layer.
[0137] Convolutional layer: Each convolutional layer scans a local area of the image through several convolutional kernels, extracting the feature patterns within that area. These features may be simple edges, textures, or higher-level complex image features.
[0138] Pooling layer: After the convolutional layer, a pooling layer is added to downsample the extracted features. The pooling layer reduces the spatial resolution of the features by selecting the maximum or average value of the local area, while retaining important information, improving computational efficiency and the robustness of the features.
[0139] Multiple convolutional and pooling operations: Through multiple convolutional and pooling operations, CNN1 gradually extracts higher-level image features, enabling the output features to represent the deep information of the image, such as shapes, contours, or texture features of specific regions.
[0140] Finally, after multiple convolutional and pooling operations, CNN1 outputs the high-order feature matrix of the image. This feature matrix retains the deep spatial features of the image and compresses the size of the original image, making it easier to process without losing important information.
[0141] The two-dimensional sensor data represents the measurement results of the sensor in space, such as the spatial distribution of numerical values such as temperature, humidity, and vibration. This data is organized into a two-dimensional matrix consistent with the image space to ensure spatial alignment.
[0142] The second two-dimensional convolutional neural network (CNN2) is used to extract high-order features of the two-dimensional sensor data. The structure of CNN2 is similar to that of CNN1, consisting of multiple convolutional layers, pooling layers, and activation functions, extracting the spatial patterns in the sensor data layer by layer.
[0143] Convolutional layer: Each convolutional layer scans a local area of the sensor data to identify the feature patterns in different spatial regions. For example, in temperature sensor data, the convolutional layer may identify patterns such as high-temperature regions and temperature gradients.
[0144] Pooling layer: After the convolutional layer, a pooling layer is added to downsample the extracted features. The pooling layer reduces the spatial dimension by selecting the local maximum or average value, retaining the main features in the sensor data.
[0145] Multiple convolutional and pooling operations: Through multiple convolutional and pooling operations, CNN2 gradually extracts higher-level sensor features, enabling the output features to reflect the deep patterns in the sensor data.
[0146] Finally, after multiple layers of convolution and pooling, CNN2 outputs a high-order feature matrix of the two-dimensional sensor data. This feature matrix reflects the deep spatial distribution of the sensor data and provides a basis for subsequent fusion with the high-order features of the image.
[0147] Fuse the first high-order feature and the second high-order feature using a convolution operation to obtain a first fused feature; based on the resonance region and the first fused feature, determine the fused feature. Specifically, it includes the following steps:
[0148] Preliminary Convolutional Fusion of High-Order Features
[0149] Input High-Order Features
[0150] The first high-order feature: The high-order feature of the image data extracted by the first two-dimensional convolutional neural network, denoted as "image high-order feature", which contains the deep spatial patterns and information of the image.
[0151] The second high-order feature: The high-order feature of the two-dimensional sensor data extracted by the second two-dimensional convolutional neural network, denoted as "sensor high-order feature", which reflects the spatial distribution and deep patterns of the sensor data.
[0152] Feature Alignment
[0153] Before fusion, ensure that the "image high-order feature" and the "sensor high-order feature" are aligned in the spatial dimension, that is, they have the same height and width. This ensures that the two can be fused pixel by pixel in space.
[0154] Feature Concatenation and Convolutional Fusion
[0155] Concatenate the "image high-order feature" and the "sensor high-order feature" in the channel dimension to form a new joint feature matrix "concatenated feature", which contains multi-level feature information of the two modalities.
[0156] Use a new two-dimensional convolutional layer to perform a convolution operation on the "concatenated feature" to extract the deep interaction patterns between the multi-modal features.
[0157] This convolutional layer captures the correlation patterns between the image features and the sensor features in space and channel by scanning the local regions of the "concatenated feature", and extracts the common information and the patterns of mutual influence between them.
[0158] After the convolution operation, apply a non-linear activation function (such as ReLU or Sigmoid) to enhance the expressive power of the features and retain the fused feature information of the positive activation.
[0159] Output of the First Fused Feature
[0160] The result obtained after the convolution operation is called the "first fusion feature", which contains the preliminary fusion information of the image data and sensor data in terms of space and channels.
[0161] This feature matrix represents the interaction pattern of multimodal data in the same space and serves as the basis for further fusion and analysis.
[0162] Determination of the fusion feature based on the resonance region
[0163] The resonance region is the synchronous activation region of multimodal data identified by the LSTM model in terms of time and space, representing the common response region of the image data and sensor data within a specific time period.
[0164] This region reflects the similar activation of the image and sensor data in space and is usually represented by a mask matrix of the same size as the "first fusion feature".
[0165] Generation of the mask for the resonance region
[0166] Generate a binary mask matrix "mask" of the same size as the "first fusion feature" to mark the position of the resonance region.
[0167] The pixel positions with a value of 1 in the mask matrix indicate belonging to the resonance region, and the pixel positions with a value of 0 indicate not belonging to the resonance region.
[0168] The generation of the mask is based on the detection result of the resonance region, ensuring alignment with the "first fusion feature" in space.
[0169] Weighting process and mask application
[0170] Multiply the mask matrix pixel by pixel with the "first fusion feature" to obtain the weighted feature matrix "weighted fusion feature".
[0171] The features within the resonance region are retained and amplified, while the features in the non-resonance region are weakened or ignored. This can highlight the strong correlation features of multimodal data in the common activation region.
[0172] This weighting process ensures the strengthening of the feature response in the resonance region in space, enabling the final fusion feature to more intensively reflect the common features of multimodal data.
[0173] Nonlinear activation and normalization
[0174] Perform nonlinear activation processing on the "weighted fusion feature", such as ReLU or Sigmoid, to further enhance the feature expression ability.
[0175] After that, the weighted features are normalized (such as batch normalization or layer normalization) to ensure the balance and stability of the fused features across different channels and spatial positions.
[0176] Output of the fused features
[0177] The finally output feature matrix is called the "fused feature", which not only contains the deep interaction information of the image and sensor features, but also highlights the influence of the resonance region.
[0178] This fused feature represents the joint response of the image data and sensor data in time and space, which is the final result of multi-modal fusion and can be used for subsequent classification, detection or prediction tasks.
[0179] On the other hand, the present invention also provides a fusion system for automatic processing of multi-source heterogeneous data based on a convolutional neural network, including:
[0180] A data acquisition module for acquiring sensor data at different positions and acquiring image data including the sensor installation position;
[0181] A mapping module for establishing a spatial mapping relationship according to the spatial correspondence between the sensor installation position and the image feature region;
[0182] A processing module for organizing the sensor data into two-dimensional sensor data according to the mapping relationship;
[0183] A detection module for determining the resonance region between the two-dimensional sensor data and the image data using an LSTM model;
[0184] An extraction module for extracting the first high-order features of the image data using a first two-dimensional convolutional neural network and extracting the second high-order features of the two-dimensional sensor data using a second two-dimensional convolutional neural network;
[0185] A fusion module for fusing the first high-order features and the second high-order features using a convolutional operation to obtain a first fused feature; determining a fused feature based on the resonance region and the first fused feature.
[0186] For the part of the module structure not specifically defined in the present invention, the content recorded in the prior art shall prevail. The prior art mentioned in the foregoing background art part and the specific embodiment part of the present invention can be used as a part of the present invention to understand the meaning of some technical features or parameters.
Claims
1. A fusion method for automatic processing of multi-source heterogeneous data based on convolutional neural network, characterized in that: The method comprises the following steps: Acquire sensor data at different positions, and acquire image data including the sensor installation position; According to the spatial correspondence between the sensor installation position and the image feature area, a spatial mapping relationship is established; Organizing the sensor data into sensor two-dimensional data according to the mapping relationship; Determine the resonance area between the two-dimensional sensor data and the image data using an LSTM model; Extracting first high-order features of the image data using a first two-dimensional convolutional neural network, and extracting second high-order features of the two-dimensional sensor data using a second two-dimensional convolutional neural network; The first high-order feature and the second high-order feature are fused using a convolution operation to obtain a first fused feature; based on the resonance area and the first fused feature, a fused feature is determined; The resonance region is specifically a region where sensor data and image data change significantly in both time and space; The two-dimensional data of the sensor is regarded as a time series feature, where the two-dimensional matrix at each time point represents the spatial distribution of the sensor at the corresponding time; Extract image data into time series form; Build a dual-input LSTM model to process the sensor two-dimensional data and image time series features in parallel. Use LSTM to extract time-related spatial features, and use another LSTM to extract time-related image features. Temporally align sensor features and image features and conduct similarity analysis; The similarity between the sensor features and the image features is calculated in the time series. When the similarity exceeds a preset threshold, the corresponding time point is determined to be a time resonance point. For the time resonance point, the common activation area of sensor data and image features in two-dimensional space is further identified; The two-dimensional features of the sensor and the image features are spatially matched, and the resonance intensity is calculated using the Hadamard product. C(u,v,t)=F s (u,v,t)⊙F i (u,v,t) Where u represents the horizontal position of the pixel, v represents the vertical position of the pixel, t represents the time parameter, ⊙ represents the Hadamard product, C(u,v,t) represents the resonance intensity, and F s (u,v,t) represents the sensor characteristics, F i (u, v, t) represents the image features; by setting the resonance threshold, the position where the resonance intensity is greater than the resonance threshold is marked as a spatial resonance point, thereby determining the resonance area in the two-dimensional space.
2. The fusion method for automatic processing of multi-source heterogeneous data based on convolutional neural network according to claim 1 is characterized in that: The establishing of the spatial mapping relationship comprises: Obtain the physical installation location of each sensor, including geographic coordinates, relative height, and direction angle information to form three-dimensional coordinates; Collect image data and obtain pixel coordinates of image feature areas; The following formula is used for projection transformation: where u i ,v i are the horizontal and vertical pixel coordinates of the image, x s ,y s ,z s are the coordinates of the sensor in three directions of the three-dimensional space; K is the intrinsic parameter matrix of the camera, and R is the rotation matrix, which represents the rotation posture of the camera in the three-dimensional space; The corresponding pixel area of each sensor in the image is determined through projection transformation, the spatial alignment of sensor data and image feature area is achieved, and a spatial mapping relationship table is generated.
3. The fusion method for automatic processing of multi-source heterogeneous data based on convolutional neural network according to claim 1 is characterized in that Organizing the sensor data into sensor two-dimensional data according to the mapping relationship comprises: According to the spatial mapping relationship, the time series data of each sensor is mapped to the corresponding image plane coordinates; Construct a two-dimensional matrix whose size is consistent with the spatial distribution of the image feature area, and each element corresponds to the sensor data at a location; Fill the measurement data of each sensor position into the corresponding pixel position to form a two-dimensional representation; The pixel positions without fill data are filled using interpolation method.
4. The fusion method for automatic processing of multi-source heterogeneous data based on convolutional neural network according to claim 1 is characterized in that Determining the fusion feature based on the resonance region and the first fusion feature includes: The high-order features of the image data are extracted by the first two-dimensional convolutional neural network to obtain the high-order features of the image; The high-order features of the two-dimensional data of the sensor are extracted by the second two-dimensional convolutional neural network to obtain the high-order features of the sensor; Aligning the image high-order features and the sensor high-order features in the spatial dimension; splicing the image high-order features and the sensor high-order features in the channel dimension to form a new joint feature matrix splicing feature; A new 2D convolutional layer is used to convolve the concatenated features to extract the deep interaction patterns between multimodal features. After the convolution operation, a nonlinear activation function is applied to enhance the expressiveness of the features and retain the fused feature information of the forward activation; The result obtained after the convolution operation is the first fusion feature; Generate a binary mask matrix mask of the same size as the first fusion feature to mark the position of the resonance area; Multiply the mask matrix pixel by pixel with the first fusion feature to obtain a weighted feature matrix weighted fusion feature; The weighted fusion features are subjected to nonlinear activation processing, the weighted features are normalized, and finally the fusion features are output.
5. A fusion system for automatic processing of multi-source heterogeneous data based on convolutional neural networks, characterized in that: The system includes the following modules: A data acquisition module, used to acquire sensor data at different positions, and acquire image data including the sensor installation position; A mapping module, used to establish a spatial mapping relationship according to the spatial correspondence between the sensor installation position and the image feature area; A processing module, used for organizing the sensor data into sensor two-dimensional data according to the mapping relationship; A detection module, used for determining a resonance area between the two-dimensional sensor data and the image data using an LSTM model; An extraction module, configured to extract a first high-order feature of the image data using a first two-dimensional convolutional neural network, and to extract a second high-order feature of the two-dimensional sensor data using a second two-dimensional convolutional neural network; A fusion module, configured to fuse the first high-order feature and the second high-order feature using a convolution operation to obtain a first fused feature; and determine a fused feature based on the resonance region and the first fused feature; Determining the resonance region between the two-dimensional sensor data and the image data using the LSTM model includes: The resonance region is specifically a region where sensor data and image data change significantly in both time and space; The two-dimensional data of the sensor is regarded as a time series feature, where the two-dimensional matrix at each time point represents the spatial distribution of the sensor at the corresponding time; Extract image data into time series form; Build a dual-input LSTM model to process the sensor two-dimensional data and image time series features in parallel. Use LSTM to extract time-related spatial features, and use another LSTM to extract time-related image features. Temporally align sensor features and image features and conduct similarity analysis; The similarity between the sensor features and the image features is calculated in the time series. When the similarity exceeds a preset threshold, the corresponding time point is determined to be a time resonance point. For the time resonance point, the common activation area of sensor data and image features in two-dimensional space is further identified; The two-dimensional features of the sensor and the image features are spatially matched, and the resonance intensity is calculated using the Hadamard product. C(u,v,t)=F s (u,v,t)⊙F i (u,v,t) Where u represents the horizontal position of the pixel, v represents the vertical position of the pixel, t represents the time parameter, ⊙ represents the Hadamard product, C(u,v,t) represents the resonance intensity, and F s (u,v,t) represents the sensor characteristics, F i (u, v, t) represents the image features; by setting the resonance threshold, the position where the resonance intensity is greater than the resonance threshold is marked as a spatial resonance point, thereby determining the resonance area in the two-dimensional space.
6. The fusion system for automatic processing of multi-source heterogeneous data based on convolutional neural network according to claim 5 is characterized in that: The establishing of the spatial mapping relationship comprises: Obtain the physical installation location of each sensor, including geographic coordinates, relative height, and direction angle information to form three-dimensional coordinates; Collect image data and obtain pixel coordinates of image feature areas; The following formula is used for projection transformation: where u i ,v i are the horizontal and vertical pixel coordinates of the image, x s ,y s ,z s are the coordinates of the sensor in three directions of the three-dimensional space; K is the intrinsic parameter matrix of the camera, and R is the rotation matrix, which represents the rotation posture of the camera in the three-dimensional space; The corresponding pixel area of each sensor in the image is determined through projection transformation, the spatial alignment of sensor data and image feature area is achieved, and a spatial mapping relationship table is generated.
7. The fusion system for automatic processing of multi-source heterogeneous data based on convolutional neural network according to claim 5 is characterized in that Organizing the sensor data into sensor two-dimensional data according to the mapping relationship comprises: According to the spatial mapping relationship, the time series data of each sensor is mapped to the corresponding image plane coordinates; Construct a two-dimensional matrix whose size is consistent with the spatial distribution of the image feature area, and each element corresponds to the sensor data at a location; Fill the measurement data of each sensor position into the corresponding pixel position to form a two-dimensional representation; The pixel positions without fill data are filled using interpolation method.
8. The fusion system for automatic processing of multi-source heterogeneous data based on convolutional neural network according to claim 5 is characterized in that Determining the fusion feature based on the resonance region and the first fusion feature includes: The high-order features of the image data are extracted by the first two-dimensional convolutional neural network to obtain the high-order features of the image; The high-order features of the two-dimensional data of the sensor are extracted by the second two-dimensional convolutional neural network to obtain the high-order features of the sensor; Aligning the image high-order features and the sensor high-order features in the spatial dimension; splicing the image high-order features and the sensor high-order features in the channel dimension to form a new joint feature matrix splicing feature; A new 2D convolutional layer is used to convolve the concatenated features to extract the deep interaction patterns between multimodal features. After the convolution operation, a nonlinear activation function is applied to enhance the expressiveness of the features and retain the fused feature information of the forward activation; The result obtained after the convolution operation is the first fusion feature; Generate a binary mask matrix mask of the same size as the first fusion feature to mark the position of the resonance area; Multiply the mask matrix pixel by pixel with the first fusion feature to obtain a weighted feature matrix weighted fusion feature; The weighted fusion features are subjected to nonlinear activation processing, the weighted features are normalized, and finally the fusion features are output.
Citation Information
Patent Citations
Target situation fusion sensing method and system based on multiple sensors
CN110866887A
High-voltage switch cabinet situation awareness method based on multi-mode deep learning
CN113780060A