Multi-source heterogeneous data fusion method and system based on edge computing

Directly processing multi-source heterogeneous data through edge computing solves the delay problem caused by centralized processing in the cloud, achieves efficient data fusion at the edge nodes, generates multimodal fusion semantic vectors, and improves the accuracy and reliability of data analysis and decision-making.

CN120541795BActive Publication Date: 2025-09-23SHANGHAI WICRENET CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511041104.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-09-23
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

In existing technologies, cloud-based centralized processing architectures produce significant network delays when transmitting multi-source heterogeneous data, resulting in reduced data timeliness and an inability to accurately capture the localized characteristics and dynamic semantic associations of data sources. As a result, the fusion results fail to fully reflect the true state of the real-time scenario and cannot improve the accuracy of data analysis and decision-making.

Method used

A multi-source heterogeneous data fusion method based on edge computing is adopted to generate multimodal fusion semantic vectors through data type identification and feature extraction, feature dimension unification and time base adjustment, modal feature synchronization verification and weight allocation, redundant feature filtering and feature merging, and perform data processing directly at the edge node.

Benefits of technology

It reduces the network bandwidth consumption of massive raw data, meets the real-time response requirements, and dynamically allocates modal synchronization weights to make the fusion results more realistically reflect the real-time scene status. By identifying and filtering redundant feature dimensions, it improves the accuracy and reliability of data analysis and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541795B_ABST
    Figure CN120541795B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data fusion technology and discloses a multi-source heterogeneous data fusion method and system based on edge computing. The method includes obtaining heterogeneous data streams and performing data type identification and data feature extraction to obtain an original feature set; based on the original feature set, unifying feature dimensions and adjusting the time base to obtain time series vector data; based on the time series vector data, filling in the feature values ​​of missing time points to obtain multimodal features; storing the multimodal features in blocks, verifying the synchronization of adjacent modalities, and assigning modal synchronization weight coefficients to ultimately generate a fused feature matrix; based on the fused feature matrix, identifying and filtering redundant feature dimensions, establishing a feature association map, and performing feature merging to obtain a simplified feature matrix; based on the simplified feature matrix, calculating feature importance scores, arranging the features in descending order of score, and assigning ranking weight coefficients to generate a multimodal fusion semantic vector. This method improves the accuracy of data analysis and decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data fusion technology, and in particular to a multi-source heterogeneous data fusion method and system based on edge computing. Background Art

[0002] Currently, driven by the explosive growth of IoT devices and the wave of digital transformation, scenarios such as smart cities, the industrial internet, and intelligent monitoring are experiencing a surge in demand for real-time fusion of multi-source data. These scenarios require the simultaneous processing of heterogeneous data sources with diverse structures and complex formats, such as video streams, sensor time series data, and text logs, and the extraction of valuable information from these data sources to support intelligent decision-making with millisecond-level responses. While traditional architectures that rely on centralized cloud processing possess powerful computing power, they face increasing challenges in real-time transmission of massive amounts of heterogeneous data, unified format conversion, and localized semantic understanding. Data transmission is prone to bandwidth bottlenecks leading to delays, cross-modal features are difficult to deeply correlate, and centralized processing struggles to meet the low-latency processing requirements of dynamic data streams at the edge. Therefore, how to efficiently achieve deep fusion of multi-source heterogeneous data on resource-constrained edge nodes has become a key technical bottleneck in improving the performance of real-time intelligent systems.

[0003] In one existing technology, a unified data fusion server is deployed in a cloud data center. All edge devices (video cameras, sensors, and log servers) transmit the collected raw, heterogeneous data (video streams, sensor readings, and text logs) to this server via a wide area network. Upon receiving the data, the server first runs a format conversion module: extracting keyframes from the video stream and encoding them into a JPEG image sequence; packaging the sensor value sequence into a CSV format table; and performing word segmentation and frequency analysis on the text logs to generate bag-of-words vectors. The converted data is then fed into a feature fusion engine, which calculates a weighted average of the feature vectors of each data type using predefined fixed weights (for example, a weight of 0.5 for video features, 0.3 for sensor features, and 0.2 for text features) to generate a fused feature vector. Finally, this fused vector is stored in a cloud database for subsequent analysis.

[0004] Existing cloud-based centralized processing architectures generate significant network latency when transmitting heterogeneous data from multiple sources, reducing data timeliness. Their coarse-grained, generalized strategies for format conversion and fixed-weight fusion fail to accurately capture the localized characteristics and dynamic semantic associations of data sources, resulting in fusion results that fail to fully reflect the true state of real-time scenarios. Consequently, existing technologies are unable to improve the accuracy of data analysis and decision-making. Summary of the Invention

[0005] The present invention provides a multi-source heterogeneous data fusion method and system based on edge computing to improve the accuracy of data analysis and decision-making.

[0006] In a first aspect, in order to solve the above technical problems, the present invention provides a multi-source heterogeneous data fusion method based on edge computing, comprising:

[0007] Acquire heterogeneous data streams from video surveillance, sensor networks, and text logs;

[0008] According to the heterogeneous data stream, data type identification and data feature extraction are performed to obtain an original feature set;

[0009] According to the original feature set, unify the feature dimensions and adjust the time base to obtain time series vector data;

[0010] Filling in the feature values ​​of the missing time points according to the time series vector data to obtain multimodal features;

[0011] The multimodal features are stored in blocks, the synchronization of adjacent modes is verified, and modal synchronization weight coefficients are assigned to finally generate a fusion feature matrix;

[0012] According to the fusion feature matrix, redundant feature dimensions are identified and filtered, a feature association map is established, and feature merging is performed to obtain a simplified feature matrix;

[0013] According to the simplified feature matrix, feature importance scores are calculated, and the features are arranged in descending order of scores and assigned sorting weight coefficients to generate a multimodal fusion semantic vector.

[0014] In an optional embodiment, the data type identification and data feature extraction are performed according to the heterogeneous data stream to obtain an original feature set, including:

[0015] According to the heterogeneous data stream, a data type identifier is identified by a regular expression matching algorithm and classified and marked to obtain a classified data stream, wherein the classified data stream includes video type data, sensor type data and text type data;

[0016] According to the video type data, frame-by-frame analysis is performed to extract a key frame sequence to obtain a video feature set;

[0017] According to the sensor type data, time series statistics are calculated by a sliding window algorithm to generate sensor time series features;

[0018] Extract key words based on the text type data through word frequency statistics to obtain lexical semantic features;

[0019] The video feature set, the sensor temporal features and the lexical semantic features are taken as an original feature set.

[0020] In an optional implementation, unifying feature dimensions and adjusting a time reference based on the original feature set to obtain time series vector data includes:

[0021] According to the original feature set, the video feature set is subjected to dimensionality compression processing by a principal component analysis algorithm to obtain a video dimensionality reduction feature;

[0022] According to the original feature set, the sensor time series features and the lexical semantic features are subjected to zero-filling and expansion processing to obtain sensor extended features and lexical extended features;

[0023] Performing time alignment on the video feature set, the sensor time series feature, and the lexical semantic feature through a network time protocol to obtain a time series alignment feature;

[0024] According to the time series alignment feature, the time series vector data is obtained by rearranging in ascending time order using a quick sorting algorithm.

[0025] In an optional embodiment, filling in the feature values ​​of missing time points according to the time series vector data to obtain multimodal features includes:

[0026] Calculating the time intervals between adjacent data points based on the time series vector data;

[0027] According to the time interval and the preset time window constraints, the optimal time window is determined by the dynamic window algorithm;

[0028] According to the time series vector data and the optimal time window, the feature values ​​of the missing time points are processed by a linear interpolation filling algorithm to obtain a complete multimodal feature;

[0029] The multimodal features include text modal features, image modal features and audio modal features.

[0030] In an optional embodiment, the multimodal features are stored in blocks, the synchronization of adjacent modalities is verified, and modal synchronization weight coefficients are assigned to finally generate a fusion feature matrix, including:

[0031] According to the data volume of the multimodal features, the data is stored in blocks using a ring buffer algorithm to obtain cache data blocks;

[0032] Extracting the modal features and the corresponding timestamp sequence in the cached data block, calculating the time difference between adjacent modalities, and comparing it with a preset time difference threshold; when the time difference between adjacent modalities is greater than the time difference threshold, adjusting the corresponding timestamp sequence using a linear compensation algorithm;

[0033] When the time differences between all adjacent modes are less than the time difference threshold, a synchronous mode feature is obtained;

[0034] According to the synchronization modal features, the complementary coefficients between different modalities are calculated by using a cosine similarity algorithm, and the complementary coefficients are normalized to obtain a synchronization weight coefficient;

[0035] The synchronization weight coefficient and the synchronization modal feature are weightedly fused to obtain a fusion feature matrix.

[0036] In an optional embodiment, identifying and filtering redundant feature dimensions based on the fused feature matrix, establishing a feature association map, and performing feature merging to obtain a streamlined feature matrix includes:

[0037] Calculating semantic similarity using a cosine similarity algorithm based on the fusion feature matrix; and when the semantic similarity exceeds a preset similarity threshold, determining that the feature dimension pair has a redundant relationship, thereby obtaining a redundant labeling feature matrix;

[0038] According to the redundant label feature matrix, the adjacency matrix storage structure is used to record the mapping relationship between modalities, and the feature association map containing redundant feature clusters is obtained;

[0039] The redundant feature clusters and the synchronization weight coefficients in the feature association map are merged into single-dimensional features through a weighted average algorithm to obtain a simplified feature matrix.

[0040] In an optional embodiment, the step of calculating feature importance scores based on the simplified feature matrix, arranging the features in descending order of scores and assigning ranking weight coefficients to generate a multimodal fusion semantic vector includes:

[0041] According to the simplified feature matrix, the important score of each dimension feature is calculated by the information entropy method;

[0042] According to the importance score, the simplified feature matrix is ​​sorted in descending order by a quick sort algorithm,

[0043] And the linear interpolation method is used to calculate the sorting weight coefficient;

[0044] According to the sorting weight coefficient, the weight value is adjusted by the gradient descent method combined with the back propagation algorithm to obtain the optimized weight coefficient;

[0045] The optimized weight coefficient and the simplified feature matrix are weightedly fused to obtain a multimodal fusion semantic vector.

[0046] In a second aspect, the present invention provides a multi-source heterogeneous data fusion system based on edge computing, comprising:

[0047] Data acquisition module, used to obtain heterogeneous data streams from video surveillance, sensor networks, and text logs;

[0048] A feature extraction module is used to identify data types and extract data features based on the heterogeneous data stream to obtain an original feature set;

[0049] A time alignment module is used to unify feature dimensions and adjust time references based on the original feature set to obtain time series vector data;

[0050] A time filling module is used to fill the feature values ​​of missing time points according to the time series vector data to obtain multimodal features;

[0051] A data fusion module is used to store the multimodal features in blocks, verify the synchronization of adjacent modes, and assign modal synchronization weight coefficients to ultimately generate a fusion feature matrix;

[0052] A data reduction module is used to identify and filter redundant feature dimensions based on the fused feature matrix, establish a feature association map, and perform feature merging to obtain a reduced feature matrix;

[0053] The data output module is used to calculate the feature importance scores based on the simplified feature matrix, arrange them in descending order of the scores and assign sorting weight coefficients to generate a multimodal fusion semantic vector.

[0054] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements any one of the above-mentioned multi-source heterogeneous data fusion methods based on edge computing.

[0055] In a fourth aspect, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned edge computing-based multi-source heterogeneous data fusion methods.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] (1) Directly identify data types and extract features from heterogeneous data streams (video, sensor, text), reducing network bandwidth consumption and latency in transmitting massive amounts of raw data to the cloud, and meeting real-time response requirements.

[0058] (2) By verifying the synchronization of adjacent modalities (calculating time difference, threshold comparison, and timestamp compensation) and dynamically allocating modal synchronization weight coefficients (based on complementary coefficient normalization), the fixed weight fusion strategy is replaced, making the fusion result more realistically reflect the real-time scene status.

[0059] (3) By identifying and filtering redundant feature dimensions (based on semantic similarity thresholds), establishing a feature association map (adjacency matrix) and performing feature merging (weighted averaging), a streamlined feature matrix is ​​obtained; then, the feature importance scores are sorted in descending order and sorting weight coefficients are assigned to reveal the core information dimensions.

[0060] (4) The sorting weight coefficient is adjusted by combining the gradient descent method with the back propagation algorithm, and a multimodal fusion semantic vector is generated by weighted fusion with the simplified feature matrix to improve the accuracy and reliability of subsequent data analysis and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is a flow chart of a multi-source heterogeneous data fusion method based on edge computing provided by the first embodiment of the present invention;

[0062] Figure 2 This is a structural diagram of a multi-source heterogeneous data fusion system based on edge computing provided by the second embodiment of the present invention. DETAILED DESCRIPTION

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0064] Reference Figure 1 The first embodiment of the present invention provides a multi-source heterogeneous data fusion method based on edge computing, comprising the following steps:

[0065] S11, acquires heterogeneous data streams from video surveillance, sensor networks, and text logs;

[0066] S12, performing data type identification and data feature extraction based on the heterogeneous data stream to obtain an original feature set;

[0067] S13, unifying feature dimensions and adjusting time bases according to the original feature set to obtain time series vector data;

[0068] S14, filling in the feature values ​​of the missing time points according to the time series vector data to obtain multimodal features;

[0069] S15, storing the multimodal features in blocks, verifying the synchronization of adjacent modalities, and assigning modal synchronization weight coefficients, and finally generating a fusion feature matrix;

[0070] S16, identifying and filtering redundant feature dimensions according to the fused feature matrix, establishing a feature association map, and performing feature merging to obtain a streamlined feature matrix;

[0071] S17, calculating feature importance scores based on the simplified feature matrix, arranging them in descending order of scores and assigning sorting weight coefficients to generate a multimodal fusion semantic vector.

[0072] In step S11, heterogeneous data streams from video surveillance, sensor networks, and text logs are obtained.

[0073] Specifically, edge computing nodes directly receive three types of raw data streams: video surveillance data encoded in H.264 format, with a "VID_" prefix identifier in its packet header; sensor network data transmitted via the Modbus protocol, containing time-series values ​​such as temperature, humidity, or vibration frequency; and text log data recorded using the Syslog protocol, containing device status event descriptions. During data acquisition, edge nodes capture the raw byte streams received from network ports in real time, avoiding the transmission of massive amounts of raw data to the cloud. This eliminates network bandwidth bottlenecks and provides a localized data foundation for subsequent millisecond-level feature extraction.

[0074] In step S12, data type identification and data feature extraction are performed based on the heterogeneous data stream to obtain an original feature set.

[0075] In a specific embodiment, the data type identification and data feature extraction are performed according to the heterogeneous data stream to obtain the original feature set, including:

[0076] According to the heterogeneous data stream, a data type identifier is identified by a regular expression matching algorithm and classified and marked to obtain a classified data stream, wherein the classified data stream includes video type data, sensor type data and text type data;

[0077] According to the video type data, frame-by-frame analysis is performed to extract a key frame sequence to obtain a video feature set;

[0078] According to the sensor type data, time series statistics are calculated by a sliding window algorithm to generate sensor time series features;

[0079] Extract key words based on the text type data through word frequency statistics to obtain lexical semantic features;

[0080] The video feature set, the sensor temporal features and the lexical semantic features are taken as an original feature set.

[0081] Specifically, a regular expression matching algorithm is first used to scan data stream packet headers. The matching rules are derived from a preset identifier rule base. This rule base is constructed by statistically analyzing frequently occurring protocol identifier patterns in historical data streams (e.g., 95% of packets in video streams begin with "VID_," 88% of packets in sensor streams contain the protocol name "MODBUS_," and 90% of log lines in text streams begin with a timestamp plus "SYSLOG_"). The algorithm automatically extracts these high-frequency identifiers to form the rule base, without manual input. Video data is then parsed frame by frame, and the pixel change rate between adjacent frames is calculated using frame differencing. The change rate threshold (e.g., 5%) is determined based on an analysis of the dynamic nature of the video content. During the training phase, the inter-frame change rate distribution is calculated for sample videos, and the 90th percentile value is used as the threshold. Only keyframe sequences with pixel changes exceeding the threshold are extracted. The SIFT feature extraction algorithm is then used to detect scale-invariant keypoints in each keyframe and generate a 1024-dimensional scale-invariant feature vector to form the video feature set. A sliding window statistical algorithm is applied to sensor data, with a window length (e.g., 500 milliseconds) based on the statistical characteristics of the sensor data sampling frequency. The algorithm calculates the median of the time interval between adjacent data points in the historical data and multiplies it by the preset number of data points in the window (set to 10 by domain experts based on data volatility requirements). Four statistical measures—the arithmetic mean, variance, maximum, and minimum—are calculated within the window to generate a 64-dimensional sensor time series feature vector. A TF-IDF weighted statistical algorithm is applied to text data. Chinese word segmentation is performed using a word segmentation dictionary expanded from an industry terminology database. The frequency of occurrence of each word in the log file is counted and an inverse document frequency weight is calculated. The weight threshold (e.g., 0.65) is automatically determined by the algorithm. The TF-IDF distribution of all words is calculated on the historical text dataset, and the 75th percentile is used as the screening threshold. Only keywords with weights above the threshold are retained, generating a 128-dimensional text keyword vector set as lexical semantic features. The above three types of feature extraction operations are all completed locally on the edge node, and the video feature set, sensor temporal features, and lexical semantic features are finally integrated into the original feature set.

[0082] In step S13, based on the original feature set, the feature dimensions are unified and the time base is adjusted to obtain time series vector data.

[0083] In a specific embodiment, unifying feature dimensions and adjusting a time reference based on the original feature set to obtain time series vector data includes:

[0084] According to the original feature set, the video feature set is subjected to dimensionality compression processing by a principal component analysis algorithm to obtain a video dimensionality reduction feature;

[0085] According to the original feature set, the sensor time series features and the lexical semantic features are subjected to zero-filling and expansion processing to obtain sensor extended features and lexical extended features;

[0086] Performing time alignment on the video feature set, the sensor time series feature, and the lexical semantic feature through a network time protocol to obtain a time series alignment feature;

[0087] According to the time series alignment feature, the time series vector data is obtained by rearranging in ascending time order using a quick sorting algorithm.

[0088] Specifically, a principal component analysis dimensionality reduction algorithm is first performed on the video feature set in the original feature set (i.e., the 1024-dimensional scale-invariant feature vector set). This algorithm compresses the dimensions to 512 to generate video dimensionality reduction features by calculating the covariance matrix of the feature vectors and retaining the principal components whose cumulative contribution rate exceeds a preset threshold (e.g., 95%). The contribution rate threshold is automatically set by the algorithm based on the statistical law of feature importance distribution in historical data analysis. At the same time, a zero-fill expansion algorithm is performed on the sensor time series feature vector set (originally 64 dimensions) and the lexical semantic feature vector set (originally 128 dimensions): elements with a value of zero are added to the tail of the feature vector until the dimension is expanded to 512, generating sensor extended features and lexical extended features, ensuring the uniformity of the feature dimensions of all modalities. Next, the time base is calibrated using the NTP protocol. Specifically, the edge node's own clock is used as the reference server time. The average offset of the video feature timestamp set, the sensor feature timestamp set, and the text feature timestamp set relative to the reference time is calculated (for example, the average delay of a text stream is 200 milliseconds). A linear time compensation algorithm is then used to shift the timestamps of all delayed features (a 200 millisecond delay results in a 200 millisecond reduction in the timestamp value), resulting in a time-aligned feature set. Finally, a quick sort algorithm is used to rearrange all feature vectors in this set in ascending timestamp order, using the millisecond timestamp value as the sort key. This outputs a strictly chronologically ordered time series vector dataset, establishing a precise timeline foundation for subsequent feature fusion.

[0089] In step S14, the feature values ​​of the missing time points are filled in according to the time series vector data to obtain multimodal features.

[0090] In a specific embodiment, filling in the feature values ​​of missing time points according to the time series vector data to obtain multimodal features includes:

[0091] Calculating the time intervals between adjacent data points based on the time series vector data;

[0092] According to the time interval and the preset time window constraints, the optimal time window is determined by the dynamic window algorithm;

[0093] According to the time series vector data and the optimal time window, the feature values ​​of the missing time points are processed by a linear interpolation filling algorithm to obtain a complete multimodal feature;

[0094] The multimodal features include text modal features, image modal features and audio modal features.

[0095] Specifically, the time interval distribution of adjacent data points in the time series vector data is first calculated (for example, it is detected that the time interval is concentrated in the range of 100-300 milliseconds), and the dynamic window adjustment algorithm is run based on the preset time window constraints (the minimum time window of 100 milliseconds is determined by the minimum sampling period of the sensor, and the maximum time window of 500 milliseconds is determined by the maximum allowable interval of the video key frame): If it is detected that the mode of the time interval distribution of the current data stream is in the range of 100-200 milliseconds, 200 milliseconds is used as the optimal time window length. This length value is automatically determined by minimizing the product of the variance of the data points in the window and the window size. The timeline is then divided based on the optimal time window. Missing time points within the window (e.g., a sensor has no data at t+150 milliseconds) are filled using a linear interpolation algorithm. The algorithm extracts valid feature values ​​immediately preceding and following the missing point (e.g., sensor extended feature vector A at t+100 milliseconds and vector B at t+200 milliseconds). A weighted average is calculated based on the temporal distance ratio (150 milliseconds is midway between 100 milliseconds and 200 milliseconds, so the weights are 0.5 for each). This fills the gap (0.5 × A + 0.5 × B). Finally, complete, time-aligned feature values ​​are generated simultaneously for the three modalities: video dimensionality reduction features, sensor extended features, and vocabulary extended features. The resulting multimodal feature matrix comprises text modality features (vocabulary extended features), image modality features (video dimensionality reduction features), and sensor modality features (sensor extended features). This process ensures feature integrity across all modalities on the same timeline, laying the foundation for subsequent cross-modal synchronous fusion.

[0096] In step S15, the multimodal features are stored in blocks, the synchronization of adjacent modalities is verified, and modal synchronization weight coefficients are allocated to finally generate a fusion feature matrix.

[0097] In a specific embodiment, the multimodal features are stored in blocks, the synchronization of adjacent modalities is verified, and modal synchronization weight coefficients are assigned to finally generate a fusion feature matrix, including:

[0098] According to the data volume of the multimodal features, the data is stored in blocks using a ring buffer algorithm to obtain cache data blocks;

[0099] Extracting the modal features and the corresponding timestamp sequence in the cached data block, calculating the time difference between adjacent modalities, and comparing it with a preset time difference threshold; when the time difference between adjacent modalities is greater than the time difference threshold, adjusting the corresponding timestamp sequence using a linear compensation algorithm;

[0100] When the time differences between all adjacent modes are less than the time difference threshold, a synchronous mode feature is obtained;

[0101] According to the synchronization modal features, the complementary coefficients between different modalities are calculated by using a cosine similarity algorithm, and the complementary coefficients are normalized to obtain a synchronization weight coefficient;

[0102] The synchronization weight coefficient and the synchronization modal feature are weightedly fused to obtain a fusion feature matrix.

[0103] Specifically, in step S15, the specific process of storing multimodal features in blocks, verifying synchronization, and generating a fusion feature matrix is ​​as follows: First, based on the total data volume of the multimodal feature matrix, the data is divided into fixed-size data blocks for circular storage (the data block size of 64MB is preset based on the memory capacity of the edge node) through a circular buffer queue algorithm, and cache data blocks are generated to adapt to resource constraints. The timestamp sequences corresponding to the text modal features, image modal features, and sensor modal features in the cache data blocks are extracted, and the absolute value of the time difference between adjacent modal features (for example, the difference between the text modal feature timestamp and the image modal feature timestamp) is calculated and compared with a preset time difference threshold of 50 milliseconds - this threshold is automatically determined by statistically analyzing the distribution of multimodal transmission delays in historical data: the inter-modal time difference of 1,000 groups of historical samples is analyzed, and the average value plus three times the standard deviation is taken as the dynamic threshold. If the time difference between adjacent modalities is detected to exceed a threshold (e.g., the time difference between text and image modalities reaches 80 milliseconds), a linear compensation algorithm is used to perform a global shift adjustment on the delayed modal timestamp sequence (for an 80 millisecond delay, all timestamp values ​​of the text features are reduced by 80 milliseconds). The time difference detection and compensation operation is repeated until the time difference between all adjacent modalities is less than 50 milliseconds. At this point, the multimodal features are marked as synchronized modal features. The cosine similarity algorithm is then used to calculate the complementary coefficients between synchronized modal features: the directional similarity between the text modal feature vector and the image modal feature vector is measured to obtain the similarity S_text-image (typical value 0.65), the similarity between the text modality and the sensor modality S_text-sensor (typical value 0.58), and the similarity between the image modality and the sensor modality S_image-sensor (typical value 0.72). The three complementary coefficients are input into a normalized exponential function for weight assignment (specifically, exponential operations are performed on the complementary coefficients and the contribution of each exponent value to the total is calculated). This generates a synchronized weight coefficient vector (text weight W_text ≈ 0.4, image weight W_image ≈ 0.35, and sensor weight W_sensor ≈ 0.25). This weight coefficient vector is then used as a row vector and combined with the synchronized modal feature matrix (1536 dimensions × the number of time points, with 512 dimensions for each of the three modalities). This involves multiplying the three modal feature vectors at each time point along the time axis by their corresponding weight coefficients and then summing them. The output is a fused feature matrix with a dimension of 512 × the number of time points, completing the dynamic weighted fusion of cross-modal features.

[0104] In step S16, redundant feature dimensions are identified and filtered based on the fused feature matrix, a feature association map is established, and feature merging is performed to obtain a simplified feature matrix.

[0105] In a specific embodiment, identifying and filtering redundant feature dimensions based on the fused feature matrix, establishing a feature association map, and performing feature merging to obtain a simplified feature matrix includes:

[0106] Calculating semantic similarity using a cosine similarity algorithm based on the fusion feature matrix; and when the semantic similarity exceeds a preset similarity threshold, determining that the feature dimension pair has a redundant relationship, thereby obtaining a redundant labeling feature matrix;

[0107] According to the redundant label feature matrix, the adjacency matrix storage structure is used to record the mapping relationship between modalities, and the feature association map containing redundant feature clusters is obtained;

[0108] The redundant feature clusters and the synchronization weight coefficients in the feature association map are merged into single-dimensional features through a weighted average algorithm to obtain a simplified feature matrix.

[0109] Specifically, redundancy detection is first performed on the fused feature matrix (dimensions are equal to the number of time points × 512). The cosine similarity algorithm is used to calculate the semantic similarity between any two feature dimension vectors (e.g., the cosine of the angle between the i-th and j-th dimension feature vectors). A similarity threshold of 0.85 is preset, automatically determined based on feature dimension correlation analysis. During the training phase, the similarity distribution of all dimension pairs is calculated, and the 95th percentile is used as the redundancy threshold. When the similarity between a feature dimension pair exceeds the threshold (e.g., the similarity between dimension A and dimension B reaches 0.92), the pair is marked as redundant in the redundant marking feature matrix. A feature association graph is then constructed using an adjacency matrix storage structure. Using the 512 feature dimensions as graph nodes, bidirectional edges are established if the similarity between two dimensions exceeds the threshold (e.g., an edge between dimensions AB and BC forms a redundant cluster containing A / B / C). Cross-modal mapping relationships are also recorded (e.g., dimension A belongs to the text modality, dimension B to the image modality). For each redundant feature cluster in the atlas (e.g., containing five similar dimensions), a merging weight is assigned based on the synchronized weight coefficient vector generated in step S15: the original weights of the modalities to which each dimension in the redundant cluster belongs are extracted (e.g., a text modality dimension weight of 0.4, an image modality weight of 0.35, and a sensor modality weight of 0.25). The dimensions of the same modality are averaged and normalized proportionally (e.g., the sum of the text dimension weights accounts for 40% of the cluster weight). The entire redundant cluster is merged into a single dimension feature using a weighted average algorithm (merging value = Σ(dimension value × normalized weight)). Finally, all dimension columns marked as redundant are deleted, and the non-redundant dimensions are recombined with the merged new dimensions to output a 768-dimensional reduced feature matrix. The compression rate is dynamically determined by the number of redundant clusters (e.g., if 20% dimension redundancy is detected, the original dimension is compressed to 80% of the original dimension).

[0110] In step S17, based on the simplified feature matrix, feature importance scores are calculated, and the features are arranged in descending order of scores and assigned sorting weight coefficients to generate a multimodal fusion semantic vector.

[0111] In a specific embodiment, the step of calculating feature importance scores based on the simplified feature matrix, arranging the features in descending order of scores and assigning ranking weight coefficients to generate a multimodal fusion semantic vector includes:

[0112] According to the simplified feature matrix, the important score of each dimension feature is calculated by the information entropy method;

[0113] According to the importance score, the simplified feature matrix is ​​sorted in descending order by a quick sort algorithm,

[0114] And the linear interpolation method is used to calculate the sorting weight coefficient;

[0115] According to the sorting weight coefficient, the weight value is adjusted by the gradient descent method combined with the back propagation algorithm to obtain the optimized weight coefficient;

[0116] The optimized weight coefficient and the simplified feature matrix are weightedly fused to obtain a multimodal fusion semantic vector.

[0117] Specifically, in step S17, the process of generating a multimodal fusion semantic vector based on the simplified feature matrix is ​​as follows: First, feature importance is assessed on the 768-dimensional simplified feature matrix. The information entropy value (entropy ranges from 0 to 10) of the eigenvalue distribution is calculated dimension by dimension using an information entropy algorithm. The entropy calculation formula calculates the uncertainty of the eigenvalue based on its probability distribution. A high importance threshold of 7.0 is preset—this threshold is automatically determined by analyzing the entropy distribution of valid and noise features in historical samples. The 90th percentile of the entropy values ​​of all dimensions is calculated on the training set as the judgment criterion. Dimensions with an entropy value exceeding 7.0 are marked as high-importance features (e.g., dimension X has an entropy value of 7.8), and a feature importance score vector (a 768-dimensional floating-point array) is generated. The simplified feature matrix is ​​then reordered in descending order of score using a quick sort algorithm. Using the importance score as the sort key, the 768 feature dimensions are reorganized into an ordered sequence (e.g., high-entropy dimension X is ranked first, low-entropy dimension Y is ranked 768th). Initial ranking weights are calculated based on the sorting position using linear interpolation: the first feature has a weight of 1.0, the last feature has a weight of 0.1, and intermediate features have weights that decrease linearly proportional to their position (e.g., the 384th feature has a weight of 0.55). These initial weights are then fed into a gradient descent optimizer (with a fixed learning rate of 0.01) for 100 iterations. Each iteration uses backpropagation to calculate the weight adjustment, and the classification accuracy of the fusion result on the validation set is used as feedback for the loss function. The final output is an optimized weight coefficient vector (a 768-dimensional floating-point array). Finally, the reduced feature matrix and the optimized weight coefficient vector are fused using a weighted linear summation algorithm: element-wise multiplication and accumulation is performed on the feature vector at each time point (output value = Σ(eigenvalue × optimized weight)). This generates a 768-dimensional multimodal fused semantic vector whose dimensions preserve the time series length of the original reduced matrix, achieving a final semantic representation that retains the core information.

[0118] This step first uses the information entropy threshold (7.0) to automatically filter high-value feature dimensions, eliminate low-information noise interference, and focus the fusion results on key information; secondly, the back-propagation mechanism based on gradient descent (learning rate 0.01) dynamically optimizes weight distribution to overcome the static defects of traditional fixed-weight fusion and adapt to environmental changes in real time (such as automatically strengthening sensor weights under abnormal conditions); finally, dimensionality compression technology is used to convert 768-dimensional features into single-dimensional semantic vectors, reducing the computing power requirements of subsequent analysis tasks by 50%, meeting the millisecond-level real-time decision-making requirements of edge nodes while retaining core semantics, and improving the accuracy and timeliness of data analysis and decision-making.

[0119] Reference Figure 2 The second embodiment of the present invention provides a multi-source heterogeneous data fusion system based on edge computing, including:

[0120] Data acquisition module, used to obtain heterogeneous data streams from video surveillance, sensor networks, and text logs;

[0121] A feature extraction module is used to identify data types and extract data features based on the heterogeneous data stream to obtain an original feature set;

[0122] A time alignment module is used to unify feature dimensions and adjust time references based on the original feature set to obtain time series vector data;

[0123] A time filling module is used to fill the feature values ​​of missing time points according to the time series vector data to obtain multimodal features;

[0124] A data fusion module is used to store the multimodal features in blocks, verify the synchronization of adjacent modes, and assign modal synchronization weight coefficients to ultimately generate a fusion feature matrix;

[0125] A data reduction module is used to identify and filter redundant feature dimensions based on the fused feature matrix, establish a feature association map, and perform feature merging to obtain a reduced feature matrix;

[0126] The data output module is used to calculate the feature importance scores based on the simplified feature matrix, arrange them in descending order of the scores and assign sorting weight coefficients to generate a multimodal fusion semantic vector.

[0127] It should be noted that the multi-source heterogeneous data fusion device based on edge computing provided in an embodiment of the present invention is used to execute all the process steps of the multi-source heterogeneous data fusion method based on edge computing in the above embodiment. The working principles and beneficial effects of the two correspond one to one, so they will not be repeated here.

[0128] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a multi-source heterogeneous data fusion program based on edge computing. When the processor executes the computer program, the steps in each of the above-mentioned embodiments of the multi-source heterogeneous data fusion method based on edge computing are implemented, such as Figure 1 Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are realized, such as a multi-source heterogeneous data fusion module based on edge computing.

[0129] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0130] The electronic device may be a computing device such as a desktop computer, notebook, PDA, or smart tablet. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the aforementioned components are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, and the like.

[0131] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the electronic device and connects various parts of the entire electronic device using various interfaces and lines.

[0132] The memory can be used to store the computer programs and / or modules. The processor implements the various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and accessing the data stored in the memory. The memory may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0133] If the module / unit integrated into the electronic device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0134] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.

[0135] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A multi-source heterogeneous data fusion method based on edge computing, characterized in that: include: Acquire heterogeneous data streams from video surveillance, sensor networks, and text logs; According to the heterogeneous data stream, data type identification and data feature extraction are performed to obtain an original feature set; According to the original feature set, unify the feature dimensions and adjust the time base to obtain time series vector data; Filling in the feature values ​​of the missing time points according to the time series vector data to obtain multimodal features; The multimodal features are stored in blocks, the synchronization of adjacent modes is verified, and modal synchronization weight coefficients are assigned to finally generate a fusion feature matrix; According to the fusion feature matrix, redundant feature dimensions are identified and filtered, a feature association map is established, and feature merging is performed to obtain a simplified feature matrix; According to the simplified feature matrix, feature importance scores are calculated, and the features are arranged in descending order of scores and assigned sorting weight coefficients to generate a multimodal fusion semantic vector.

2. The multi-source heterogeneous data fusion method based on edge computing according to claim 1 is characterized in that: The method of performing data type identification and data feature extraction according to the heterogeneous data stream to obtain an original feature set includes: According to the heterogeneous data stream, a data type identifier is identified by a regular expression matching algorithm and classified and marked to obtain a classified data stream, wherein the classified data stream includes video type data, sensor type data and text type data; According to the video type data, frame-by-frame analysis is performed to extract a key frame sequence to obtain a video feature set; According to the sensor type data, time series statistics are calculated by a sliding window algorithm to generate sensor time series features; Extract key words based on the text type data through word frequency statistics to obtain lexical semantic features; The video feature set, the sensor temporal features and the lexical semantic features are taken as an original feature set.

3. The multi-source heterogeneous data fusion method based on edge computing according to claim 2 is characterized in that: The step of unifying feature dimensions and adjusting the time base according to the original feature set to obtain time series vector data includes: According to the original feature set, the video feature set is subjected to dimensionality compression processing by a principal component analysis algorithm to obtain video dimensionality reduction features; According to the original feature set, the sensor time series features and the lexical semantic features are subjected to zero-filling and expansion processing to obtain sensor extended features and lexical extended features; Performing time alignment on the video feature set, the sensor time series feature, and the lexical semantic feature through a network time protocol to obtain a time series alignment feature; According to the time series alignment feature, the time series vector data is obtained by rearranging in ascending time order using a quick sorting algorithm.

4. The multi-source heterogeneous data fusion method based on edge computing according to claim 1 is characterized in that: Filling the feature values ​​of missing time points according to the time series vector data to obtain multimodal features includes: Calculating the time intervals between adjacent data points based on the time series vector data; According to the time interval and the preset time window constraints, the optimal time window is determined by the dynamic window algorithm; According to the time series vector data and the optimal time window, the feature values ​​of the missing time points are processed by a linear interpolation filling algorithm to obtain a complete multimodal feature; The multimodal features include text modal features, image modal features and audio modal features.

5. The multi-source heterogeneous data fusion method based on edge computing according to claim 1 is characterized in that: The multimodal features are stored in blocks, the synchronization of adjacent modes is verified, and modal synchronization weight coefficients are allocated to finally generate a fusion feature matrix, including: According to the data volume of the multimodal features, the data is stored in blocks using a ring buffer algorithm to obtain cache data blocks; Extracting the modal features and the corresponding timestamp sequence in the cached data block, calculating the time difference between adjacent modalities, and comparing it with a preset time difference threshold; when the time difference between adjacent modalities is greater than the time difference threshold, adjusting the corresponding timestamp sequence using a linear compensation algorithm; When the time differences between all adjacent modes are less than the time difference threshold, a synchronous mode feature is obtained; According to the synchronization modal features, the complementary coefficients between different modalities are calculated by using a cosine similarity algorithm, and the complementary coefficients are normalized to obtain a synchronization weight coefficient; The synchronization weight coefficient and the synchronization modal feature are weightedly fused to obtain a fusion feature matrix.

6. The multi-source heterogeneous data fusion method based on edge computing according to claim 5 is characterized in that: The method of identifying and filtering redundant feature dimensions based on the fused feature matrix, establishing a feature association map, and performing feature merging to obtain a streamlined feature matrix includes: Calculating semantic similarity using a cosine similarity algorithm based on the fusion feature matrix; and determining that a redundant relationship exists between feature dimension pairs when the semantic similarity exceeds a preset similarity threshold, thereby obtaining a redundant labeling feature matrix. According to the redundant label feature matrix, the adjacency matrix storage structure is used to record the mapping relationship between modalities, and the feature association map containing redundant feature clusters is obtained; The redundant feature clusters and the synchronization weight coefficients in the feature association map are merged into single-dimensional features through a weighted average algorithm to obtain a simplified feature matrix.

7. The multi-source heterogeneous data fusion method based on edge computing according to claim 1 is characterized in that: The step of calculating feature importance scores based on the simplified feature matrix, arranging the features in descending order and assigning ranking weight coefficients to generate a multimodal fusion semantic vector includes: According to the simplified feature matrix, the important score of each dimension feature is calculated by the information entropy method; According to the importance score, the simplified feature matrix is ​​sorted in descending order by a quick sort algorithm, And the linear interpolation method is used to calculate the sorting weight coefficient; According to the sorting weight coefficient, the weight value is adjusted by the gradient descent method combined with the back propagation algorithm to obtain the optimized weight coefficient; The optimized weight coefficient and the simplified feature matrix are weightedly fused to obtain a multimodal fusion semantic vector.

8. A multi-source heterogeneous data fusion system based on edge computing, characterized in that: include: Data acquisition module, used to obtain heterogeneous data streams from video surveillance, sensor networks, and text logs; A feature extraction module is used to identify data types and extract data features based on the heterogeneous data stream to obtain an original feature set; A time alignment module is used to unify feature dimensions and adjust time references based on the original feature set to obtain time series vector data; A time filling module is used to fill the feature values ​​of missing time points according to the time series vector data to obtain multimodal features; A data fusion module is used to store the multimodal features in blocks, verify the synchronization of adjacent modes, and assign modal synchronization weight coefficients to ultimately generate a fusion feature matrix; A data reduction module is used to identify and filter redundant feature dimensions based on the fused feature matrix, establish a feature association map, and perform feature merging to obtain a reduced feature matrix; The data output module is used to calculate the feature importance scores based on the simplified feature matrix, arrange them in descending order of the scores and assign sorting weight coefficients to generate a multimodal fusion semantic vector.

9. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the multi-source heterogeneous data fusion method based on edge computing as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the multi-source heterogeneous data fusion method based on edge computing as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-granularity extraction and enhancement method, device and equipment for multi-source heterogeneous data

    CN119513569A

  • Multi-modal multi-source heterogeneous data fusion method

    CN120372540A