A multi-source data fusion distribution processing and analysis method and system
By constructing multi-dimensional data structures and memory state reconstruction, combined with distributed processing, the problems of missing and abnormalities of heterogeneous multi-source data are solved, and efficient and reliable data completion and fusion analysis are achieved.
Patent Information
- Application Number
- CN202510808295.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The prior art is difficult to effectively deal with the differences between heterogeneous multi-source data, insufficient completion accuracy of data loss and exceptions, and failure to achieve efficient and balanced distributed processing, resulting in insufficient accuracy and reliability of multi-source data fusion analysis.
Build a multi-dimensional data structure, and optimize the reliability and efficiency of the data completion process by identifying missing or abnormal locations, using associated neighborhoods and historical observation data.
The data completion accuracy is improved, the reliability of the completion process is quantified, and the data processing efficiency and resource utilization are significantly optimized through the distributed processing mechanism, which improves the accuracy and reliability of the fusion analysis results.
Smart Images

Figure CN120316727B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method and system for multi-source data fusion distribution processing and analysis. Background Art
[0002] With the rapid development of the Industrial Internet of Things (IIoT), smart manufacturing, intelligent transportation, and smart cities, distributed heterogeneous sensor networks have been widely used in a variety of fields, including environmental monitoring, target identification, and equipment status monitoring. Typically, practical application scenarios require the deployment of multiple sensor nodes, including visual sensors, radars, lidars, and inertial measurement units, to achieve accurate multi-dimensional and multi-modal data collection.
[0003] However, existing technologies for fusion analysis of multi-source heterogeneous sensor data often rely on single-modality independent analysis or simple data splicing, ignoring the differences and inherent correlations between different sensor modalities and failing to fully leverage the advantages of collaborative multimodal data analysis. Furthermore, due to node failures, signal interference, and data transmission issues in actual deployment environments, data loss or anomalies are common. Traditional data completion methods, such as simple interpolation or local averaging, fail to fully utilize the temporal continuity and spatial correlation of data, resulting in insufficient completion accuracy and even the introduction of additional errors. Furthermore, existing data completion methods lack reliability evaluation of the completion process, which affects the accuracy and confidence of subsequent analysis.
[0004] Furthermore, as sensor networks expand, traditional centralized data processing approaches face challenges such as excessive computational load, long response delays, and insufficient scalability. Existing distributed processing methods suffer from uneven node computational load distribution and poor cross-node data synchronization and fusion, significantly limiting the efficiency and accuracy of data fusion analysis.
[0005] In summary, existing technologies struggle to effectively address the differences between heterogeneous multi-source data, lack the accuracy to complete missing and anomaly data, and fail to implement an efficient and balanced distributed processing mechanism, resulting in significant deficiencies in the accuracy and reliability of multi-source data fusion analysis. Therefore, there is an urgent need for a technical solution that can efficiently and accurately process heterogeneous sensor data, accurately complete missing or anomaly data, quantify the reliability of the completion process, and optimize distributed processing to significantly improve the reliability and overall efficiency of data fusion analysis results. Summary of the Invention
[0006] In response to the deficiencies of the existing technology, the present application provides a multi-source data fusion distributed processing and analysis method and system.
[0007] In a first aspect, the present application provides a multi-source data fusion distribution processing and analysis method, comprising:
[0008] Acquire heterogeneous sensor data from multiple distributed collection nodes, wherein the heterogeneous sensor data includes collection time, collection channel, and source node identification information;
[0009] Constructing a multidimensional data structure for representing the heterogeneous sensor data, and identifying missing or abnormal sampling locations in the multidimensional data structure;
[0010] Based on the missing or abnormal sampling positions identified in the multidimensional data structure, determining the associated neighborhood of the complementary data, and selecting, based on a preset neighborhood threshold, neighborhood data with a correlation degree greater than or equal to the neighborhood threshold as complementary reference data, and then obtaining historical continuous observation data of the complementary reference data in the time dimension and the space dimension, and constructing a memory state sequence;
[0011] Based on the memory state sequence, perform memory state reconstruction, calculate the completion candidate data and corresponding confidence weight for each missing or abnormal position, fuse multiple completion candidate data based on the confidence weight, generate a multidimensional data structure after completion processing, and calculate indicator information used to reflect the reliability of the completion process;
[0012] Performing distributed processing on the completed multidimensional data structure, the distributed processing comprising: dividing the multidimensional data structure into different processing nodes and performing cross-node synchronization;
[0013] Based on the completed multidimensional data structure and the indicator information, a multi-source fusion analysis operation is performed to output fusion analysis result data and fusion analysis confidence indicators.
[0014] As an optional implementation manner, constructing a multidimensional data structure for representing the heterogeneous sensor data includes:
[0015] Generate a first index key value corresponding to the heterogeneous sensor data based on the acquisition time, acquisition channel, and source node identifier carried in the heterogeneous sensor data;
[0016] performing a hash mapping operation on the heterogeneous sensor data based on the first index key value to determine a storage location of the heterogeneous sensor data in a preset sparse tensor;
[0017] According to the storage locations, the heterogeneous sensor data are written one by one into corresponding locations of the preset sparse tensor to generate a multidimensional data structure representing the heterogeneous sensor data.
[0018] As an optional implementation manner, determining the associated neighborhood of the complementary data includes:
[0019] Determine a target acquisition node identifier corresponding to the missing or abnormal sampling position according to a first index key value corresponding to the missing or abnormal sampling position in the multidimensional data structure;
[0020] Based on the target acquisition node identifier and a preset physical distance threshold or signal correlation threshold, selecting other acquisition node identifiers whose physical distance or signal correlation degree meets the preset threshold from the multidimensional data structure to generate a neighborhood candidate node set;
[0021] Based on the consistency of the sampling channel types of the neighborhood candidate node set and the sampling channel types corresponding to the missing or abnormal sampling position, candidate nodes with the same channel type as the missing or abnormal sampling position are screened from the neighborhood candidate node set as associated neighbors of the supplementary data, and the associated neighbors are output for selection and use by the supplementary reference data.
[0022] As an optional implementation, the step of constructing a memory state sequence includes:
[0023] Determining, based on the first index key value corresponding to the supplemented reference data, a collection node identifier corresponding to the supplemented reference data;
[0024] Retrieving, from the multidimensional data structure, historical data that appears continuously with the supplementary reference data within a preset time range according to the acquisition node identifier and the first index key value, to generate a first historical data sequence;
[0025] Retrieving, from the multidimensional data structure, the historical data of nodes adjacent to the complementary reference data in a spatial or channel dimension according to the physical location or channel characteristics of the acquisition node corresponding to the complementary reference data, to generate a second historical data sequence;
[0026] The first historical data sequence and the second historical data sequence are combined into a memory state sequence corresponding to the supplementary reference data, and the memory state sequence is output.
[0027] As an optional implementation, performing memory state reconstruction and calculating the candidate data for completion of each missing or abnormal position and the corresponding confidence weight includes:
[0028] Determining, based on the first historical data sequence and the second historical data sequence in the memory state sequence, scale features corresponding to different sensor modalities in the historical data sequence, and calculating a corresponding confidence level of each modality historical data sequence at a current moment;
[0029] Calculating the historical window importance weight of the historical data sequence relative to the current missing or abnormal position based on the time difference between the acquisition time of the historical data sequence and the current missing or abnormal position;
[0030] Calculate the memory weight for each historical data sequence according to the scale feature, the corresponding confidence level, and the historical window importance weight;
[0031] Performing memory state reconstruction on the historical data sequence based on the memory weights to generate completion candidate data of different modalities, and calculating a fusion confidence weight corresponding to each of the completion candidate data based on a combination of the memory weights;
[0032] Output the completion candidate data of the different modalities and the fusion confidence weight.
[0033] As an optional implementation, generating the multidimensional data structure after completion processing includes:
[0034] Dividing the multidimensional data structure into a plurality of preset local spatial regions according to the fusion confidence weights, and respectively calculating the variance value of the fusion confidence weights in each of the local spatial regions;
[0035] Comparing the variance value with a preset island threshold, and determining a first local area exceeding the island threshold from the local spatial area;
[0036] For the first local area, according to a preset extended neighborhood rule, determining a second local area adjacent to the first local area in the multidimensional data structure, and obtaining candidate modality data and a fusion confidence weight of the second local area;
[0037] Merging the candidate modality data and fusion confidence weight of the second local area with the candidate modality data and fusion confidence weight of the first local area to generate an extended fusion data set, and recalculating fusion data of a corresponding position of the first local area based on the extended fusion data set;
[0038] The recalculated fusion data are written one by one into the positions corresponding to the first local areas in the multidimensional data structure to generate a multidimensional data structure after extended neighborhood fusion processing.
[0039] As an optional implementation, the distributed processing of the completed multidimensional data structure includes:
[0040] Calculating node allocation parameters for each spatial region based on the fusion confidence weight corresponding to each spatial region in the completed multidimensional data structure;
[0041] Performing spatial region division on the completed multidimensional data structure according to the node allocation parameters to generate local data subsets corresponding to a plurality of processing nodes;
[0042] Sending the local data subsets to corresponding processing nodes respectively, and having the processing nodes perform local calculations on the local data subsets respectively to generate local calculation result data;
[0043] Based on the preset cross-node synchronization rules, the local calculation result data is transmitted to the preset central coordination node, and the central coordination node performs the synchronous fusion calculation of the local calculation result data to generate a multi-dimensional data structure after cross-node synchronous fusion.
[0044] As an optional implementation, performing a multi-source fusion analysis operation based on the completed multidimensional data structure and the indicator information, and outputting fusion analysis result data and fusion analysis confidence indicators includes:
[0045] Calculating feature vectors of multiple heterogeneous sensor modalities based on the completed multidimensional data structure, and appending the indicator information to the feature vectors respectively to generate a multimodal feature set with confidence indicators;
[0046] Constructing a preset multimodal analysis model for fusion analysis based on the multimodal feature set, and determining input weight parameters of the multimodal analysis model according to the indicator information;
[0047] Inputting the multimodal feature set into the multimodal analysis model, performing a multimodal joint analysis operation, and obtaining fusion analysis result data;
[0048] Calculating a fusion analysis confidence index of the fusion analysis result data according to the model output parameters corresponding to the fusion analysis result data;
[0049] The fusion analysis result data and the fusion analysis confidence index are output as the final output of the multi-source data fusion distribution processing and analysis method.
[0050] In a second aspect, the present application provides a multi-source data fusion distributed processing and analysis system, comprising:
[0051] An acquisition module is used to acquire heterogeneous sensor data from multiple distributed acquisition nodes, wherein the heterogeneous sensor data includes acquisition time, acquisition channel, and source node identification information;
[0052] A construction module, configured to construct a multidimensional data structure for representing the heterogeneous sensor data and identify missing or abnormal sampling locations in the multidimensional data structure;
[0053] a first completion module, configured to determine, based on the missing or abnormal sampling locations identified in the multidimensional data structure, an associated neighborhood of the completion data, and select, based on a preset neighborhood threshold, neighborhood data with a degree of association greater than or equal to the neighborhood threshold as completion reference data, thereby obtaining historical continuous observation data of the completion reference data in time and space dimensions, and constructing a memory state sequence;
[0054] a second completion module, configured to perform memory-based state reconstruction based on the memory state sequence, calculate completion candidate data and corresponding confidence weights for each missing or abnormal position, fuse multiple completion candidate data based on the confidence weights, generate a completed multidimensional data structure, and calculate indicator information reflecting the reliability of the completion process;
[0055] a processing module, configured to perform distributed processing on the multidimensional data structure that has undergone the completion processing, wherein the distributed processing includes: dividing the multidimensional data structure into different processing nodes and performing cross-node synchronization;
[0056] The output module is used to perform multi-source fusion analysis operations based on the completed multidimensional data structure and the indicator information, and output fusion analysis result data and fusion analysis confidence indicators.
[0057] Compared with the existing technology, the present application constructs a multidimensional data structure for representing heterogeneous sensor data and accurately identifies missing or abnormal locations, and uses associated neighborhoods and historical observation data to construct a memory state sequence to achieve high-quality data completion. This method effectively improves the accuracy of data completion and quantifies the reliability of the completion process through memory state reconstruction and confidence weight fusion mechanism. In addition, the present application also proposes a distributed processing mechanism for the fused multidimensional data structure, which significantly optimizes data processing efficiency and resource utilization through precise node load distribution and cross-node synchronous fusion strategy. Finally, through multimodal joint analysis operations and output of fusion analysis results and confidence indicators, the accuracy and reliability of the fusion analysis results are further improved. The present application effectively solves the problems of insufficient data completion accuracy, lack of reliability evaluation of the completion process, and unbalanced distributed data processing in the existing technology, and realizes efficient fusion processing and analysis of multi-source heterogeneous data. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is a flowchart of a multi-source data fusion distribution processing and analysis method provided in this application;
[0059] Figure 2 A flowchart of a method for constructing a multidimensional data structure for representing heterogeneous sensor data provided by the present application;
[0060] Figure 3A flowchart of a method for determining associated neighborhoods for complementary data provided in this application;
[0061] Figure 4 This is a schematic diagram of a multi-source data fusion distributed processing and analysis system provided in this application. DETAILED DESCRIPTION
[0062] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.
[0063] See also Figure 1 FIG. 1 is a flowchart of a multi-source data fusion distribution processing and analysis method provided in an embodiment of the present application. The method includes steps S101 to S106, wherein:
[0064] S101: Acquire heterogeneous sensor data from multiple distributed collection nodes, where the heterogeneous sensor data includes collection time, collection channel, and source node identification information;
[0065] S102: Constructing a multidimensional data structure for representing the heterogeneous sensor data, and identifying missing or abnormal sampling locations in the multidimensional data structure;
[0066] S103: Based on the missing or abnormal sampling positions identified in the multidimensional data structure, determining the associated neighborhood of the supplementary data, and selecting, based on a preset neighborhood threshold, neighborhood data with a correlation degree greater than or equal to the neighborhood threshold as the supplementary reference data, and then obtaining historical continuous observation data of the supplementary reference data in the time dimension and the space dimension, and constructing a memory state sequence;
[0067] S104: Based on the memory state sequence, perform memory state reconstruction, calculate the completion candidate data and the corresponding confidence weight for each missing or abnormal position, fuse multiple completion candidate data based on the confidence weight, generate a multidimensional data structure after completion processing, and calculate indicator information used to reflect the reliability of the completion process;
[0068] S105: performing distributed processing on the multidimensional data structure that has been completed, wherein the distributed processing includes: dividing the multidimensional data structure into different processing nodes and performing cross-node synchronization;
[0069] S106: Based on the completed multidimensional data structure and the indicator information, perform a multi-source fusion analysis operation, and output fusion analysis result data and fusion analysis confidence indicators.
[0070] Regarding S101 above:
[0071] In a specific implementation, heterogeneous sensor data is acquired through multiple distributed collection nodes deployed at different spatial locations. These distributed collection nodes include, but are not limited to, vision sensor nodes, radar sensor nodes, lidar sensor nodes, inertial measurement unit sensor nodes, and temperature and humidity sensor nodes. Each of these collection nodes collects corresponding sensor data at a preset sampling frequency and adds a collection timestamp to the data using a predetermined, unified time base to facilitate subsequent data synchronization.
[0072] Specifically, the sensor data collected by each distributed collection node is generated in the form of data packets. Each data packet clearly contains the following metadata:
[0073] Collection timestamp: used to accurately identify the time when the data was collected, which can be an absolute time point under Coordinated Universal Time (UTC) or other unified time reference;
[0074] Acquisition channel number: used to identify specific data acquisition channel information, such as the video channel corresponding to a visual sensor, the RF channel corresponding to a radar sensor, or the specific channel number of other types of sensors;
[0075] Source node identification information: used to uniquely identify the specific collection node from which the data packet comes. The node identification information can be a node number, a node hardware device address, or a preset unique node code.
[0076] In the above-mentioned specific embodiments, the data acquired by the distributed collection nodes can be transmitted to a predetermined data aggregation node or data center via wired or wireless communication. For example, each collection node can transmit the sensor data carrying the aforementioned metadata to the data aggregation node in real time or near real time via a network interface such as an industrial Ethernet interface, a CAN bus, a wireless local area network (Wi-Fi), or a low-power wide area network (LoRaWAN), thereby completing the data acquisition operation.
[0077] After receiving the data packets sent by each distributed collection node, the data aggregation node automatically performs preliminary data alignment and sorting based on the collection timestamp, collection channel number and source node identification information carried in the data packet, generates a heterogeneous sensor data set, and provides basic data input for the subsequent construction of a multidimensional data structure.
[0078] The above-mentioned data acquisition process can effectively ensure the data consistency and synchronization in the subsequent data analysis stage, significantly reduce the spatiotemporal dislocation and data deviation caused by distributed node collection, and provide high-quality data input for the subsequent data fusion and processing links of the technical solution of this application.
[0079] Regarding S102 above:
[0080] In a specific implementation, after completing the preliminary alignment and organization of the heterogeneous sensor data set, the data aggregation node further constructs a multidimensional data structure representing the heterogeneous sensor data based on the acquisition timestamp, acquisition channel number and source node identification information carried in the heterogeneous sensor data set to support subsequent data completion and fusion analysis.
[0081] Specifically, the data aggregation node determines a unified timeline index based on the acquisition timestamps carried in each data packet; determines a node dimension index based on the source node identification information carried in the data packet; and determines a channel dimension index based on the acquisition channel number carried in the data packet. Subsequently, based on these timeline indexes, node dimension indexes, and channel dimension indexes, the data aggregation node establishes a sparse multidimensional data structure for storing heterogeneous sensor data. Specifically, this structure can be implemented as a pre-defined sparse tensor to efficiently support heterogeneous data storage and retrieval.
[0082] When constructing the sparse multidimensional data structure, the data aggregation node first generates a one-to-one corresponding index key value based on each data packet's acquisition timestamp, acquisition channel number, and source node identification information. Using a preset hash mapping algorithm, the index key value is mapped to the corresponding storage location in the sparse tensor to achieve precise data location. The data aggregation node then writes the sensor data in the data packet to the corresponding location in the sparse tensor based on the index key value, completing the construction of the multidimensional data structure for the heterogeneous sensor data.
[0083] After completing the construction of the multidimensional data structure, the data aggregation node further identifies the missing or abnormal sampling positions in the multidimensional data structure to support subsequent data completion processing. Specifically, the data aggregation node first performs a continuity check on the time dimension of the multidimensional data structure. When the interval between two adjacent sampling timestamps exceeds a preset sampling time interval threshold, the corresponding position is marked as a missing sampling position. In addition, the data aggregation node further performs a local change rate analysis on each data point in the multidimensional data structure, by calculating the change amplitude with adjacent moments or adjacent spatial positions. When the calculated change amplitude exceeds a preset anomaly detection threshold, the corresponding position is marked as an abnormal sampling position.
[0084] This marking process generates a mask matrix with the same dimensions as the multidimensional data structure, where each element indicates whether the data at the corresponding location is missing or abnormal. The data aggregation node outputs the multidimensional data structure and the corresponding mask matrix as data input for subsequent steps to support further completion of missing or abnormal locations.
[0085] Through the above-mentioned multi-dimensional data structure construction and missing and abnormal position identification process, high-quality and clear data input can be provided for subsequent data completion and fusion processing, effectively ensuring the accuracy and stability of subsequent data fusion analysis.
[0086] Regarding S103 above:
[0087] In the specific implementation, the data aggregation node further determines the associated neighborhood of the completion data corresponding to the position to be completed based on the missing or abnormal sampling position identified in the multidimensional data structure, so as to obtain highly correlated completion reference data and construct the memory state sequence required for subsequent data completion processing.
[0088] Specifically, the data aggregation node first extracts the node identifier and channel identifier corresponding to each missing or abnormal position, and retrieves the pre-stored node layout information and channel topology information based on the node identifier and channel identifier to determine the spatial or signal correlation neighborhood range.
[0089] In actual engineering implementation, the above-mentioned associated neighborhoods can be determined by using spatial or topological similarity analysis tools, such as spatial neighborhood search tools implemented in the Python environment (such as the NearestNeighbors module of Scikit-learn), or customized topological graph search algorithms, to quickly calculate the spatial or topological distance between missing or abnormal locations and surrounding nodes, and generate a preliminary set of neighborhood candidates.
[0090] Subsequently, the data aggregation node screens the neighborhood candidate set based on a preset neighborhood threshold and selects neighborhood data with a correlation degree greater than or equal to the neighborhood threshold as the complementary reference data. Specifically, the data aggregation node pre-sets the neighborhood threshold, for example, the physical distance threshold is set to 10 meters, and the signal correlation threshold is set to a predefined inter-channel correlation coefficient greater than or equal to 0.7, and uses numerical computing libraries such as NumPy and SciPy to perform the correlation threshold calculation. Node data in the neighborhood candidate nodes that meets the above neighborhood threshold conditions is determined as complementary reference data for subsequent memory state construction.
[0091] After determining the complementary reference data, the data aggregation node further performs extended-view processing to obtain historical continuous observation data in the time and space dimensions of the complementary reference data and construct a memory state sequence. Specifically, the data aggregation node uses a sliding time window mechanism based on the historical data stored in the multidimensional data structure, and searches back through the historical data one by one with a set time window length, for example, a historical observation window length of 10 sampling periods.
[0092] In terms of implementation tools, data aggregation nodes can use timeline-based data retrieval tools (such as InfluxDB or TimescaleDB) to quickly perform back-dated queries on historical data. Furthermore, by searching for node locations in the spatial dimension, the efficiency of historical data retrieval can be improved through the GeoHash algorithm or spatial tree indexes (such as KD-tree and R-tree).
[0093] After acquiring historical data, the data aggregation node further constructs a memory state sequence. For example, the data aggregation node first sorts the historical data in chronological order, then adds spatial distance and time interval information to the historical data of each node, forming a memory state data packet sequence for each complementary reference data. Each data packet clearly contains the observation value, timestamp, spatial location, and corresponding node identification information.
[0094] The construction of the above memory state sequence can be implemented through the DataFrame structure of the Pandas library or the NumPy array in the Python environment, which facilitates subsequent rapid processing and retrieval.
[0095] Through the above-mentioned determination of associated neighborhoods, neighborhood screening, extended field of view processing and construction of memory state sequences, high-quality data support can be provided for subsequent memory state reconstruction and data fusion, further ensuring the stability and accuracy of the data completion processing process.
[0096] Regarding S104 above:
[0097] In the specific implementation, after completing the construction of the memory state sequence, the data aggregation node further performs memory state reconstruction to complete the sampling data at missing or abnormal positions in the multidimensional data structure, and calculates the corresponding completion candidate data and the corresponding confidence weights, and then performs fusion processing to generate a multidimensional data structure after completion processing, and calculates the corresponding completion process reliability index information.
[0098] During implementation, the data aggregation node first performs state reconstruction calculations based on the constructed memory state sequence. Based on each missing or abnormal location, the data aggregation node uses a pre-built memory state reconstruction model, such as a prediction model based on a Transformer network or recurrent neural network (RNN), to perform temporal or spatial sequence modeling on the memory state sequence to obtain candidate data for the missing or abnormal location.
[0099] For example, data aggregation nodes can implement Transformer models or long short-term memory (LSTM) neural networks based on deep learning frameworks such as PyTorch or TensorFlow for efficient modeling and prediction. The preprocessed memory state sequence data is used as input, and the model outputs predictions based on the historical state data, resulting in multiple candidate completion data for the corresponding position.
[0100] The data aggregation node then further calculates the confidence weights corresponding to the candidate data. Specifically, this node can compare the model's prediction results with the statistical characteristics of the historical data in the memory state sequence, using numerical analysis tools (such as SciPy or NumPy libraries) to calculate the prediction error or variance for each candidate data, and then calculate the corresponding confidence weight based on this error or variance.
[0101] For example, the confidence weight can be determined by using the inverse error mapping method. The smaller the error, the higher the confidence weight. The specific calculation method can be selected according to the project implementation requirements.
[0102] After calculating the confidence weights, the data aggregation node further performs a fusion operation on the multiple candidate completion data. Specifically, the data aggregation node fuses the candidate completion data based on a preset fusion algorithm, such as a weighted average fusion algorithm or a weighted Kalman filter fusion algorithm. Taking the weighted average fusion algorithm as an example, the data aggregation node uses the confidence weights corresponding to the candidate completion data as fusion weights, performs weight normalization, and generates fused completion data values. The fused completion data values are then written one by one into the corresponding missing or abnormal locations in the original multidimensional data structure.
[0103] In addition, the data aggregation node further calculates indicators reflecting the reliability of the completion process. Specifically, the mean, variance, or fusion residual of the confidence weights corresponding to each position in the fusion process can be calculated as a quantitative indicator of the reliability of the completion process. For example, the confidence weight variance can be used to represent the stability of the completion process, while the fusion residual can be used to represent the accuracy of the completed data. These indicator calculations can be quickly implemented using statistical functions in the SciPy or Pandas libraries and output to subsequent analysis steps in a standardized data structure, such as an array or data frame.
[0104] After the above-mentioned memory state reconstruction, confidence weight calculation and fusion processing, the data aggregation node generates a complete multi-dimensional data structure after completion processing, and comes with clear reliability indicator information of the completion process, which serves as the data basis input for subsequent distributed processing and fusion analysis steps.
[0105] Regarding S106 above:
[0106] In practice, after generating a complete, fully-filled multidimensional data structure, the data aggregation node further performs distributed processing to improve the efficiency and scalability of the overall analysis process. Specifically, the data aggregation node first spatially partitions the completed multidimensional data structure and distributes the partitioned data subsets to multiple processing nodes for subsequent cross-node synchronous processing.
[0107] Specifically, the data aggregation node calculates node allocation parameters based on the fusion confidence weights corresponding to each spatial region in the completed multidimensional data structure. These node allocation parameters can be calculated based on metrics such as the amount of data in the spatial region and the mean or variance of the fusion confidence weights. For example, the data aggregation node can use the NumPy or Pandas libraries in Python to implement these parameter calculations. By performing statistical analysis on the fusion confidence weights, the data processing load characteristics of different spatial regions can be determined, and a node allocation strategy can be generated.
[0108] Based on the node allocation parameters, the data aggregation node performs spatial region partitioning. Specifically, the data aggregation node can partition a multidimensional data structure into multiple data subsets using a spatial indexing algorithm, such as a space-filling curve (Hilbert curve or Z-order curve) or a KD-tree spatial partitioning algorithm. Taking the Hilbert curve as an example, the data aggregation node uses a Python spatial partitioning tool library (such as hilbertcurve or a custom Hilbert curve implementation module) to map the multidimensional data structure into a one-dimensional space and then perform data partitioning to generate multiple data subsets.
[0109] The data aggregation node then sends the aforementioned data subsets to the corresponding multiple processing nodes via pre-established distributed data communication links. During implementation, each processing node can be a computing node in a distributed computing framework, such as a computing node implemented based on Apache Spark, Dask, or other distributed computing engines.
[0110] After receiving the assigned data subset, each processing node performs local data analysis and computation. Specifically, each processing node can perform local feature extraction, local statistics calculation, or other pre-defined data analysis tasks during the local computation process. In terms of implementation, each processing node can use the Pandas or NumPy libraries in a Python environment to complete local data processing tasks, or perform computations using a pre-trained local data analysis model (such as a local inference model implemented in PyTorch or TensorFlow) to output local computation result data.
[0111] After completing local computations, each processing node sends the local computation results to a pre-set central coordination node according to pre-set cross-node synchronization rules, achieving cross-node data synchronization and converged computation. In practice, cross-node synchronization can be achieved through distributed message queue systems (such as Apache Kafka and RabbitMQ) or remote procedure call mechanisms (such as gRPC and REST APIs) to achieve efficient data exchange and synchronization.
[0112] After receiving the local computation results from each processing node, the central coordinating node further performs synchronous fusion calculations to generate a unified multidimensional data structure after cross-node synchronous fusion. In specific implementations, the central coordinating node can use a fusion calculation algorithm, such as weighted average fusion, Kalman filter fusion, or a confidence-based adaptive fusion algorithm, to complete the fusion operation using the received local computation results and output the final unified data structure.
[0113] After the above-mentioned distributed processing and cross-node synchronization process, the data aggregation node generates a unified and complete multi-dimensional data structure, and uses this structure for the fusion analysis operations in subsequent steps to support the efficient execution of the final data analysis.
[0114] Regarding S106 above:
[0115] In the specific implementation, after completing the unified multi-dimensional data structure after cross-node synchronous fusion and the calculation of the reliability index information of the completion process, the data aggregation node further performs multi-source fusion analysis operations to achieve comprehensive analysis of multimodal heterogeneous data, and finally outputs the fusion analysis result data and the corresponding fusion analysis confidence index.
[0116] Specifically, the data aggregation node first calculates feature vectors corresponding to multiple heterogeneous sensor modalities based on the completed multidimensional data structure. Taking visual sensors, radar sensors, and inertial measurement units as examples, the data aggregation node can extract video frame features, for example, using convolutional neural networks to extract image depth feature vectors and radar point cloud features. For example, clustering algorithms and distance distributions are used to calculate radar feature vectors, and motion state features of inertial measurement unit data are used, such as through motion trajectory filtering and motion modeling to estimate feature parameters.
[0117] In actual engineering implementation, the above feature vector extraction process can be completed based on deep learning models built on the TensorFlow or PyTorch framework, such as ResNet, YOLO, or PointNet, or based on the preset feature extraction tools in the Scikit-learn, NumPy, or SciPy libraries to ensure that the feature extraction process is efficient and accurate.
[0118] The data aggregation node then applies the aforementioned indicator information, i.e., the completion process reliability indicator calculated in the previous step, to each modal feature vector, generating a multimodal feature set with explicit confidence indicators. Specifically, the data aggregation node can use Pandas DataFrames or NumPy arrays in a Python environment to organize the feature vectors and their confidence indicators, facilitating fast and efficient subsequent fusion analysis.
[0119] After generating a multimodal feature set, the data aggregation node further constructs a pre-defined multimodal analysis model to perform multimodal joint analysis. This implementation can employ models based on deep learning or statistical learning, such as the multimodal fusion Transformer model, a multi-input LSTM network, a Kalman filter-based multimodal state fuser, or a graph neural network (GNN) with a graph attention mechanism. The construction and deployment of these analysis models can be efficiently accomplished using machine learning frameworks such as TensorFlow, PyTorch, or Scikit-learn.
[0120] During the specific fusion analysis operation, the data aggregation node first determines the input weight parameters of the multimodal analysis model based on the reliability index information of the completion process. For example, data with higher reliability indexes are assigned higher input weights in the fusion operation, while data with lower reliability indexes are assigned lower input weights, thus achieving the model's adaptive weighted fusion of input features.
[0121] The data aggregation node then inputs the multimodal feature set into the multimodal analysis model, performs fusion analysis, and outputs fusion analysis results. This fusion analysis results may include comprehensive analysis results such as fused scene recognition results, target detection or tracking results, and environmental state estimation results. The specific analysis results are determined based on the actual engineering application scenario requirements.
[0122] After obtaining the fusion analysis result data, the data aggregation node further calculates the fusion analysis confidence index corresponding to the fusion analysis result data to characterize the reliability of the fusion analysis result. In the specific implementation process, the fusion analysis confidence index can be calculated based on the posterior probability distribution of the analysis model output result, the residual statistics, or the model prediction uncertainty estimate. For example, the data aggregation node can use the Softmax probability, output variance, or prediction residual of the model output as the basis for calculating the fusion analysis confidence index, and quickly implement the confidence index calculation through the SciPy or NumPy library in the Python environment.
[0123] Finally, the data aggregation node outputs the fusion analysis results and the corresponding fusion analysis confidence index, forming the final data analysis output. The fusion analysis results and confidence index data can be further stored in a database (such as MySQL, MongoDB, or a time series database such as InfluxDB) or file system for subsequent application access and analysis display.
[0124] As an optional implementation, see Figure 2 , is a flow chart of a method for constructing a multidimensional data structure for representing heterogeneous sensor data provided by the present application, including S201 to S203, wherein:
[0125] S201: Generate a first index key value corresponding to the heterogeneous sensor data based on the acquisition time, acquisition channel, and source node identifier carried in the heterogeneous sensor data;
[0126] S202: Performing a hash mapping operation on the heterogeneous sensor data based on the first index key value to determine a storage location of the heterogeneous sensor data in a preset sparse tensor;
[0127] S203: Writing the heterogeneous sensor data into corresponding positions of the preset sparse tensor one by one according to the storage positions to generate a multi-dimensional data structure representing the heterogeneous sensor data.
[0128] In some implementations, when constructing the multidimensional data structure in step S102, due to the heterogeneous and large-scale data sources, directly using a conventional dense data structure may result in low storage space utilization, insufficient indexing efficiency, and difficulty locating data. Therefore, this application proposes an alternative implementation that effectively improves the indexing and storage efficiency of heterogeneous sensor data, enhances data location accuracy, and reduces data storage redundancy by constructing a sparse tensor structure based on a hash map.
[0129] Specifically, in this embodiment, the data aggregation node first constructs a corresponding unique index key based on the acquisition timestamp, acquisition channel number, and source node identification information carried in each heterogeneous sensor data packet. In practice, the data aggregation node can combine the acquisition timestamp, acquisition channel number, and source node identification information using string concatenation or binary encoding. For example, a unique key value such as "node ID_channel number_timestamp" can be generated using string concatenation to ensure that the index key value for each data packet is unique.
[0130] The data aggregation node then performs a preset hash mapping operation on the index key value to determine the specific storage location of the data in the sparse tensor. In a specific implementation, the hash mapping operation can be implemented using a hash function commonly used in engineering, such as using a hash algorithm such as MD5, SHA-1, or SHA-256 to map the index key value to a fixed-length hash value, which is then further mapped to a pre-defined storage address in the sparse tensor space.
[0131] For example, data aggregation nodes can use the hashlib hash library in Python to perform hash mapping operations, converting index key strings into hash values, and then using modulo operations to map the hash values to the dimensional coordinates of the sparse tensor. This gives each data packet a clear and unique multi-dimensional coordinate location, avoiding the occurrence of duplicate or redundant data in the storage structure.
[0132] The data aggregation node further writes the heterogeneous sensor data one by one to the corresponding locations of the sparse tensor based on the storage locations determined above, thereby generating a multidimensional data structure representing the heterogeneous sensor data. In specific implementations, the data aggregation node can be based on a sparse data storage engine, such as the sparse matrices provided by the SciPy library in the Python environment (such as the CSR or COO format of scipy.sparse) or the SparseTensor data structure of frameworks such as TensorFlow and PyTorch, effectively supporting efficient storage and fast access of heterogeneous sensor data.
[0133] Through the above-mentioned hash mapping operation and sparse tensor construction method, this implementation effectively solves the indexing and storage efficiency problems of heterogeneous sensor data in multidimensional data structures, improves the performance of large-scale heterogeneous data processing, and thus provides a more efficient and accurate data structure foundation for subsequent data completion and fusion analysis.
[0134] For example, a data aggregation node acquires data collected by a visual sensor node, a radar sensor node, and an inertial measurement unit (IMU) node at the timestamp "20240610103015." For the visual sensor, for example, its source node identifier is "cam_03," its acquisition channel number is "channel_01," and its acquisition timestamp is "20240610103015." The data aggregation node first generates a unique index key based on these three pieces of information.
[0135] Subsequently, the data aggregation node performs a hash mapping operation to determine the storage location of the visual sensor data in a preset sparse tensor. In a specific implementation, the hash operation can be performed using the Python hash library hashlib.
[0136] The data aggregation node then writes the raw data collected by the vision sensor into a sparse tensor structure based on the determined location. For example, the data can be stored using the COO sparse matrix method in Python's SciPy library.
[0137] Similarly, for the data of the radar sensor node (source node identifier "radar_02", channel number "channel_03", also collected at timestamp "20240610103015") and the inertial measurement unit node (source node identifier "imu_01", channel number "channel_02", collected at timestamp "20240610103015"), the data aggregation node performs the same index key value generation, hash mapping, and data storage operations:
[0138] Radar data example index key value: "radar_02_channel_03_20240610103015";
[0139] Inertial measurement unit data example index key: "imu_01_channel_02_20240610103015".
[0140] As an optional implementation, see Figure 3 , is a flow chart of a method for determining an associated neighborhood of complementary data provided by the present application, including S301 to S303, wherein:
[0141] S301: Determine a target acquisition node identifier corresponding to the missing or abnormal sampling position according to a first index key value corresponding to the missing or abnormal sampling position in the multidimensional data structure;
[0142] S302: Based on the target collection node identifier and a preset physical distance threshold or signal correlation threshold, select other collection node identifiers whose physical distance or signal correlation degree meets the preset threshold from the multidimensional data structure to generate a neighborhood candidate node set;
[0143] S303: Based on the consistency of the sampling channel types of the neighborhood candidate node set and the sampling channel types corresponding to the missing or abnormal sampling position, candidate nodes with the same channel type as the missing or abnormal sampling position are screened from the neighborhood candidate node set as associated neighborhoods of the supplementary data, and the associated neighborhoods are output for selection of the supplementary reference data.
[0144] In some implementations, due to the high density and diverse nature of acquisition nodes, determining the associated nodes for supplementary data based solely on general topological relationships or simple spatial distances can result in insufficient association between the selected supplementary nodes and the actual data at the missing or abnormal location, thereby reducing data supplementation accuracy. Therefore, this application provides a further optimized method for determining the associated neighborhood of supplementary data to improve the accuracy of the association between the selected supplementary reference data and the target data.
[0145] During the specific implementation, the data aggregation node first accurately determines the corresponding target collection node identifier based on the first index key value of the missing or abnormal sampling location in the multidimensional data structure. In the specific implementation, the data aggregation node parses the source node identifier field in the index key value to determine the specific collection node corresponding to the missing or abnormal location.
[0146] The data aggregation node then determines the physical or signal relationships between nodes based on a pre-stored database of node locations or topology maps, combined with the aforementioned target node identifiers. This can be accomplished using the GeoPandas library in Python or a graph-based networking library (such as NetworkX) to analyze physical locations or signal topology relationships, allowing for rapid calculation of spatial or signal relationships between the target node and other nodes.
[0147] In a specific project implementation, the data aggregation node predefines a physical distance threshold and a signal correlation threshold. For example, the physical distance threshold is set to 15 meters, and the signal correlation threshold is set to a channel signal correlation coefficient greater than 0.8. Based on these preset thresholds, the data aggregation node then quickly filters out other collection node identifiers that meet these threshold conditions from the deployment location database or topology map database to generate a set of candidate neighborhood nodes.
[0148] The data aggregation node then performs a more precise screening based on the sampling channel types of the nodes in the candidate neighborhood node set. Specifically, the data aggregation node checks the data channel type of each node in the candidate neighborhood node set one by one, identifying and screening nodes with the same channel type as the missing or abnormal sampling location, thereby obtaining a more accurate neighborhood for the completed data association.
[0149] For example, in one specific implementation example, a radar node (identified as "radar_05") experiences missing or abnormal data at timestamp "20240610104500" on channel number "channel_01." The data aggregation node first identifies the target node as "radar_05," then deploys a location database using pre-existing nodes, and uses Python's GeoPandas tool or NetworkX graph analysis tool to identify other node sets within a 15-meter physical distance from the "radar_05" node or with a signal correlation greater than 0.8. For example, a preliminary set of candidate nodes is generated:
[0150] Neighborhood candidate node set: ["radar_03", "radar_08", "cam_04", "imu_02"]
[0151] Subsequently, the data aggregation node further checks the channel type of each node and finds that the data channel types of "radar_03" and "radar_08" in the candidate node set are consistent with the target node "radar_05" (both are radar channel types), while the channel types of "cam_04" (visual sensor node) and "imu_02" (inertial measurement unit node) are inconsistent with the target node. Therefore, the latter two are removed from the candidate node set, and finally a more accurate association neighborhood is obtained:
[0152] Determine the set of associated neighboring nodes: ["radar_03", "radar_08"]
[0153] The data aggregation node then outputs the above-determined set of associated neighboring nodes as the data input basis for subsequent steps such as completing the reference data selection and building the memory state sequence.
[0154] The above-mentioned optimized method for determining the associated neighborhood can effectively improve the accuracy and stability of subsequent data completion, avoid the interference of irrelevant or weakly relevant data in the completion process, and improve the quality and reliability of the overall fusion analysis.
[0155] As an optional implementation, the step of constructing a memory state sequence includes:
[0156] Determining, based on the first index key value corresponding to the supplemented reference data, a collection node identifier corresponding to the supplemented reference data;
[0157] Retrieving, from the multidimensional data structure, historical data that appears continuously with the supplementary reference data within a preset time range according to the acquisition node identifier and the first index key value, to generate a first historical data sequence;
[0158] Retrieving, from the multidimensional data structure, the historical data of nodes adjacent to the complementary reference data in a spatial or channel dimension according to the physical location or channel characteristics of the acquisition node corresponding to the complementary reference data, to generate a second historical data sequence;
[0159] The first historical data sequence and the second historical data sequence are combined into a memory state sequence corresponding to the supplementary reference data, and the memory state sequence is output.
[0160] During the construction of the memory state sequence, the historical data corresponding to the supplementary reference data may be discontinuous or insufficient in time sequence, and the selection of historical data fails to fully consider spatial or channel correlations, which may reduce the reliability and accuracy of the subsequent data completion process. This embodiment effectively improves the integrity and relevance of the memory state sequence by more carefully selecting and combining the historical data sequences for the supplementary reference data, thereby improving the accuracy of data completion.
[0161] During the specific implementation process, the data aggregation node first clearly determines the specific collection node identifier corresponding to the completed reference data based on the first index key value of the completed reference data to ensure the accuracy of historical data retrieval.
[0162] Subsequently, based on the determined acquisition node identifier and the corresponding index key value, the data aggregation node retrieves and obtains historical data of the supplementary reference data that appears continuously within a preset historical time window from a pre-constructed multidimensional data structure, thereby forming a first historical data sequence. In a specific implementation, the data aggregation node can quickly retrieve the above historical data from a time series database or an internal index data structure based on the node identifier and time index. The time series length of the data can be set to, for example, the most recent 10 to 20 consecutive sampling periods to provide sufficient data support.
[0163] Furthermore, to fully account for spatial or channel correlations, the data aggregation node also identifies other acquisition nodes adjacent to the supplemented reference data in the spatial or channel dimensions based on the physical deployment location or sensor channel type of the acquisition node corresponding to the supplemented reference data. The data aggregation node rapidly obtains the identities of spatially adjacent nodes or nodes with similar channels from a node deployment database or a predefined node topology map. Based on this, the data aggregation node further retrieves historical data from these adjacent nodes within the same historical time window from the multidimensional data structure to form a second historical data sequence.
[0164] During the formation of the above-mentioned second historical data sequence, the data of the selected nodes can be further weighted or filtered according to spatial or channel correlation to highlight highly correlated data and improve the accuracy of subsequent memory state sequences.
[0165] Finally, the data aggregation node combines the first and second historical data sequences based on precise timestamp alignment to form a complete memory state sequence corresponding to the completed reference data. This combined memory state sequence contains the historical continuous observation data of the completed reference data itself and the historical observation data of adjacent nodes in space or channels. It is output to the next completion processing step in a clear data structure such as a table, array, or standard data object.
[0166] In this way, the problem of insufficient or irrelevant historical data sequences is effectively overcome, and the quality and reliability of the memory state sequences used in the data completion process are significantly improved.
[0167] As an optional implementation, performing memory state reconstruction and calculating the candidate data for completion of each missing or abnormal position and the corresponding confidence weight includes:
[0168] Determining, based on the first historical data sequence and the second historical data sequence in the memory state sequence, scale features corresponding to different sensor modalities in the historical data sequence, and calculating a corresponding confidence level of each modality historical data sequence at a current moment;
[0169] Calculating the historical window importance weight of the historical data sequence relative to the current missing or abnormal position based on the time difference between the acquisition time of the historical data sequence and the current missing or abnormal position;
[0170] Calculate the memory weight for each historical data sequence according to the scale feature, the corresponding confidence level, and the historical window importance weight;
[0171] Performing memory state reconstruction on the historical data sequence based on the memory weights to generate completion candidate data of different modalities, and calculating a fusion confidence weight corresponding to each of the completion candidate data based on a combination of the memory weights;
[0172] Output the completion candidate data of the different modalities and the fusion confidence weight.
[0173] When performing memory state reconstruction, the data aggregation node further refines the data completion process for each missing or abnormal location to improve the accuracy and reliability of the data completion results.
[0174] During implementation, the data aggregation node first determines the scale characteristics of different sensor modalities based on the acquired memory state sequence, which includes the historical sequence of the reference data itself and the historical sequences of spatial or channel neighboring nodes. For the visual modality, the scale characteristics can be selected as the stability of image feature points or the grayscale variation range. For the radar modality, the scale characteristics can be selected as the target detection range or the echo intensity fluctuation range. For the inertial measurement unit modality, the scale characteristics can be determined based on the statistical fluctuations of acceleration or angular velocity. Determination of scale characteristics can generally be accomplished using standard statistical analysis methods.
[0175] Subsequently, the data aggregation node further calculates the corresponding confidence level of each historical data sequence at the current moment. The corresponding confidence level is determined based on factors such as the stability of the historical data, the missing ratio or the abnormal data ratio. The higher the stability and the fewer abnormal data, the higher the corresponding confidence level.
[0176] In addition, to reflect the timeliness of the data, the data aggregation node calculates the importance weight of the historical window based on the time difference between the collection time of the historical data sequence and the current missing or abnormal position. The weight of data closer to the current moment is relatively high, while the weight of earlier data gradually decreases. The weight of the historical window can be determined based on a predefined attenuation law.
[0177] Based on the above-mentioned scale characteristics, corresponding confidence levels, and historical window importance weights, the data aggregation node further calculates the comprehensive memory weight for each historical data sequence. The memory weight comprehensively reflects the influence of data scale, modal reliability, and time factors, and can be implemented using linear combinations or more advanced adaptive weighting methods.
[0178] Next, the data aggregation node performs memory state reconstruction on the historical data sequence based on the calculated memory weights to obtain completion candidate data corresponding to different modalities. The memory state reconstruction method used may include but is not limited to a prediction network based on a Transformer structure, a multi-input LSTM neural network, or a multimodal attention network, etc., to effectively capture the correlation between historical data of different modalities and generate accurate completion candidate data.
[0179] The completion candidate data generated by each modality are further calculated based on the corresponding memory weight combination to obtain their own fusion confidence weight, where the fusion confidence weight reflects the weight of each modality candidate data in the final completion fusion process to support the accurate calculation of the final fusion result.
[0180] For example, when the data aggregation node processes the missing data of a certain visual sensor, it is clear that the scale feature fluctuation amplitude of the historical data of the visual modality is small, the corresponding confidence level is high, and the historical window importance weight is large; while the scale feature fluctuation amplitude of the adjacent radar modality is large, the modal confidence level is relatively low, and the historical window weight is relatively small. Therefore, the memory weight calculated after comprehensive calculation is more inclined to the visual modality data, so that the visual modality completion candidate data plays a greater role in fusion, thereby improving the reliability of the final completion data.
[0181] It is understood that relative terms such as "smaller," "higher," and "larger" are qualitative expressions commonly used by those skilled in the art in data quality assessment and multimodal fusion practice. Their meanings can be understood in conjunction with algorithmic parameters such as scale characteristics, confidence levels, and historical window weights as described in the context. Those skilled in the art can determine appropriate thresholds or comparison benchmarks based on specific application scenarios through conventional statistical analysis, error assessment, or experimental calibration.
[0182] Finally, the data aggregation node outputs the above-generated candidate data for each modality completion and the corresponding fusion confidence weight. The output data structure can adopt a standard array, data table or other standardized data formats commonly used in other projects for subsequent fusion analysis and other system calls.
[0183] It is understandable that the above-mentioned statistical analysis, weight calculation, model construction and fusion processing can be achieved through common numerical analysis tools, statistical function libraries and deep learning frameworks.
[0184] In this way, the present application improves the accuracy and reliability of the data completion process and optimizes the overall effect of multi-source data fusion analysis.
[0185] As an optional implementation, generating the multidimensional data structure after completion processing includes:
[0186] Dividing the multidimensional data structure into a plurality of preset local spatial regions according to the fusion confidence weights, and respectively calculating the variance value of the fusion confidence weights in each of the local spatial regions;
[0187] Comparing the variance value with a preset island threshold, and determining a first local area exceeding the island threshold from the local spatial area;
[0188] For the first local area, according to a preset extended neighborhood rule, determining a second local area adjacent to the first local area in the multidimensional data structure, and obtaining candidate modality data and a fusion confidence weight of the second local area;
[0189] Merging the candidate modality data and fusion confidence weight of the second local area with the candidate modality data and fusion confidence weight of the first local area to generate an extended fusion data set, and recalculating fusion data of a corresponding position of the first local area based on the extended fusion data set;
[0190] The recalculated fusion data are written one by one into the positions corresponding to the first local areas in the multidimensional data structure to generate a multidimensional data structure after extended neighborhood fusion processing.
[0191] In order to further improve the continuity and stability of local area data fusion in the multidimensional data structure after completion processing, the data aggregation node executes a local area expansion fusion method based on fusion confidence weight to effectively avoid the local discontinuity of the fused data caused by excessive differences in local data confidence, namely the "island" phenomenon.
[0192] In the specific implementation, the data aggregation node first divides the entire multidimensional data structure into several preset local spatial regions based on the fusion confidence weights calculated in the previous steps. The local spatial regions can be predefined based on the spatial location distribution of nodes or the layout of data channels, for example, according to a certain spatial range or data collection node group.
[0193] Subsequently, the data aggregation node calculates the variance of all fusion confidence weights within each local spatial region to reflect the consistency level of fusion data confidence within the local spatial region. Among them, areas with higher variance values indicate that the data confidence in the region is significantly different, and there may be potential data fusion discontinuity.
[0194] The data aggregation node then compares the calculated variance values for each local spatial region with a predefined island threshold. This threshold can be determined based on historical experience with data fusion, actual environmental requirements, or statistical analysis. If the variance value of the fusion confidence weight for a local region exceeds the predefined island threshold, the data aggregation node marks that region as the first local region requiring additional fusion processing.
[0195] After determining the first local area, the data aggregation node determines a second local area that is spatially adjacent to or signal-linked to the first local area based on a preset extended neighborhood rule. The extended neighborhood rule can be predefined based on physical spatial proximity or sensor data channel correlation.
[0196] After determining the second local area, the data aggregation node obtains the data in the second local area from the multidimensional data structure, including the candidate modal data and the corresponding fusion confidence weights of the second local area, and merges these data with the existing data in the first local area (including the candidate modal data and the fusion confidence weights) to form a complete extended fusion data set.
[0197] The data aggregation node then recalculates the fused data for each location within the first local area based on the expanded fused data set. When recalculating the fused data, the aforementioned weighted fusion algorithm or a more sophisticated adaptive fusion method can be used to fully utilize all the data features and confidence weight information in the expanded data set, thereby reducing local discontinuities in the fusion of the original regional data.
[0198] After the recalculation is completed, the data aggregation node writes the updated fused data back to the position corresponding to the first local area in the multidimensional data structure one by one, thereby obtaining an updated version of the multidimensional data structure that has undergone extended neighborhood fusion processing.
[0199] For example, within a radar node distribution area within a multidimensional data structure, the data aggregation node discovers that the variance of the fusion confidence weight within the area exceeds the island threshold, indicating that the data confidence level in this area varies significantly. The data aggregation node then marks this area as a first local area. Subsequently, based on the principle of physical proximity, the node determines the spatial area to which other radar and visual sensor nodes adjacent to the first local area belong as a second local area, extracts the data and confidence weights of these neighboring nodes, and merges them with the data from the first local area to form an extended fusion data set. Next, the data aggregation node uses this extended data set to recalculate the fusion data within the first local area, thereby reducing the local differences in the data fusion in this area and ultimately generating a multidimensional data structure with better fusion continuity.
[0200] The various statistical analysis calculations and data fusion algorithms in the above implementation process can be implemented through commonly used numerical analysis tool libraries, statistical tools or deep learning frameworks in engineering, such as numerical libraries or machine learning frameworks in the Python environment. The choice of specific tools depends on actual needs.
[0201] Through this extended neighborhood fusion implementation method, the present application effectively solves the "island" problem of local area data fusion in multidimensional data structures, significantly enhances the spatial consistency and fusion reliability of data, and thus improves the quality of overall multi-source data fusion processing and analysis.
[0202] As an optional implementation, the distributed processing of the completed multidimensional data structure includes:
[0203] Calculating node allocation parameters for each spatial region based on the fusion confidence weight corresponding to each spatial region in the completed multidimensional data structure;
[0204] Performing spatial region division on the completed multidimensional data structure according to the node allocation parameters to generate local data subsets corresponding to a plurality of processing nodes;
[0205] Sending the local data subsets to corresponding processing nodes respectively, and having the processing nodes perform local calculations on the local data subsets respectively to generate local calculation result data;
[0206] Based on the preset cross-node synchronization rules, the local calculation result data is transmitted to the preset central coordination node, and the central coordination node performs the synchronous fusion calculation of the local calculation result data to generate a multi-dimensional data structure after cross-node synchronous fusion.
[0207] In order to efficiently and accurately process large-scale multi-dimensional data structures after completion processing, this application further provides a refined distributed data processing method, which improves the parallelism and fusion consistency of data processing through a clear node allocation and synchronization mechanism.
[0208] During implementation, the data aggregation node first determines the node allocation parameters for each spatial region based on the fusion confidence weights of each spatial region within the completed multidimensional data structure. Specifically, the data aggregation node performs a refined statistical analysis of the fusion confidence weights of each region, including but not limited to calculating the average, maximum, minimum, standard deviation, or data magnitude characteristics of the regional confidence weights, and uses this to determine the scale of computing resources required for each region. The higher the regional fusion confidence weight or the greater the data volatility, the higher the corresponding allocated computing resource parameters to fully ensure computational accuracy and efficiency.
[0209] Furthermore, the data aggregation node performs a clear spatial region partitioning operation on the completed multidimensional data structure based on the node allocation parameters calculated above, thereby forming a series of local data subsets corresponding to the multiple processing nodes. The specific implementation method for spatial region partitioning can adopt a region partitioning algorithm based on a space-filling curve (such as a Hilbert curve or a Z-order curve) or a clustering-based algorithm (such as a K-means clustering algorithm). Through this method, the data aggregation node can accurately map each spatial region in the data structure to a specific processing node.
[0210] After completing the spatial area division, the data aggregation node sends the local data subsets obtained by division to the corresponding multiple processing nodes respectively, wherein the data subset transmission process can adopt a preset efficient data transmission link, including but not limited to Ethernet, distributed file system or remote procedure call mechanism, such as gRPC or RESTful API interface. Each processing node immediately executes a clear local computing task after receiving the data subset, wherein the local computing task specifically includes but is not limited to local extraction of data features, local data statistical analysis or model reasoning based on local data. Among them, local data statistical analysis can, for example, calculate the mean, variance, and peak characteristics of regional data; model reasoning based on local data can, for example, use a locally pre-deployed deep learning model for target recognition or anomaly detection.
[0211] In specific implementations, each processing node performs local computations on the received local data subset and generates corresponding local computation result data. Specifically, the local computation result data can include feature vectors of each region's data, local statistical analysis results, or inference output data from a local model. These data are stored locally in standardized data structures such as data arrays or data tables for subsequent synchronous fusion.
[0212] Furthermore, after completing their local computation tasks, each processing node transmits the local computation results to a pre-designated central coordination node according to pre-set cross-node synchronization rules. These cross-node synchronization rules may include, but are not limited to, specifying synchronization time windows, data transmission formats, and coordination strategies for the synchronization process. For example, each processing node may be required to transmit data to a central node within a specific period, or a distributed message queue (such as Kafka or RabbitMQ) may be used to achieve efficient and real-time data synchronization and message notifications.
[0213] After receiving the local computation results from all processing nodes, the central coordination node further performs cross-node synchronous fusion calculations to achieve comprehensive integration of the overall data structure. Specific fusion calculation implementation methods may include, but are not limited to, weighted averaging algorithms, Kalman filtering algorithms, confidence-based adaptive fusion algorithms, or other specific fusion algorithms to fully utilize the local computation results provided by each node.
[0214] Exemplarily, a data aggregation node first determines node allocation parameters based on the standard deviation of each region's fusion confidence weights and the data volume after completing the processed data from visual sensors, radar sensors, and lidar sensors distributed across a large area. The node then uses the Hilbert curve space partitioning method to map the region's data into four sub-regions, distributing them to four independent processing nodes. Each processing node performs local statistical analysis and feature extraction on the received data subset. For example, the visual node performs image feature point extraction, the radar node performs local calculations of target position and velocity information, and the lidar node performs local density analysis of the point cloud. After the calculations are complete, each node transmits the local calculation results to the central coordination node within a predetermined synchronization period based on preset synchronization rules. The central coordination node then uses a weighted average fusion algorithm to perform comprehensive calculations on the data from each node, resulting in an integrated multidimensional data structure with more comprehensive fusion and higher confidence.
[0215] It can be understood that the specific implementation methods of various statistical analyses, spatial divisions, data synchronization and fusion calculations involved in the above implementation process can be implemented through statistical analysis tool libraries, spatial division algorithm libraries, distributed data communication frameworks or data fusion algorithm tools commonly used in engineering. Among them, the choice of specific implementation tools depends on the actual system requirements and the specific implementation environment.
[0216] Through the above-mentioned specific and refined distributed processing and synchronous fusion implementation methods, this application effectively solves the data processing bottleneck problem in large-scale multi-source data fusion analysis, realizes efficient, accurate and consistent multi-dimensional data fusion processing, and greatly improves the efficiency and quality of overall data fusion analysis.
[0217] As an optional implementation, performing a multi-source fusion analysis operation based on the completed multidimensional data structure and the indicator information, and outputting fusion analysis result data and fusion analysis confidence indicators includes:
[0218] Calculating feature vectors of multiple heterogeneous sensor modalities based on the completed multidimensional data structure, and appending the indicator information to the feature vectors respectively to generate a multimodal feature set with confidence indicators;
[0219] Constructing a preset multimodal analysis model for fusion analysis based on the multimodal feature set, and determining input weight parameters of the multimodal analysis model according to the indicator information;
[0220] Inputting the multimodal feature set into the multimodal analysis model, performing a multimodal joint analysis operation, and obtaining fusion analysis result data;
[0221] Calculating a fusion analysis confidence index of the fusion analysis result data according to the model output parameters corresponding to the fusion analysis result data;
[0222] The fusion analysis result data and the fusion analysis confidence index are output as the final output of the multi-source data fusion distribution processing and analysis method.
[0223] In order to achieve efficient and accurate multi-source fusion analysis of the multi-dimensional data structure after completion and synchronous fusion, this application further implements a refined multimodal fusion analysis method to improve the reliability, comprehensiveness and accuracy of data analysis.
[0224] During the specific implementation process, the data aggregation node first performs refined feature vector calculations on various heterogeneous sensor modal data, such as vision, radar, lidar, and inertial measurement units, based on the complete multidimensional data structure obtained through cross-node synchronous fusion processing. Feature vectors for the visual modality can be extracted using convolutional neural networks or traditional image processing algorithms, including but not limited to the target's edge features, texture features, or semantic features; feature vectors for the radar modality can be extracted using target detection and tracking algorithms, including features such as the target's distance, speed, and azimuth; three-dimensional spatial features can be extracted for the lidar modality through point cloud clustering, density analysis, or spatial morphology analysis algorithms; and motion features can be extracted for the inertial measurement unit modality through trajectory fitting or motion state analysis algorithms. The feature vectors for each modality have a clear physical meaning and a standardized numerical expression.
[0225] Furthermore, the data aggregation node appends the metrics calculated in the previous steps to the feature vectors of the corresponding modalities, thereby generating a multimodal feature set with explicit confidence indicators. For example, the metrics can be the fusion residual of the completion process, confidence weights, or data stability indicators.
[0226] These feature sets with accompanying indicator information enable subsequent fusion analysis to be effectively weighted according to the reliability of the data.
[0227] Subsequently, the data aggregation node constructs or selects a preset multimodal fusion analysis model based on the multimodal feature set generated above. The multimodal analysis model may include but is not limited to a multimodal Transformer model, a multi-input long short-term memory network, or a graph neural network model based on a graph attention mechanism. The specific selection depends on the application scenario requirements and the form of feature data. When constructing the model, the data aggregation node determines the input weight parameters of the model based on the accompanying indicator information, wherein modal features with higher confidence are assigned larger weight parameters, and modal features with lower confidence correspond to smaller input weight parameters, thereby realizing adaptive weighted fusion of input data.
[0228] Next, the data aggregation node inputs the aforementioned multimodal feature set with confidence indicators into the pre-defined multimodal fusion analysis model, performs a specific multimodal joint analysis operation, and obtains the fusion analysis result data. The multimodal joint analysis operation process includes, but is not limited to, specific analysis tasks such as target detection, recognition, classification, tracking, environmental state estimation, or event prediction. The specific analysis task is determined based on the actual application scenario requirements.
[0229] Furthermore, the data aggregation node calculates a fusion analysis confidence index for the fusion analysis results based on the model output parameters corresponding to the fusion analysis result data. Specifically, the fusion analysis confidence index can be determined based on the posterior probability distribution, output variance, output residual, or prediction uncertainty of the model output. The fusion analysis confidence index clearly represents the reliability and trustworthiness of the fusion analysis result data; higher values indicate greater reliability of the result data.
[0230] For example, when a data aggregation node processes a multi-source data fusion analysis task within a certain area, it first extracts feature vectors for the visual data, radar data, and lidar data, and appends previously acquired confidence indicator information. Subsequently, a multimodal Transformer analysis model is constructed, in which the input feature weight of the visual modality is the highest (e.g., a confidence weight of 0.7), followed by the radar modality (e.g., 0.2), and the lidar modality has a lower weight (e.g., 0.1). After a clear multimodal joint analysis operation, the model outputs fusion analysis result data, such as the recognition and classification results of specific targets and the estimation results of environmental conditions. Finally, the data aggregation node further calculates the fusion analysis confidence indicators, such as a target recognition confidence level of 95% and a variance of the environmental condition estimation result below a preset threshold, indicating that the fusion analysis result has a high degree of reliability.
[0231] Ultimately, the data aggregation node outputs the fusion analysis result data and the corresponding fusion analysis confidence indicators in a clear standardized data format, such as outputting to a standardized data table, JSON object or a specific database system, so that other system components or applications can call, display or further analyze and process them.
[0232] It is understandable that the specific technical means involved in the above process, such as feature extraction, indicator calculation, model construction, fusion analysis and confidence indicator calculation, can be implemented through feature extraction algorithm libraries, numerical analysis tools, deep learning frameworks or statistical analysis tools commonly used in engineering. The selection of specific tools is flexibly determined according to the actual deployment environment and needs.
[0233] In this way, the present application effectively realizes the precision and automation of the multi-source data fusion analysis process, significantly improves the reliability and comprehensiveness of the fusion analysis results, ensures the efficiency, accuracy and stability of the final data analysis output, and overcomes the defects of one-sided information and insufficient confidence in traditional single modal analysis methods.
[0234] Based on the same inventive concept, the embodiment of the present application also provides a multi-source data fusion distributed processing and analysis system corresponding to a multi-source data fusion distributed processing and analysis method. Since the principle of solving the problem by the system in the embodiment of the present application is similar to the above-mentioned multi-source data fusion distributed processing and analysis method in the embodiment of the present application, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be repeated.
[0235] Reference Figure 4 FIG. 1 is a schematic diagram of a multi-source data fusion distribution processing and analysis system provided by the present application, the system comprising:
[0236] The acquisition module 10 is used to acquire heterogeneous sensor data from multiple distributed acquisition nodes, wherein the heterogeneous sensor data includes acquisition time, acquisition channel, and source node identification information;
[0237] A construction module 20 is configured to construct a multidimensional data structure for representing the heterogeneous sensor data and identify missing or abnormal sampling locations in the multidimensional data structure;
[0238] A first completion module 30 is configured to determine, based on the missing or abnormal sampling locations identified in the multidimensional data structure, an associated neighborhood of the completion data, and select, based on a preset neighborhood threshold, neighborhood data with a correlation degree greater than or equal to the neighborhood threshold as completion reference data, thereby obtaining historical continuous observation data of the completion reference data in the time and space dimensions, and constructing a memory state sequence;
[0239] A second completion module 40 is configured to perform memory-based state reconstruction based on the memory state sequence, calculate candidate completion data and corresponding confidence weights for each missing or abnormal position, fuse multiple candidate completion data based on the confidence weights, generate a completed multidimensional data structure, and calculate indicator information reflecting the reliability of the completion process;
[0240] A processing module 50 is configured to perform distributed processing on the multidimensional data structure that has been completed, wherein the distributed processing includes: dividing the multidimensional data structure into different processing nodes and performing cross-node synchronization;
[0241] The output module 60 is used to perform a multi-source fusion analysis operation based on the completed multi-dimensional data structure and the indicator information, and output fusion analysis result data and fusion analysis confidence indicators.
[0242] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
Claims
1. A multi-source data fusion distribution processing and analysis method, characterized in that: include: Acquire heterogeneous sensor data from multiple distributed collection nodes, wherein the heterogeneous sensor data includes collection time, collection channel, and source node identification information; Constructing a multidimensional data structure for representing the heterogeneous sensor data, and identifying missing or abnormal sampling locations in the multidimensional data structure; Based on the missing or abnormal sampling positions identified in the multidimensional data structure, determining the associated neighborhood of the complementary data, and selecting, based on a preset neighborhood threshold, neighborhood data with a correlation degree greater than or equal to the neighborhood threshold as complementary reference data, and then obtaining historical continuous observation data of the complementary reference data in the time dimension and the space dimension, and constructing a memory state sequence; Based on the memory state sequence, perform memory state reconstruction, calculate the completion candidate data and corresponding confidence weight for each missing or abnormal position, fuse multiple completion candidate data based on the confidence weight, generate a multidimensional data structure after completion processing, and calculate indicator information used to reflect the reliability of the completion process; Performing distributed processing on the completed multidimensional data structure, the distributed processing comprising: dividing the multidimensional data structure into different processing nodes and performing cross-node synchronization; Based on the completed multidimensional data structure and the indicator information, a multi-source fusion analysis operation is performed to output fusion analysis result data and fusion analysis confidence indicators.
2. A multi-source data fusion distribution processing and analysis method according to claim 1, characterized in that: The constructing of a multidimensional data structure for representing the heterogeneous sensor data comprises: Generate a first index key value corresponding to the heterogeneous sensor data based on the acquisition time, acquisition channel, and source node identifier carried in the heterogeneous sensor data; performing a hash mapping operation on the heterogeneous sensor data based on the first index key value to determine a storage location of the heterogeneous sensor data in a preset sparse tensor; According to the storage locations, the heterogeneous sensor data are written one by one into corresponding locations of the preset sparse tensor to generate a multidimensional data structure representing the heterogeneous sensor data.
3. The multi-source data fusion distribution processing and analysis method according to claim 2 is characterized in that: Determining the associated neighborhood of the complementary data includes: Determine a target acquisition node identifier corresponding to the missing or abnormal sampling position according to a first index key value corresponding to the missing or abnormal sampling position in the multidimensional data structure; Based on the target acquisition node identifier and a preset physical distance threshold or signal correlation threshold, selecting other acquisition node identifiers whose physical distances or signal correlation degrees meet the preset physical distance threshold or signal correlation threshold from the multidimensional data structure to generate a neighborhood candidate node set; Based on the consistency of the sampling channel types of the neighborhood candidate node set and the sampling channel types corresponding to the missing or abnormal sampling position, candidate nodes with the same channel type as the missing or abnormal sampling position are screened from the neighborhood candidate node set as associated neighbors of the supplementary data, and the associated neighbors are output for selection and use by the supplementary reference data.
4. The multi-source data fusion distribution processing and analysis method according to claim 3 is characterized in that: The construction memory state sequence includes: Determining, based on the first index key value corresponding to the supplemented reference data, a collection node identifier corresponding to the supplemented reference data; Retrieving, from the multidimensional data structure, historical data that appears continuously with the supplementary reference data within a preset time range according to the acquisition node identifier and the first index key value, to generate a first historical data sequence; Retrieving, from the multidimensional data structure, the historical data of nodes adjacent to the complementary reference data in a spatial or channel dimension according to the physical location or channel characteristics of the acquisition node corresponding to the complementary reference data, to generate a second historical data sequence; The first historical data sequence and the second historical data sequence are combined into a memory state sequence corresponding to the supplementary reference data, and the memory state sequence is output.
5. The multi-source data fusion distribution processing and analysis method according to claim 4 is characterized in that: The execution of memory state reconstruction and calculation of candidate data for completion of each missing or abnormal position and the corresponding confidence weight includes: Determining, based on the first historical data sequence and the second historical data sequence in the memory state sequence, scale features corresponding to different sensor modalities in the historical data sequence, and calculating a corresponding confidence level of each modality historical data sequence at a current moment; Calculating the historical window importance weight of the historical data sequence relative to the current missing or abnormal position based on the time difference between the acquisition time of the historical data sequence and the current missing or abnormal position; Calculate the memory weight for each historical data sequence according to the scale feature, the corresponding confidence level, and the historical window importance weight; Performing memory state reconstruction on the historical data sequence based on the memory weights to generate completion candidate data of different modalities, and calculating a fusion confidence weight corresponding to each of the completion candidate data based on a combination of the memory weights; Output the completion candidate data of the different modalities and the fusion confidence weight.
6. The multi-source data fusion distribution processing and analysis method according to claim 5 is characterized in that: Generating a multidimensional data structure after completion processing includes: Dividing the multidimensional data structure into a plurality of preset local spatial regions according to the fusion confidence weights, and respectively calculating the variance value of the fusion confidence weights in each of the local spatial regions; Comparing the variance value with a preset island threshold, and determining a first local area exceeding the island threshold from the local spatial area; For the first local area, according to a preset extended neighborhood rule, determining a second local area adjacent to the first local area in the multidimensional data structure, and obtaining candidate modality data and a fusion confidence weight of the second local area; Merging the candidate modality data and fusion confidence weight of the second local area with the candidate modality data and fusion confidence weight of the first local area to generate an extended fusion data set, and recalculating fusion data of a corresponding position of the first local area based on the extended fusion data set; The recalculated fusion data are written one by one into the positions corresponding to the first local areas in the multidimensional data structure to generate a multidimensional data structure after extended neighborhood fusion processing.
7. The multi-source data fusion distribution processing and analysis method according to claim 6 is characterized in that: The distributed processing of the completed multidimensional data structure includes: Calculating node allocation parameters for each spatial region based on the fusion confidence weight corresponding to each spatial region in the completed multidimensional data structure; Performing spatial region division on the completed multidimensional data structure according to the node allocation parameters to generate local data subsets corresponding to a plurality of processing nodes; Sending the local data subsets to corresponding processing nodes respectively, and having the processing nodes perform local calculations on the local data subsets respectively to generate local calculation result data; Based on the preset cross-node synchronization rules, the local calculation result data is transmitted to the preset central coordination node, and the central coordination node performs the synchronous fusion calculation of the local calculation result data to generate a multi-dimensional data structure after cross-node synchronous fusion.
8. The multi-source data fusion distribution processing and analysis method according to claim 7 is characterized in that: The performing of multi-source fusion analysis operations based on the completed multidimensional data structure and the indicator information, and outputting fusion analysis result data and fusion analysis confidence indicators includes: Calculating feature vectors of multiple heterogeneous sensor modalities based on the completed multidimensional data structure, and appending the indicator information to the feature vectors respectively to generate a multimodal feature set with confidence indicators; Constructing a preset multimodal analysis model for fusion analysis based on the multimodal feature set, and determining input weight parameters of the multimodal analysis model according to the indicator information; Inputting the multimodal feature set into the multimodal analysis model, performing a multimodal joint analysis operation, and obtaining fusion analysis result data; Calculating a fusion analysis confidence index of the fusion analysis result data according to the model output parameters corresponding to the fusion analysis result data; The fusion analysis result data and the fusion analysis confidence index are output as the final output of the multi-source data fusion distribution processing and analysis method.
9. A multi-source data fusion distribution processing and analysis system, characterized in that: include: An acquisition module is used to acquire heterogeneous sensor data from multiple distributed acquisition nodes, wherein the heterogeneous sensor data includes acquisition time, acquisition channel, and source node identification information; A construction module, configured to construct a multidimensional data structure for representing the heterogeneous sensor data and identify missing or abnormal sampling locations in the multidimensional data structure; a first completion module, configured to determine, based on the missing or abnormal sampling locations identified in the multidimensional data structure, an associated neighborhood of the completion data, and select, based on a preset neighborhood threshold, neighborhood data with a degree of association greater than or equal to the neighborhood threshold as completion reference data, thereby obtaining historical continuous observation data of the completion reference data in time and space dimensions, and constructing a memory state sequence; a second completion module, configured to perform memory-based state reconstruction based on the memory state sequence, calculate completion candidate data and corresponding confidence weights for each missing or abnormal position, fuse multiple completion candidate data based on the confidence weights, generate a completed multidimensional data structure, and calculate indicator information reflecting the reliability of the completion process; a processing module, configured to perform distributed processing on the multidimensional data structure that has undergone the completion processing, wherein the distributed processing includes: dividing the multidimensional data structure into different processing nodes and performing cross-node synchronization; The output module is used to perform multi-source fusion analysis operations based on the completed multidimensional data structure and the indicator information, and output fusion analysis result data and fusion analysis confidence indicators.
Citation Information
Patent Citations
Crop growth prediction method based on multi-source data fusion analysis
CN119398284A
System detection method and apparatus based on multi-source heterogeneous data
WO2024148880A1