A spatiotemporal reference unified multi-source heterogeneous data real-time processing and collaborative computing method
By constructing a unified spatiotemporal coordinate system, multimodal feature encoding, and edge-cloud collaborative computing, the problems of inconsistent spatiotemporal benchmarks and high processing latency of multi-source heterogeneous data in complex dynamic scenarios are solved, realizing accurate fusion and real-time scheduling of multi-source data, and improving scenario operation efficiency and decision accuracy.
Patent Information
- Application Number
- CN202610732366.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-06-23
AI Technical Summary
Existing technologies suffer from problems such as inconsistent spatiotemporal benchmarks, inconsistent feature representations, high processing latency, weak concurrency support capabilities, and poor versatility in processing and collaborative computing of multi-source heterogeneous data in complex dynamic scenarios, making it difficult to achieve accurate fusion and real-time scheduling of multi-source data.
By constructing a unified spatiotemporal coordinate system, multimodal feature coding, edge-cloud collaborative computing, and digital twin dynamic scheduling, we achieve unified spatiotemporal reference, low-latency processing, and high-concurrency support for multi-source heterogeneous data. We use BeiDou spatiotemporal reference calibration, lightweight dynamic feature coding network, edge-cloud collaborative computing, and digital twin engine for data processing and scheduling.
It achieves millisecond-level spatiotemporal synchronization of multi-source heterogeneous data, improves the accuracy and robustness of cross-modal data processing, reduces processing latency, enhances concurrency capabilities, supports precise full-domain control and intelligent decision-making, and has good versatility and scalability.
Smart Images

Figure CN122262491A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of multi-source data processing, collaborative computing, and spatiotemporal benchmark calibration, and in particular to a method for real-time processing and collaborative computing of multi-source heterogeneous data for complex dynamic scenarios. It is applicable to multi-source data fusion, real-time computing, and precise control in various complex dynamic scenarios, and can be widely used in multiple fields such as port scheduling, large-scale scenario control, and industrial plant management, providing core technical support for the efficient operation and safe control of scenarios. Background Technology
[0002] With the rapid development of technologies such as the Internet of Things (IoT), artificial intelligence (AI), and satellite positioning, the number of sensing devices deployed in various complex and dynamic scenarios (such as ports, large public areas, and industrial plants) is constantly increasing, forming a multi-source heterogeneous data acquisition system covering various types of devices, including radar, cameras, IoT sensors, and positioning terminals. This multi-source data contains rich information such as target status, environmental parameters, and equipment operation status within the scenario, and is the core foundation for achieving precise scenario control and efficient scheduling.
[0003] However, significant technical bottlenecks still exist in the processing and collaborative computing of multi-source heterogeneous data in complex dynamic scenarios: First, various sensing devices have different spatiotemporal references and inconsistent coordinate systems and timestamps, forming data silos. This can easily lead to mismatches in cross-modal target matching and distortions in scene state assessment, directly affecting the accuracy of scheduling decisions. Second, the feature dimensions and representation forms of multimodal data such as radar point clouds, video streams, and IoT sensing vary greatly. Existing technologies lack efficient and unified encoding and fusion methods, resulting in poor performance in cross-modal target matching, continuous tracking, and accurate identification. Third, traditional centralized architectures require uploading massive amounts of raw data to the cloud, resulting in high bandwidth consumption and processing latency, making it difficult to meet the real-time scheduling and response requirements of dynamic scenarios and easily causing instruction lag and security risks. Fourth, traditional architectures have limited concurrent carrying capacity, making it difficult to support the simultaneous access of tens of thousands of sensing devices, easily leading to data congestion and a sharp drop in processing efficiency. Fifth, existing solutions are mostly designed for single, dedicated scenarios, lacking versatility and portability. Core technologies are difficult to reuse across scenarios, resulting in high R&D costs and limited applicability. Existing improvement methods can only partially optimize a single problem and cannot simultaneously meet the comprehensive needs of unified spatiotemporal benchmarks, low-latency real-time computing, high-concurrency access support, and cross-scenario universal adaptation, highlighting their technical shortcomings.
[0004] Therefore, developing a multi-source heterogeneous data processing and collaborative computing method that can achieve unified spatiotemporal references, efficient fusion of multimodal features, low-latency collaborative computing, high concurrency support, and broad applicability has become a pressing technical challenge for those skilled in the art. Based on accumulated expertise in port multi-source data processing, this invention abstracts and optimizes core technologies such as spatiotemporal reference calibration, dynamic feature encoding, and edge collaborative computing to form a universal multi-source heterogeneous data processing solution. This solution effectively addresses the aforementioned technical bottlenecks and adapts to the application needs of various complex and dynamic scenarios. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing methods for handling multi-source heterogeneous data in complex dynamic scenarios, such as inconsistent spatiotemporal references, inconsistent feature representations, high processing latency, weak concurrency support, and poor versatility. This invention proposes a real-time processing and collaborative computing method for multi-source heterogeneous data with unified spatiotemporal references. By constructing a unified spatiotemporal coordinate system, unifying the encoding of multimodal features, edge-cloud collaborative computing, and dynamic scheduling of digital twins, this method achieves accurate fusion, low-latency processing, and efficient full-domain management of multi-source heterogeneous data, thereby improving scenario operation efficiency and decision-making accuracy.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for real-time processing and collaborative computing of multi-source heterogeneous data with unified spatiotemporal reference includes the following steps:
[0007] 1) Construct a unified spatiotemporal coordinate system, perform spatial coordinate calibration on multi-source heterogeneous sensing devices based on the BeiDou spatiotemporal reference, and convert the coordinates of each device into the WGS84 coordinate system; use the BeiDou time signal as a reference to calibrate the timestamps of each device's data to eliminate time offset between devices; establish a spatiotemporal correlation index to bind multi-source data to the unified spatiotemporal coordinate system, and achieve millisecond-level spatiotemporal accurate alignment of multi-source data.
[0008] 2) Multimodal data cross-modal feature encoding: Lightweight dynamic feature encoding network is used to extract features from radar point cloud, video stream, IoT sensing and BeiDou positioning data respectively; combined with an adaptive weight allocation mechanism, the weight of each modal data is dynamically adjusted according to the scene changes; cross-modal feature weighted fusion is completed through attention mechanism, and a unified dimension fusion feature vector is output to achieve unified representation of multi-source heterogeneous data and accurate cross-modal target matching.
[0009] 3) Edge-cloud collaborative low-latency computing: deploying edge computing nodes to complete local data preprocessing, feature encoding and hierarchical feature compression; the cloud adopts a distributed architecture to achieve parallel fusion of data from multiple edge nodes; adopts an asynchronous federated learning strategy to complete global model optimization while protecting data privacy; and enables concurrent access of devices to reduce end-to-end data processing latency.
[0010] 4) Dynamic scene control and scheduling: Based on fused data under a unified spatiotemporal benchmark, a digital twin engine is integrated to build a real-time scene mapping model; combined with a dynamic operations research optimization algorithm library, scene state deduction and resource scheduling optimization are completed; scheduling instructions are generated and executed through edge nodes to form a closed-loop control, achieving precise scheduling and intelligent decision-making across the entire domain.
[0011] Preferably, the multi-source heterogeneous sensing devices in step 1) include millimeter-wave radar, high-definition camera equipment, IoT sensing terminals, location positioning terminals, and other categories; the spatial coordinate calibration deviation is strictly controlled within a small range, and the time synchronization accuracy achieves a high-precision level standard.
[0012] Preferably, the dynamic feature encoding network in step 2) includes a feature extraction module, an adaptive weight allocation module, and a cross-modal feature fusion module, outputting 128-dimensional unified fused features; cosine similarity is used for target matching; after pruning and quantization optimization, the model can be lightweightly deployed on edge nodes.
[0013] Preferably, in step 3), the edge nodes are set with a moderate coverage range and are responsible for local data preprocessing and feature compression; the cloud uses a distributed cluster to achieve high-concurrency processing; asynchronous federated learning completes global model aggregation and update without uploading the original data.
[0014] Preferably, in step 4), the dynamic operations research optimization algorithm library supports path planning, resource allocation, and risk prediction functions; the scheduling instructions are transmitted in encrypted form, the executing devices respond in real time and provide feedback on their status, and support full-domain one-map visualization management and control.
[0015] Preferably, the present invention also includes data security protection steps, employing a triple mechanism of data encryption, access control, and anomaly monitoring to ensure the security and reliability of multi-source data throughout the entire process of collection, transmission, storage, and processing.
[0016] A real-time processing and collaborative computing system for multi-source heterogeneous data with unified spatiotemporal reference includes: Data input module: used to acquire data from multi-source sensing devices, spatiotemporal parameters, and device configuration information; Spatiotemporal reference calibration module: used to complete equipment coordinate calibration, timestamp calibration and spatiotemporal correlation index establishment; Multimodal feature encoding module: used for feature extraction, adaptive weight allocation, and cross-modal feature fusion in dynamic feature encoding networks; Edge-cloud collaborative computing module: used for data preprocessing, feature compression, global data fusion, and model optimization; Digital Twin and Scheduling Module: Used for scene twin modeling, dynamic simulation, and generation and issuance of scheduling instructions; Constraints and Security Management Module: Used for data anomaly detection, security encryption, and access control; The results output module is used to output fusion features, target matching results, digital twin scenarios, and scheduling instructions.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0018] This invention achieves millisecond-level spatiotemporal synchronization of multi-source heterogeneous data through the unification of BeiDou spatiotemporal reference and timestamp calibration, solving the problem of data spatiotemporal asynchrony and eliminating data silos.
[0019] This invention employs a lightweight dynamic feature encoding network and an adaptive weighting mechanism to achieve unified representation of multimodal data, thereby improving the accuracy and robustness of target matching in complex scenarios.
[0020] This invention achieves low-latency, high-concurrency, and high-privacy-security data processing capabilities through an edge-cloud collaborative architecture, hierarchical feature compression, and asynchronous federated learning.
[0021] This invention combines digital twins with dynamic operations optimization to form a closed-loop scheduling system, supporting full-domain visualized and precise control, and improving the level of intelligent decision-making in various scenarios.
[0022] This invention adopts a universal design, which can be quickly adapted to various complex and dynamic scenarios such as ports, industrial parks, and factories, and has good versatility, practicality and scalability. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the overall process of the method of the present invention; Figure 2 This is a schematic diagram illustrating the construction process of the unified spatiotemporal coordinate system in this invention; Figure 3 This is a schematic diagram of the lightweight dynamic feature encoding network in this invention; Figure 4 This is a schematic diagram illustrating the deployment of the edge-cloud collaborative computing architecture in this invention; Figure 5 This is a schematic diagram of the closed-loop process of scene dynamic control and scheduling in this invention; Figure 6 This is a schematic diagram of a digital twin model of a port scene in an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] This invention provides a technical solution: A method for real-time processing and collaborative computing of multi-source heterogeneous data with unified spatiotemporal reference, the specific implementation of which includes the following steps:
[0026] 1) Construct a unified spatiotemporal coordinate system, perform spatial coordinate calibration on multi-source heterogeneous sensing devices based on the BeiDou spatiotemporal reference, and convert the coordinates of each device into the WGS84 coordinate system; use the BeiDou time signal as a reference to calibrate the timestamps of each device's data to eliminate time offset between devices; establish a spatiotemporal correlation index to bind multi-source data to the unified spatiotemporal coordinate system, and achieve millisecond-level spatiotemporal accurate alignment of multi-source data.
[0027] 2) Multimodal data cross-modal feature encoding: Lightweight dynamic feature encoding network is used to extract features from radar point cloud, video stream, IoT sensing and BeiDou positioning data respectively; combined with an adaptive weight allocation mechanism, the weight of each modal data is dynamically adjusted according to the scene changes; cross-modal feature weighted fusion is completed through attention mechanism, and a unified dimension fusion feature vector is output to achieve unified representation of multi-source heterogeneous data and accurate cross-modal target matching.
[0028] 3) Edge-cloud collaborative low-latency computing: deploying edge computing nodes to complete local data preprocessing, feature encoding and hierarchical feature compression; the cloud adopts a distributed architecture to achieve parallel fusion of data from multiple edge nodes; adopts an asynchronous federated learning strategy to complete global model optimization while protecting data privacy; and enables concurrent access of devices to reduce end-to-end data processing latency.
[0029] 4) Dynamic scene control and scheduling: Based on fused data under a unified spatiotemporal benchmark, a digital twin engine is integrated to build a real-time scene mapping model; combined with a dynamic operations research optimization algorithm library, scene state inference and resource scheduling optimization are completed; scheduling instructions are generated and executed through edge nodes, forming a closed-loop control of "collection-processing-inference-scheduling-feedback", realizing precise scheduling and intelligent decision-making across the entire domain.
[0030] Step 1) specifically includes the following steps:
[0031] 1.1 Coordinate Acquisition from Multi-Source Devices For the i-th sensing device, its original coordinates are collected: Simultaneously obtain BeiDou positioning coordinates: in, For the first i The original coordinates of the sensing device. x i y i, z i : No. i The three-dimensional spatial coordinate components of a device in its own local coordinate system. For the first iGeographic coordinates of the device in the BeiDou / WGS coordinate system; It is the first i Longitude of each device It is the first i The dimensions of each device It is the first i The altitude of each device.
[0032] 1.2 WGS84 Unified Coordinate Transformation The device's local coordinates are converted to global coordinates in the WGS84 coordinate system using a homogeneous coordinate transformation matrix. in, For the first i WGS84 global coordinates after device transformation; homogeneous transformation matrix T Defined as: in, T The homogeneous coordinate transformation matrix contains rotation and translation information; These are elements of the rotation matrix R, used to achieve rotation and alignment of the coordinate system; These are translation vector components used to compensate for coordinate system offsets. The transformed device coordinate error meets the following constraints: in, This refers to the coordinate transformation error, which is the deviation between the transformed coordinates and the BeiDou positioning coordinates.
[0033] 1.3 Time Synchronization Calibration Calculate the offset of the device time from the standard time: in, For the first i Time offset of each device; For the first i The original data acquisition time of each device; This is the standard time for BeiDou navigation. The calibrated synchronization time is: in, For the first i The synchronized timestamps of each device after calibration. The system time synchronization accuracy meets the constraints: in, This represents the maximum time synchronization error between devices.
[0034] 1.4 Linear Interpolation Compensation For asynchronously sampled data, a linear interpolation algorithm is used for time alignment at any given time. t The calibrated data is as follows: in, The data is the device data after interpolation compensation at time t; , These are two adjacent valid sampling times; , For a moment , The original equipment data at the location.
[0035] 1.5 Spatiotemporal Relation Index Construction Define a uniform spatiotemporal index for each data entry: Where I is a unified spatiotemporal index, used to identify the unique location of data in the spatiotemporal dimension; The corresponding WGS84 3D coordinates for the data; This is the synchronization timestamp corresponding to the data.
[0036] Establish a mapping relationship between data streams and unified spatiotemporal coordinates: Wherein is the mapping function from the data stream to the spatiotemporal index; For the first i Data stream of each device ={ d 1 ,d 2 ,…, d n}, d j This is a single data entry.
[0037] Step 2) specifically includes the following steps:
[0038] 2.1 Video Stream Feature Extraction The input video frame is represented as: in, For input video frames; H, W, C These represent the height, width, and number of channels of a video frame, respectively.
[0039] Extracting video features using a CNN network: in, This represents the video modal feature vector. A convolutional neural network for video feature extraction. The output feature vector dimension satisfies:
[0040] 2.2 Point Cloud Feature Extraction The point cloud set is represented as: in, A collection of point cloud data; It is the first i The three-dimensional coordinates of each point.
[0041] Point cloud features are encoded using PointNet network features: in, The feature vectors of the point cloud modes; This is the PointNet network used for point cloud feature extraction.
[0042] 2.3 Internet of Things Data Encoding The IoT state vector is represented as: in, A vector of state data collected from IoT devices; For the first i Each state data component.
[0043] Mapped to feature vectors via fully connected layers: in, This represents the feature vector of the Internet of Things (IoT) mode. This is the weight matrix of the fully connected layer; For bias variables of fully connected layers; This is the activation function.
[0044] 2.4 Adaptive Weight Calculation Define the credibility of each modal feature: in, For the first i The credibility of each modal feature; For the first i Feature vectors of each modality; This is a credibility evaluation function used to dynamically calculate modal credibility based on data quality and interference levels.
[0045] Adaptive weights are obtained by normalization using the Softmax function: in, For the first i Adaptive weights for each modality; M This represents the total number of modes participating in the fusion. The weights satisfy the normalization constraint:
[0046] 2.5 Attention Fusion First, the features of each modality are weighted and fused: in, This is the initial weighted fusion feature vector.
[0047] Furthermore, an attention mechanism is introduced to enhance key features, and the attention coefficient is calculated as follows: in, For the first i Attention coefficients of each feature component; For the first i Attention query value for each feature component.
[0048] Final output: fused feature vector in, This is the final fused feature vector after being enhanced by the attention mechanism.
[0049] 2.6 Target Matching Cosine similarity is used to determine cross-modal target consistency: in, For feature vectors and Cosine similarity; These are target feature vectors of two different modalities; They are the feature vectors and The length of the module.
[0050] When the similarity is greater than a preset threshold, they are determined to be the same target: in, For feature vectors The corresponding target object.
[0051] Step 3) specifically includes the following steps:
[0052] 3.1 Edge Node Preprocessing Noise filtering is applied to the raw data: in, Clean data after noise filtering; Raw data collected for edge nodes; This represents the noise component in the data.
[0053] Normalize the data: in, The data is after normalization; This is the original data; These are the minimum and maximum values of the data in this dimension.
[0054] 3.2 Wavelet Compression For characteristic signals Wavelet transform is used to compress data. in, For signal wavelet transform coefficients; Scale factor; The translation factor; These are wavelet basis functions.
[0055] Define the data compression ratio: in, For data compression ratio; This represents the original data volume. This represents the compressed data volume; the compression ratio must meet the target constraints:
[0056] 3.3 Asynchronous Federated Learning Set edge nodes i The local model parameters are: The cloud-based FedAvg algorithm is used to aggregate and update the models of each edge node. in, These are the global model parameters aggregated in the cloud; For the first i The number of samples at each edge node; This represents the total number of samples across all edge nodes. For the first i Local model parameters for each edge node.
[0057] 3.4 Delay Optimization Model The total end-to-end latency of the system is the sum of the latency of edge processing, data transmission, and cloud processing. in, Total end-to-end delay; Address latency issues at edge nodes; For data transmission delay; To handle latency for cloud nodes.
[0058] The system must meet the low latency constraint:
[0059] Step 4) specifically includes the following steps: Assume the real-world scenario is as follows: The digital twin model is in the following state: Scene mapping error is defined as: in, The error in the state mapping between the real scene and the digital twin model; The actual scene state at time t; The state of the digital twin model at time t; Let be the norm of the state vector. The mapping error must satisfy the following constraints: 4.2 Risk Prediction Define the scenario risk assessment function: in, For a moment t Scenario risk value; No. j One risk factor; No. j The weighting coefficients of each risk factor.
[0060] 4.3 Dynamic Optimization Scheduling Construct a multi-objective scheduling optimization model with the objective function as follows: in, For scheduling comprehensive cost function; For scheduling time; For scheduling costs; For resource consumption. These are the weighting coefficients for each objective.
[0061] 4.4 Closed-loop feedback Define device feedback status: The scheduling instructions are dynamically updated based on equipment feedback to achieve closed-loop control. in, For a moment t The scheduling status; To based on equipment feedback The calculated scheduling state adjustment amount.
[0062] Through the above steps, this invention constructs a method for real-time processing and collaborative computing of multi-source heterogeneous data that integrates BeiDou spatiotemporal reference unification, cross-modal feature coding, edge-cloud collaborative computing, digital twins, and operations optimization scheduling. This method can achieve spatiotemporal synchronization of multi-source data, unified feature representation, low-latency processing, and precise control across the entire domain in complex dynamic scenarios.
[0063] Example 2: Another example of the present invention provides a real-time processing and collaborative computing system for multi-source heterogeneous data with unified spatiotemporal reference, comprising: The data input module is used to acquire data from multi-source sensing devices, spatiotemporal parameters, and device configuration information. The spatiotemporal reference calibration module is used to complete equipment coordinate calibration, timestamp calibration, and spatiotemporal correlation index establishment; from The multimodal feature encoding module is used for feature extraction, adaptive weight allocation, and cross-modal feature fusion in dynamic feature encoding networks. Edge-cloud collaborative computing module for data preprocessing, feature compression, global data fusion and model optimization; The digital twin and scheduling module is used for scene twin modeling, dynamic simulation, and generation and issuance of scheduling instructions; The constraint and security management module is used for data anomaly detection, security encryption, and access control. The results output module is used to output fusion features, target matching results, digital twin scenarios, and scheduling instructions.
[0064] Example 3: such as Figure 6 As shown, a multi-source heterogeneous data processing and collaborative computing application environment for complex dynamic scenarios is constructed, including multi-source sensing devices such as millimeter-wave radar, high-definition cameras, IoT sensors, and Beidou positioning terminals, as well as corresponding edge computing nodes and cloud computing nodes.
[0065] In this scenario, heterogeneous data from multiple sources, including radar point clouds, video streams, IoT sensing, and BeiDou positioning, are collected. Each device has independent spatial coordinates and system timestamps, leading to issues such as inconsistent spatiotemporal references, heterogeneous data features, and high processing latency. Traditional data processing methods are initially employed without spatiotemporal unification, cross-modal feature fusion, or edge-cloud collaborative computing to obtain baseline performance indicators, including spatiotemporal synchronization accuracy, cross-modal target matching effectiveness, data processing latency, and concurrency support capabilities.
[0066] Subsequently, the processing is performed according to the method proposed in this invention: First, spatial coordinate calibration and timestamp calibration are completed based on the BeiDou spatiotemporal reference to achieve millisecond-level spatiotemporal synchronization of multi-source data; then, multimodal data feature extraction, adaptive weight allocation, and cross-modal feature fusion are completed through a lightweight dynamic feature coding network to achieve unified representation and target matching; next, an edge-cloud collaborative architecture is adopted to complete data preprocessing, feature compression, and asynchronous federated learning to ensure low latency and high concurrency support; finally, based on the fused data, a digital twin engine and an operations research optimization algorithm library are integrated to complete scenario simulation and resource scheduling optimization.
[0067] After processing, the system outputs spatiotemporal synchronization data, cross-modal fusion features, target matching results, digital twin scenarios, and scheduling instructions, and calculates corresponding performance indicators, including spatiotemporal synchronization accuracy, cross-modal target matching accuracy, end-to-end latency, and device concurrent access capability.
[0068] By comparing the performance indicators before and after optimization, it can be verified that the method of the present invention can achieve unified spatiotemporal benchmarks for multi-source heterogeneous data, improve cross-modal data processing and target matching capabilities, reduce processing latency, enhance concurrency support capabilities, and realize precise control and efficient collaborative scheduling across the entire domain, under the premise of meeting the application constraints of complex dynamic scenarios.
[0069] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention do not depart from the essence and scope of the technical solution of the present invention.
Claims
1. A method for real-time processing and collaborative computing of multi-source heterogeneous data with unified spatiotemporal reference, characterized in that, Includes the following steps: 1) Construct a unified spatiotemporal coordinate system, perform spatial coordinate calibration of multi-source heterogeneous sensing devices based on the BeiDou spatiotemporal reference, and perform timestamp calibration on the multi-source data collected by each device to achieve millisecond-level spatiotemporal synchronization of multi-source data; 2) Multimodal data cross-modal feature encoding: Through a lightweight dynamic feature encoding network and an adaptive weight allocation mechanism, unified feature extraction and encoding are performed on radar point cloud data, video stream data, and IoT sensing data to achieve unified representation and target matching of cross-modal data; 3) Edge-cloud collaborative low-latency computing: deploy edge computing nodes and cloud nodes. Edge nodes complete local data preprocessing and feature compression, while cloud nodes realize global data fusion, dynamic inference and scheduling instruction generation. Hierarchical feature compression and asynchronous federated learning strategies are adopted to ensure concurrent access of devices and reduce end-to-end data processing latency. 4) Scene dynamic control and scheduling: Based on fused data under a unified spatiotemporal coordinate system, it integrates a digital twin engine and a dynamic operations optimization algorithm library to complete scene state deduction, resource scheduling optimization, generate scheduling instructions and issue them for execution through edge nodes, and achieve precise control across the entire domain.
2. The method according to claim 1, characterized in that, In step 1), the multi-source heterogeneous sensing devices include millimeter-wave radar, high-definition camera, IoT sensor, and Beidou positioning terminal, and the multi-source data includes radar point cloud data, video stream data, IoT status data, and positioning data.
3. The method according to claim 1, characterized in that: The process of constructing a unified spatiotemporal coordinate system in step 1) is as follows: the geographical coordinates of the device are obtained based on the BeiDou satellite positioning system, and a three-dimensional coordinate transformation algorithm is used to uniformly convert them into the WGS84 coordinate system; the timestamp is calibrated using a linear interpolation algorithm based on the BeiDou time signal. Establish a spatiotemporal correlation index to bind multi-source data to a unified spatiotemporal coordinate system.
4. The method according to claim 1, characterized in that: In step 2), the lightweight dynamic feature encoding network sequentially performs feature extraction, adaptive weight allocation, and cross-modal feature fusion; CNN, PointNet, and fully connected networks are used to extract features from video, point cloud, and IoT data, respectively. The credibility weights are calculated in real time based on the dynamic changes of the scene, and then weighted and fused after Softmax normalization to output a 128-dimensional unified feature vector.
5. The method according to claim 1, characterized in that: In step 2), cross-modal target matching uses a cosine similarity algorithm. When the similarity is greater than 0.85, the targets are determined to be the same.
6. The method according to claim 1, characterized in that: In step 3), the edge nodes are set with an appropriate coverage range. After noise filtering, normalization, and outlier removal, wavelet transform is used to compress the data volume. The cloud adopts a distributed architecture and an asynchronous federated learning strategy.
7. The method according to claim 1, characterized in that: In step 4), the digital twin engine enables real-time scene mapping with a mapping error of ≤1%; the dynamic operations research and optimization algorithm library supports path planning, resource allocation, and risk prediction; and the scheduling instructions are encrypted and transmitted to form a closed-loop management system of collection, processing, deduction, scheduling, and feedback.
8. The method according to claim 1, characterized in that: It also includes data security measures, employing a triple mechanism of data encryption, access control, and anomaly monitoring. Transmission uses AES-256 encryption, and access is managed using role-based access control.
9. The method according to claim 1, characterized in that: The lightweight dynamic feature encoding network, after pruning and quantization optimization, can be deployed lightweightly at edge nodes.
10. The method according to claim 1, characterized in that: The dynamic operations research and optimization algorithm library supports dynamic adaptation to scenarios and can automatically update the scheduling scheme according to changes in the number of devices and the target status.
11. A real-time processing and collaborative computing system for multi-source heterogeneous data with unified spatiotemporal reference, characterized in that, include: Data input module: used to acquire data from multi-source sensing devices, spatiotemporal parameters, and device configuration information; Spatiotemporal reference calibration module: used to complete equipment coordinate calibration, timestamp calibration and spatiotemporal correlation index establishment; Multimodal feature encoding module: used for dynamic feature encoding network feature extraction, adaptive weight allocation and cross-modal feature fusion; Edge-cloud collaborative computing module: used for data preprocessing, feature compression, global data fusion and model optimization; Digital twin and scheduling module: used for scene twin modeling, dynamic simulation and scheduling instruction generation and issuance; Constraint and security management module: used for data anomaly detection, security encryption and access control; Result output module: used to output fused features, target matching results, digital twin scenes and scheduling instructions.