Cross-scale spatiotemporal registration and alignment methods
By preprocessing and feature analyzing multimodal data, selecting appropriate spatiotemporal registration strategies, and combining learning algorithms to optimize parameters, the problem of multimodal data fusion requiring a unified coordinate system is solved, efficient multi-source data fusion is achieved, and the accuracy and system adaptability of multi-source data fusion are improved.
Patent Information
- Application Number
- CN202511000017.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-21
AI Technical Summary
In related technologies, multimodal data fusion requires the mandatory unification of coordinate systems and time bases, which solidifies the fusion method of multimodal data and results in poor multi-source data fusion effects.
By preprocessing multimodal data to generate feature vectors and knowledge elements, analyzing the time update characteristics and spatial position characteristics, determining the data type, and selecting the target strategy for spatiotemporal registration and alignment based on the data type, the alignment and registration parameters are optimized by combining evaluation indicators and learning algorithms, and the distributed feature extraction and adaptive fusion mechanism are used to dynamically adjust the alignment and registration strategies.
It improves the accuracy of multi-source data fusion, ensures high-precision spatiotemporal registration and alignment under different data characteristics and application scenarios, reduces errors in spatiotemporal data conversion, and improves the system's adaptability and alignment accuracy.
Smart Images

Figure CN120524097B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method for cross-scale spatiotemporal registration and alignment. Background Art
[0002] In the critical process of driving the development of smart cities in the AI era, the intelligent fusion of large-scale, multi-source data has become an indispensable core technical support and digital infrastructure cornerstone for achieving dynamic optimization of urban traffic, precise scheduling of energy grids, and instant emergency response. However, related technologies for multimodal data fusion require a unified coordinate system and time base, which rigidifies the fusion method and leads to poor multi-source data fusion results.
[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a method for cross-scale spatiotemporal registration and alignment, aiming to solve the technical problem in related technologies that multimodal data fusion requires the forced unification of coordinate systems and time bases, solidifies the fusion method of multimodal data, and thus leads to poor multi-source data fusion effects.
[0005] To achieve the above objectives, this application proposes a method for cross-scale spatiotemporal registration and alignment, which includes:
[0006] Preprocess the acquired multimodal data to generate feature vectors and knowledge elements;
[0007] Analyzing time update characteristics and spatial position characteristics based on the feature vector and the corresponding knowledge element to determine the data type of the multimodal data;
[0008] According to the data type, screening in the association relationship between the data type and the alignment and registration strategy to determine the target strategy corresponding to the multimodal data;
[0009] Based on the target strategy, the feature vectors and the corresponding knowledge elements are fused, and the fused data are spatially and temporally aligned to generate a target result;
[0010] Based on the evaluation indicators and the target results, combined with the learning algorithm, target parameters are generated to optimize the relevant parameters of alignment and registration.
[0011] In one embodiment, in the time sequence field obtained by parsing the feature vector, the time identification field and the space identification field corresponding to the feature vector are extracted according to the preset index bits;
[0012] Associating the time identification field and the space identification field according to the knowledge elements, and determining the spatiotemporal label corresponding to the feature vector based on the associated field relationship;
[0013] Based on the data type judgment rule, the spatiotemporal label matching is judged to determine the data type of the multimodal data.
[0014] In one embodiment, the feature vector is transformed and projected into a three-dimensional hierarchical coordinate system, and the knowledge element is used as a constraint function to fuse the projected feature vector and its constraint function to generate a fusion result;
[0015] Aligning the timestamps and scales of the fusion results according to the alignment method of the target strategy, and completing the missing timestamps according to the interpolation strategy of the target strategy to generate a time alignment result;
[0016] The initial alignment and fine alignment of the time alignment results are performed according to the geometric transformation and precision algorithm of the target strategy, and the target result is generated by aligning the time alignment results at various spatial scales.
[0017] In one embodiment, the timestamps of the data in the fusion result are calibrated to determine the missing parts of the timestamps corresponding to the fusion result;
[0018] According to the time interpolation algorithm, reasonable data is added to the time points corresponding to the missing parts of the timestamp to generate an initial result;
[0019] Alignment operations are performed on the initial results at different time scales to generate the time-aligned results.
[0020] In one embodiment, at least one evaluation indicator is obtained based on monitoring the alignment and registration effect of the target result;
[0021] Quantitatively analyzing the evaluation indicators, evaluating the alignment and registration effect of the target result, and generating an evaluation result;
[0022] A target optimization function is constructed for the evaluation results according to a learning algorithm, and the target parameters are generated according to a function solution of the target optimization function to optimize the relevant parameters of alignment and registration.
[0023] In one embodiment, based on the evaluation result, the dynamic level and spatial resolution level corresponding to the target result are analyzed to generate feedback information;
[0024] generating a feedback strategy for adjusting the time alignment frequency and spatial registration mode of the fusion result according to the feedback information;
[0025] The feedback strategy and the target strategy are fused and adjusted through a learning algorithm to generate target parameters.
[0026] In one embodiment, global feature extraction and distributed encoding are performed on the multimodal data to generate parent attributes, and local feature extraction and distributed encoding are performed on the multimodal data to generate child attributes;
[0027] Based on the parent attribute and the child attribute, time alignment and redundant information compression are performed to generate the feature vector.
[0028] In one embodiment, a global feature of the multimodal data is extracted, the global feature is encoded through distributed clustering, and a parent attribute corresponding to the global feature is generated;
[0029] Extract local features of the target area corresponding to the multimodal data, perform gradient calculation and encoding on the local features through computing nodes, and generate sub-attributes corresponding to the local features.
[0030] In one embodiment, after synchronizing the time axes of the parent attribute and the child attribute according to a preset protocol, the synchronized parent attribute and child attribute are concatenated to generate a high-dimensional vector;
[0031] The high-dimensional vector is subjected to principal component analysis, and is compressed and reduced in dimension according to a preset dimension to generate the feature vector.
[0032] In one embodiment, after performing target analysis on the multimodal data, the entity corresponding to the feature vector is identified;
[0033] Determine a relationship network between the entities based on the spatial topological association and the event logical association in the semantic association;
[0034] Rules and elements are extracted from the relationship network, and the rules and elements are associated and aggregated to generate the knowledge elements.
[0035] In addition, to achieve the above-mentioned purpose, the present application also proposes a device for spatiotemporal registration and alignment, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for cross-scale spatiotemporal registration and alignment as described above.
[0036] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the method for cross-scale spatiotemporal registration and alignment as described above are implemented.
[0037] The present application provides a method for cross-scale spatiotemporal registration and alignment. To address the accuracy issues of registration and alignment, in the multi-scale feature extraction stage, distributed feature extraction is used, covering global feature extraction and local feature extraction, to accurately extract basic attributes such as color, texture, and shape, generate low-dimensional representations, and provide a high-quality feature basis for subsequent alignment operations. In the spatiotemporal alignment process, steps such as timestamp alignment, time interpolation supplementation, multi-scale alignment, and fine alignment are introduced to effectively reduce the errors introduced during spatiotemporal data conversion and improve the accuracy of spatiotemporal data conversion. At the same time, the present application provides an adaptive fusion mechanism that dynamically adjusts alignment and registration strategies based on the characteristics of multi-source heterogeneous data, provides real-time feedback and optimizes alignment parameters, further improving the alignment accuracy and system adaptability, and ensuring that high-precision spatiotemporal registration and alignment can be achieved under different data characteristics and application scenarios.
[0038] To summarize, this application solves the problem in related technologies that multimodal data fusion requires the mandatory unification of coordinate systems and time bases by defining a new paradigm of joint temporal and spatial registration, solidifies the fusion method of multimodal data, and thus leads to poor results in multi-source data fusion, thereby improving the accuracy of multi-source data fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0041] Figure 1 This is a flowchart of the first embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application;
[0042] Figure 2 This is a flow chart of the second embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application;
[0043] Figure 3 This is a flowchart of the third embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application;
[0044] Figure 4 This is a flowchart of the fourth embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application;
[0045] Figure 5 This is a flowchart of the fifth embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application;
[0046] Figure 6 This is a schematic diagram of the architecture of the fifth embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application;
[0047] Figure 7 This is a flowchart of a sixth embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application;
[0048] Figure 8 This is a schematic diagram of the structure of the device for spatiotemporal registration and alignment in this application.
[0049] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0050] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0051] In related technologies, multimodal data fusion requires the unification of coordinate systems and time bases, which solidifies the fusion method of multimodal data and leads to poor results in multi-source data fusion.
[0052] The present application provides a solution: first, the acquired multimodal data is preprocessed to generate feature vectors and knowledge elements; then, based on the feature vectors and the corresponding knowledge elements, the time update characteristics and spatial position characteristics are analyzed to determine the data type of the multimodal data; then, based on the data type, the association between the data type and the alignment strategy is screened to determine the target strategy corresponding to the multimodal data; then, based on the target strategy, the feature vectors and the corresponding knowledge elements are fused, and the fused data are spatially and temporally aligned to generate the target result; finally, based on the evaluation index and the target result, the target parameters are generated in combination with the learning algorithm to optimize the relevant parameters of the alignment and registration.
[0053] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, a spatiotemporal registration and alignment device, etc. The following uses a spatiotemporal registration and alignment device as an example to illustrate this embodiment and the following embodiments.
[0054] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0055] The present invention provides a method for cross-scale spatiotemporal registration and alignment. Figure 1 , Figure 1This is a flowchart of the first embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application.
[0056] In this embodiment, the cross-scale spatiotemporal registration and alignment method includes steps S10 to S50:
[0057] Step S10: pre-process the acquired multimodal data to generate feature vectors and knowledge elements.
[0058] In this embodiment, a data source refers to raw data input from external sensors or databases. Multimodal data refers to a dataset containing at least two heterogeneous data types. A feature vector is a numerical sequence generated after dimensionality reduction, representing the core attributes of an entity. Knowledge elements refer to the associations between entities in a relational network and their related parameters.
[0059] As an optional implementation method of feature vectors, heterogeneous multimodal data is collected synchronously through various data sources. Global feature extraction is performed on multimodal data, including color, texture, and shape. At the same time, the global features are encoded through distributed clusters and aggregated into parent attribute encoding to represent entity categories. Local feature extraction is performed on the demarcated target area, including color, texture, and shape. Local key point detection and texture gradient calculation are performed in parallel at the edge computing node, and sub-attribute vectors are generated by combining sensor cascade coding under timestamps. The entity time base anchor point is defined based on the parent attribute, the sub-attribute time axis is aligned through the dynamic time warping algorithm, and principal component analysis is used to compress the redundant dimensions in the sub-attribute high-dimensional matrix whose variance contribution rate is lower than the preset contribution rate, and finally a feature vector is generated.
[0060] As an optional implementation method for knowledge elements, this approach identifies entity instances corresponding to multimodal data, completes the instantiation tagging of physical objects, and assigns unique entity identifiers and category labels. Spatial topological associations are calculated based on a predefined semantic association rule library. Through entity polygon coordinate inclusion verification and event logical association reasoning, a directed relationship network between entities is constructed, forming a relationship network of node entity identifiers and associated edge attributes. Finally, rules and elements are extracted from the relationship network. By identifying high-frequency association patterns and combining them with an expert experience library to generalize the rules, the extracted association rules and related elements are logically normalized to generate knowledge elements.
[0061] As an optional implementation method for multimodal data acquisition, heterogeneous multimodal data is collected synchronously through multiple data sources such as satellite sensors, traffic monitoring probes, and drone aerial photography, or multimodal data is obtained from the corresponding database of the smart city or public data source.
[0062] Step S20 : analyzing the time update characteristics and spatial position characteristics of the multimodal data based on the feature vector and the corresponding knowledge element, and determining the data type of the multimodal data.
[0063] In this embodiment, the data type refers to an identifier of the essential category of the data, and the type is determined based on the data update frequency and spatial position relationship.
[0064] As an optional implementation, the time and space identifier fields are extracted by parsing the binary coded fields in the feature vector. These fields are then combined with the temporal constraint rules and spatial structure rules in the precompiled knowledge element library to generate spatiotemporal labels. The spatiotemporal labels are then matched against the data type library to select the data type corresponding to the spatiotemporal labels.
[0065] As an optional implementation, a weighted calculation is performed on the time field, while simultaneously matching the spatial rule verification. The two are weighted and fused to output the target weight corresponding to the spatiotemporal tag. Based on the target weight and the weight judgment rule, the data type corresponding to the spatiotemporal tag is determined.
[0066] As an optional implementation, the time update characteristics and spatial position characteristics of the multimodal data are determined by analyzing the time update characteristics corresponding to the feature vector and combining the spatial position characteristics analysis of the corresponding knowledge elements. If the time update characteristics corresponding to the multimodal data are high frequency and the spatial position characteristics are dynamic, then the data type corresponding to the multimodal data is determined to be high frequency data. If the time update characteristics corresponding to the multimodal data are low frequency and the spatial position characteristics are static, then the data type corresponding to the multimodal data is determined to be low frequency data. If the characteristic analysis corresponding to the multimodal data is high resolution, then the data type corresponding to the multimodal data is determined to be high resolution data. If the characteristic analysis corresponding to the multimodal data is low resolution, then the data type corresponding to the multimodal data is determined to be low resolution data.
[0067] Step S30 , based on the data type, screening is performed in the association relationship between the data type and the alignment and registration strategy to determine a target strategy corresponding to the multimodal data.
[0068] In this embodiment, the alignment strategy refers to a set of technical methods for achieving spatiotemporal synchronization. The association relationship refers to the preset mapping rules between data types and strategies. The target strategy refers to the optimal alignment method and parameter set for the current data.
[0069] As an optional implementation, based on the data type, a table of associations between data types and alignment strategies is filtered to select an initial strategy corresponding to the data type. Alignment testing of the multimodal data is performed based on the initial strategy. Based on the test feedback, a feedback compensation strategy is determined in the feedback rule library and added to the initial strategy to generate the target strategy.
[0070] As an optional implementation, if the data type is high-frequency data, a real-time alignment and registration strategy is adopted to ensure the timeliness and accuracy of the data. If the data type is low-frequency data, a periodic alignment and registration strategy is adopted to reduce computational costs and ensure data stability. If the characteristic analysis corresponding to the multimodal data is high-resolution, an alignment and registration algorithm based on feature point matching is used to achieve high-precision spatial alignment. For low-resolution multimodal data, a statistical-based method is used for alignment and registration to improve processing efficiency.
[0071] Step S40: Based on the target strategy, the feature vectors and the corresponding knowledge elements are fused, and the fused data are aligned in time and space to generate a target result.
[0072] In this embodiment, fused data refers to the structured data volume generated by combining features and rules. Spatiotemporal registration refers to the unification of spatiotemporal references through mathematical transformations. The target result refers to the spatiotemporal synchronized fusion output.
[0073] As an optional implementation, the feature vectors and their corresponding knowledge elements are associated and fused to generate a fusion result. The fusion result is first time-aligned to synchronize the results in the temporal dimension. Then, the fusion result is spatially registered to align the results in spatial position, generating the target result.
[0074] As an alternative implementation, a 3D hierarchical coordinate system is constructed, and the feature vectors are transformed into a fusion result. Time alignment is activated based on the target strategy, and temporal interpolation compensation is performed to generate a time-aligned result. Based on the time-aligned result, spatial registration is performed to global coordinates, and the target result is output.
[0075] Step S50: generating target parameters based on the evaluation index and the target result in combination with a learning algorithm to optimize parameters related to alignment and registration.
[0076] In this embodiment, the evaluation metric refers to a numerical parameter that quantifies the registration effect. The learning algorithm refers to a model that automatically optimizes parameters through data training. The target parameter refers to the key variable that needs to be adjusted. Optimization refers to the process of improving parameters to enhance registration accuracy.
[0077] As an optional implementation method, the evaluation indicators and the trajectory point sequence in the target result are input, the learning algorithm is used to evaluate and analyze the target result, the objective function and parameter space are constructed, the optimal solution is selected through a preset number of Gaussian process regression iterations, and the target parameters corresponding to the optimal solution are output.
[0078] As another optional implementation, the motion state vector in the combined evaluation index and target result is input into a genetic algorithm, including chromosome encoding and fitness function, to perform selection, crossover, mutation and evolution for a preset number of generations, and the optimal individual is decoded to generate target parameters and output.
[0079] For example, in a smart city traffic dispatch scenario, road cameras and radars collect visible light images and millimeter wave point clouds of vehicles. The RGB color moment distribution is calculated for the entire image, and SIFT keypoints and velocity vectors are extracted within the vehicle detection frame. PCA dimensionality reduction is then performed to generate a 16-dimensional feature vector. Entities E01 and E02 are identified based on the feature vectors, and a relationship network is constructed based on spatial topology rules to generate the knowledge element "When the emergency vehicle is less than 200 meters from the intersection, extend the green light by 15 seconds." The time stamp field and spatial flag in the feature vector are analyzed, combined with the knowledge element, to determine if it is a dynamic data type. The target policy ID_D004 is bound based on the "dynamic to real-time filtering" association. The feature vector and knowledge element are fused to generate a fusion result [vehicle ID, coordinates, speed, green light extension time stamp]. Time alignment is performed, UTC timing is interpolated to 10 Hz, and spatial registration is performed. The trajectory is then adsorbed based on the road network, generating the target result: an instruction to extend the green light in the N direction of the intersection to t+15 seconds, improving emergency vehicle traffic efficiency during peak hours by 33%.
[0080] For example, during multi-source data fusion, to ensure data consistency and accuracy across both spatial and temporal dimensions, we designed a comprehensive alignment process encompassing two core components: spatial alignment and temporal alignment. Static data is spatially consistent but has low temporal correlation. Therefore, temporal alignment primarily targets dynamic data, focusing on ensuring its temporal integrity, continuity, and correlation, thereby guaranteeing data synchronization. Dynamic data is subject to change in spatial terms but highly correlated with time. Based on this, in this framework, we first perform temporal alignment followed by spatial alignment to reduce computational resource consumption. Temporal alignment aims to calibrate data synchronization across the temporal dimension, ensuring temporal coherence between static and dynamic data. First, time alignment involves timestamp alignment: First, the data timestamps are aligned to ensure accurate correspondence across all data. This process takes into account factors such as time synchronization errors in the data acquisition equipment and time delays during data transmission, laying the foundation for subsequent alignment operations. Temporal interpolation supplements: In some cases, data may contain temporal gaps or discontinuities. Temporal interpolation uses interpolation algorithms to generate reasonable data at missing time points, making the data more complete and continuous in time. This stage primarily targets dynamic data, supplementing its temporal gaps and ensuring data coherence, thus providing a continuous time series for subsequent analysis. Multiscale alignment: Similar to spatial alignment, temporal alignment also involves multiscale alignment. By performing alignment operations at different time scales, data consistency is ensured across different time scales, improving the accuracy and stability of the temporal alignment. While static data primarily focuses on the consistency of the temporal base, multiscale alignment of dynamic data requires performing alignment at different time scales (such as hours, days, and months) to ensure temporal consistency across these scales to meet the needs of diverse application scenarios. Spatial alignment, on the other hand, involves calibrating the spatial position of data. Its goal is to ensure accurate spatial correspondence between data collected from different sensors or at different time points. Initial alignment: This is the first step in spatial alignment, primarily through rough geometric transformations to initially align the data to a rough spatial position. This stage typically uses simple operations such as translation and rotation to quickly minimize spatial differences between the data. For static data, since it is relatively stable in space, it can be based on a fixed reference frame. For example, using a geographic coordinate system as a reference, the data can be aligned to the same coordinate system through simple translation and rotation, providing a basis for subsequent fine alignment. Fine alignment: Based on the initial alignment, further improve the alignment accuracy. Through more complex algorithms, such as those based on feature point matching or bundle adjustment, the data is fine-tuned to achieve higher spatial matching accuracy. Static data is based on feature points for high-precision spatial matching, while the fine alignment of dynamic data needs to consider changes in time and space at the same time to ensure accurate correspondence of the data in the temporal and spatial dimensions.Multi-scale alignment: Considering that data may have different characteristics and errors at different scales, multi-scale alignment performs alignment operations at multiple scales to ensure good data consistency across all scales. This process helps improve the robustness and accuracy of the alignment results. Multi-scale alignment of static data focuses on maintaining data consistency at different spatial resolutions, while multi-scale alignment of dynamic data requires multi-scale alignment at different temporal and spatial scales to ensure spatial consistency across different time scales to accommodate the spatiotemporal variations of dynamic data. Third, the joint alignment and optimization module: After completing spatial and temporal alignment, the spatially and temporally aligned data is input into the joint alignment and optimization module. This module's primary task is to comprehensively optimize the spatial and temporal alignment results to further improve the overall alignment accuracy. This optimized alignment provides a more accurate and reliable data foundation for subsequent data analysis and applications, thereby improving the performance and effectiveness of the entire data processing pipeline. The entire alignment process is designed to achieve high-precision alignment of data in both spatial and temporal dimensions, providing strong support for multi-source data fusion and, in turn, ensuring high-quality data for research and applications in related fields.
[0081] By defining a new paradigm of joint temporal and spatial registration, the problem of multimodal data fusion in related technologies requiring the unification of coordinate systems and time bases is solved, the fusion method of multimodal data is solidified, which in turn leads to poor multi-source data fusion effects, and the accuracy of multi-source data fusion is improved.
[0082] Based on any of the above embodiments, in the second embodiment of the present application, refer to Figure 2 , Figure 2 This is a flow chart of the second embodiment of the method for cross-scale spatiotemporal registration and alignment of this application. Step S20 includes steps A11 to A13:
[0083] Step A11 : extracting the time identification field and the space identification field corresponding to the feature vector according to preset index bits from the time series field obtained by parsing the feature vector.
[0084] In this embodiment, time series parsing refers to analyzing data patterns related to time evolution in feature vectors. A time identification field is a binary segment in a feature vector that encodes temporal attributes. A spatial identification field is a numerical segment in a feature vector that identifies spatial characteristics. A time series field is a continuous data field in a feature vector that stores time-related parameters.
[0085] As an optional implementation, based on the timing field corresponding to the feature vector, the index preset bit binary segment is extracted as the time identification field, converted into the time update frequency through the frequency decoder, the index preset bit value is synchronously extracted, parsed into the space identification field, and the time identification field and the space identification field are output.
[0086] As another optional implementation, in the feature vector, a timestamp interval sequence field consisting of a preset index bit is located, and the mean time interval is calculated as the time identification field. The preset index bit is a space flag bit, and the flag bit field is directly extracted as the space identification field.
[0087] Step A12: Associate the time identification field and the space identification field according to the knowledge elements, and determine the space-time label corresponding to the feature vector based on the associated field relationship.
[0088] In this embodiment, the spatiotemporal label refers to a classification identifier that characterizes the spatiotemporal attributes of data.
[0089] As an optional implementation, extract the time identification field from the first preset position to the second preset position of the feature vector index, and the space identification field from the third preset position to the fourth preset position. Based on the time identification field and the space identification field, match the corresponding knowledge elements and generate a spatiotemporal label.
[0090] Step A13: Based on the data type judgment rule, the spatiotemporal label matching is judged to determine the data type of the multimodal data.
[0091] In this embodiment, the data type judgment rule refers to a set of preset logical conditions. Matching judgment refers to the process of comparing the label with the rule conditions.
[0092] As an optional implementation, the spatiotemporal label is matched and judged according to the preset data type judgment rule, and the spatiotemporal label is used as an index to search in the data type judgment rule to determine the data type corresponding to the multimodal data.
[0093] For example, in the context of smart city traffic scheduling, vehicle motion feature vectors are obtained from millimeter-wave radar streams at intersections. The time field ("01011") in indexes 3 to 7 is decoded as a frequency of 22 Hz, and the spatial field ("110") in indexes 8 to 12 is labeled as a dynamic trajectory. Combined with the knowledge element library rule "Frequency > 20 Hz and dynamic is marked as a dynamic label," the spatiotemporal label "Dynamic" is generated. Based on the data type judgment rule ("Dynamic label + trajectory continuity confidence > 0.8 is considered dynamic data"), the data is determined to be dynamic.
[0094] For example, before performing alignment and registration, the multi-source heterogeneous data is first classified and analyzed for characteristics. Based on the results of feature extraction and element extraction, the data is divided into two categories: dynamic data and static data. Its temporal and frequency properties are then analyzed in depth, and its characteristics in both spatial and temporal dimensions are evaluated. This process aims to fully understand the data characteristics and provide a basis for subsequent adaptive alignment and registration strategies. The frequency and method of alignment and registration are flexibly adjusted based on the dynamic nature of the data. For high-frequency dynamic data, such as traffic flow, a real-time alignment and registration strategy is adopted to ensure data timeliness and accuracy. For low-frequency static data, such as geographic basemap data, a periodic alignment and registration strategy is used to reduce computational costs and ensure data stability. Furthermore, the appropriate alignment and registration algorithm is selected based on the spatial resolution of the data. For high-resolution image data, an alignment and registration algorithm based on feature point matching is used to achieve high-precision spatial alignment. For low-resolution data, a statistical-based alignment and registration method is used to improve processing efficiency.
[0095] By classifying multimodal data, dynamic data and static data are processed with different strategies respectively, and corresponding methods are used for alignment and registration to improve processing efficiency and reliability of data processing.
[0096] Based on any of the above embodiments, in the third embodiment of the present application, refer to Figure 3 , Figure 3 This is a flow chart of the third embodiment of the method for cross-scale spatiotemporal registration and alignment of this application. Step S40 includes steps B11 to B13:
[0097] Step B11: transform and project the feature vector into a three-dimensional hierarchical coordinate system, and use the knowledge element as a constraint function to fuse the projected feature vector and its constraint function to generate a fusion result.
[0098] In this embodiment, the three-dimensional hierarchical coordinate system refers to a three-level coordinate reference frame consisting of a geocentric system, a projection system, and an object system. Unifying the levels means converting the data to the same coordinate system level.
[0099] As an optional implementation, the feature vector is projected from the geocentric coordinate system to a three-dimensional hierarchical coordinate system using parameter transformation, and then transformed to the object coordinate system with the entity's center of mass as the origin. The knowledge elements are compiled into object constraint functions. The projected feature vector, the entity in the object coordinate system, and the object constraint function are fused to generate a fusion result.
[0100] As another optional implementation, the feature vectors are mapped from a three-dimensional hierarchical coordinate system to a projection system grid via a coordinate transformation chain. A local system is established with the object system's motion direction as the X-axis. The knowledge elements are converted into object system motion constraints. The projected feature vectors, object system motion constraints, and local system are fused to produce a fusion result.
[0101] Step B12: aligning the timestamp and scale of the fusion result according to the alignment method of the target strategy, and completing the missing timestamp part according to the interpolation strategy of the target strategy to generate a time alignment result.
[0102] In this embodiment, timestamp alignment refers to unifying time tags to a standard clock reference. Time interpolation refers to filling in missing time periods with data. Multi-scale alignment refers to checking consistency at different resolution levels. The time alignment result is a data sequence with synchronized output time axes.
[0103] As an optional implementation, the target strategy processes the fusion results, aligns the timestamps with preset second pulses, generates trajectory points by linear interpolation for the missing intervals, verifies the displacement continuity at the meter scale, checks the velocity mutation at the centimeter scale, and outputs the time alignment results.
[0104] As an alternative implementation, a targeted strategy processes the hybrid fusion results, aligning timestamps for dynamic data using Kalman filtering and synchronizing static data using atomic clocks. Linear interpolation is performed for dynamic segments, while the original time scale is retained for static segments. The spatial inclusion relationship between trajectories and the road network is verified at the meter level, while the displacement variance of static points is verified at the centimeter level. The time-aligned results are then output.
[0105] Step B13: performing initial alignment and fine alignment of the time alignment results according to the geometric transformation and precision algorithm of the target strategy, and generating the target result by aligning the time alignment results at various spatial scales.
[0106] In this embodiment, initial alignment refers to rough registration achieved through basic spatial transformations. Fine alignment refers to high-precision feature matching to optimize local positions. The target result is a high-precision fusion output that is unified in time and space.
[0107] As an optional implementation method, the target strategy loads the time alignment results, performs projection transformation for initial alignment, uses point cloud registration for fine alignment, and verifies the road network connectivity at the meter level and the building corner spacing error at the centimeter level for multi-scale alignment. After verification, the target result is generated.
[0108] As an alternative implementation, the target strategy decomposes the time alignment results. Initial alignment of the static portion uses an affine transformation, while initial alignment of the dynamic portion uses dynamic time warping. Static fine alignment utilizes feature matching, while dynamic fine alignment uses particle filtering. Multi-scale alignment verifies the spatial topology of static and dynamic entities at the meter level and the vibration spectrum energy distribution at the millimeter level to generate the target results.
[0109] For example, in a smart city traffic dispatch scenario, vehicle trajectory point clouds and video streams are captured by millimeter-wave radar and cameras at intersections. These multimodal data are then processed in a three-dimensional hierarchical coordinate system. The feature vector [speed = 62 km / h, acceleration = 1.3 m / s²] is fused with the knowledge element "emergency priority" in the geocentric coordinate system to produce {coordinates (longitude, latitude, altitude), velocity vector, priority label}. This is then projected to the UTM 50N coordinate system using a seven-parameter transformation. An object coordinate system is then established with the traffic light as the origin, generating the fused result {local coordinates (x', y', z'), motion parameters, priority}. Using the dynamic target strategy ID_D007 (Kalman filter Q = 0.015), the fused result is timestamp aligned using GPS timing ±1ms and linear interpolation to fill in missing frames (10 Hz → 30 hr). The trajectory-road network topology connectivity is verified at the meter level, and acceleration continuity (|Δa / Δt| ≤ 3 m / s³) is verified at the centimeter level to produce the time-aligned result. Based on the strategic spatial registration instructions, the initial alignment is registered to the global coordinates of the road network through affine transformation. The fine alignment uses ICP iterative optimization to match wheel feature points. Multi-scale alignment verifies the spatial relationship between the vehicle and the traffic light at the meter level and verifies the trajectory curvature radius at the centimeter level. The target results are generated: the green light at the driving intersection is extended by 15 seconds, and the emergency vehicle travel delay is reduced by 37%.
[0110] Data fusion can combine the advantages of different data sources and improve the accuracy and reliability of fusion result analysis.
[0111] Based on any of the above embodiments, in the fourth embodiment of the present application, refer to Figure 4 , Figure 4 This is a flow chart of the fourth embodiment of the method for cross-scale spatiotemporal registration and alignment of this application. Step B12 includes steps C11 to C13:
[0112] Step C11: calibrate the timestamps of each data in the fusion result, and determine the missing parts of the timestamps corresponding to the fusion result.
[0113] In this embodiment, timestamp calibration refers to aligning the time stamp to a standard clock reference. Timestamp missing portions refer to uncollected or lost time slots in the data stream. Determining missing portions refers to the process of identifying and locating discontinuities in the time axis.
[0114] As an optional implementation, the fusion results are input and the timestamps are calibrated to nanoseconds using a precision clock protocol. Missing segments with time intervals greater than a preset interval are detected, and a third-order spline interpolation algorithm is used to generate a continuous temperature field and coordinate sequence, outputting the initial results.
[0115] As another optional implementation, the timestamps of the point cloud frame sequence in the fusion result are aligned by presetting the second pulse. For the missing period, the transition point cloud is generated based on the motion model of the adjacent frames and the initial result is output.
[0116] Step C12: supplementing the time points corresponding to the missing parts of the timestamp with reasonable data according to a time interpolation algorithm to generate an initial result.
[0117] In this embodiment, time scale refers to the temporal granularity of data observation. Alignment refers to the process of unifying data at different time scales to a common time base. Time interpolation refers to the computational method used to generate reasonable data for missing time periods. Time points refer to uncollected time slots in a data stream. Reasonable data refers to interpolation results that conform to physical laws and statistical characteristics. Initial results refer to data sequences with a complete time axis.
[0118] As an optional implementation, the initial results of the input sampling trajectory sequence are aligned at the minute time scale. Taking the whole minute as the reference point, the second-level data is linearly interpolated to generate minute alignment points, and the time alignment result is output.
[0119] As another optional implementation, the initial result of asynchronous multi-source data is input, and the first data source is segmented and downsampled based on the timestamp of the second data source, and the mean difference between adjacent segments is constrained to be less than the preset mean difference, and the time alignment result is output.
[0120] Step C13: performing alignment operations on the initial results at different time scales to generate the time alignment results.
[0121] In this embodiment, different time scales refer to the granularity of data processing. Alignment refers to the process of unifying data to a reference time point. The time alignment result is a synchronized data sequence at multiple scales.
[0122] As an optional implementation, the trajectory sequence corresponding to the initial result is input and aligned at the minute time scale. Taking zero second as the reference point, the second-level data is linearly interpolated to generate coordinate points. When the distance between adjacent reference points is a preset multiple of the sampling interval, cubic spline interpolation is enabled to output the time alignment result of multi-scale alignment.
[0123] As another optional implementation, based on the initial results corresponding to the asynchronous data source, in multi-scale alignment, the timestamp of the first source data is used as the reference point, and the average of the second source data segments is taken. If the mean change rate within the segment is greater than the preset change rate, it is split into sub-segments and the time alignment result of the synchronous sequence is output.
[0124] For example, in a smart city traffic dispatch scenario, a millimeter-wave radar (data source A) and a camera (data source B) at an intersection capture an emergency vehicle's trajectory point cloud and video stream, generating a fusion result: radar coordinates (x1, y1, t1) and image coordinates (x2, y2, t2). Timestamps are calibrated to nanosecond UTC using the PTP precision clock protocol. Trajectory points are supplemented for missing radar frames using cubic spline interpolation, constraining acceleration to ≤ 3 m / s². This generates an initial result: a continuous trajectory sequence. At the second-scale, the 30Hz video coordinates are downsampled to 1Hz to match the traffic light phase. At the millisecond-scale, the radar trajectory is interpolated to 1000Hz to synchronize with the video optical flow, generating a time-aligned result.
[0125] Since data are integrated in the spatial and temporal dimensions and the scale differences between different fields are coordinated, the synchronization of data is guaranteed and the accuracy of multi-source data fusion is improved.
[0126] Based on any of the above embodiments, in the fifth embodiment of the present application, refer to Figure 5 , Figure 5 This is a flowchart of the fifth embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application. Step S50 includes steps D11 to D13:
[0127] Step D11: obtaining at least one evaluation indicator based on monitoring the alignment and registration effects of the target result.
[0128] In this embodiment, the registration effect refers to the spatial position consistency. The evaluation index refers to the numerical parameter that quantifies the performance.
[0129] As an optional implementation, the acceleration mutation values of adjacent trajectory points in the target result are monitored and the time alignment effect index is calculated.
[0130] As another optional implementation, the coordinates of the static building corner points in the target result are extracted, and the root mean square error of the cloud computing space distance with the reference point is used to generate a registration effect index.
[0131] As another optional implementation, the target result output delay is monitored, and when the delay is greater than a preset delay, the timeliness indicator is recorded.
[0132] Therefore, the evaluation index may be one of a time alignment effect index, a registration effect index, and a recording timeliness index.
[0133] Step D12: quantitatively analyzing the evaluation indicators, evaluating the alignment and registration effect of the target result, and generating an evaluation result.
[0134] In this embodiment, the evaluation result refers to the conclusive output of whether the performance meets the standard.
[0135] As an optional implementation, a set of evaluation indicators is input for quantitative analysis. The alignment is evaluated as either unsatisfactory for temporal synchronization or marginal for spatial registration, with the corresponding evaluation results generated. The temporal jitter exceeds a preset threshold, and the spatial jitter exceeds a preset percentage.
[0136] As another optional implementation, an evaluation index is obtained and quantitative analysis is performed. When the evaluation acceleration index exceeds a limit, an evaluation result is output: the time alignment effect fails, and the spatial registration effect meets the standard.
[0137] Step D13: constructing a target optimization function for the evaluation result according to a learning algorithm, and generating the target parameters according to a function solution of the target optimization function to optimize the relevant parameters of alignment and registration.
[0138] In this embodiment, a learning algorithm refers to a model that automatically optimizes parameters through data training. Target parameters refer to key variables that need to be adjusted. Optimization refers to the process of improving performance by iteratively refining parameters.
[0139] As an optional implementation, based on the evaluation results, particle swarm optimization is driven to generate global optimal particles, and target parameters are generated based on the global optimal particles.
[0140] As another optional implementation, multiple evaluation results are integrated to construct a multi-objective optimization, and the Pareto front is solved by an algorithm to weigh the solution to the corresponding target parameters.
[0141] For example, in a smart city traffic dispatch scenario, real-time monitoring of an emergency vehicle target using radar at an intersection revealed that its three-dimensional trajectory was aligned with the phase of the traffic light, resulting in evaluation metrics such as spatial conflict rate of 1.8% and temporal jitter σ of 0.9 ns. Quantitative analysis showed that the spatial conflict rate exceeded the threshold of 20%, while the temporal jitter was acceptable. The evaluation results were: "Spatial Registration," "Deviation Exceeded," "Temporal Alignment," and "Excellent," resulting in an overall rating of "B." Parameters optimized using reinforcement learning were: state = spatial conflict rate 1.8%, action space = [Kalman filter Q ± 0.005, interpolation step size ± 0.1 s], and reward function R = 1 / conflict rate. After executing the action "Increase Q by 0.005," the conflict rate dropped to 0.7%, resulting in the target parameter: noise covariance Q = 0.015.
[0142] For example, alignment and registration effect monitoring: real-time monitoring of the quality of alignment and registration results, focusing on the consistency, accuracy and completeness of the data. By introducing evaluation indicators such as error rate, alignment accuracy, response time, etc., the alignment and registration effects are quantitatively analyzed, so as to comprehensively and objectively evaluate the performance of alignment and registration. Dynamic parameter adjustment: Dynamically adjust the alignment and registration parameters based on the monitoring results of the alignment and registration effects. For example, if monitoring finds that the alignment accuracy does not meet expectations, the parameters of the feature extraction method or alignment algorithm can be adjusted in time to optimize the alignment and registration effects. In addition, with the help of machine learning algorithms, such as reinforcement learning, the alignment and registration parameters are further optimized, so that the system can automatically learn and adapt to different data characteristics and application scenarios, thereby achieving continuous optimization of the alignment and registration strategy.
[0143] Further, refer to Figure 6 , Figure 6 This is a schematic diagram of the architecture of the fifth embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application. Visible light cameras, lidars, and on-board GPS deployed on the roadside: multi-source heterogeneous data input, synchronously collect vehicle RGB images, point cloud coordinates, and positioning trajectory streams. In the data classification module, data channels are divided according to sensor type and spatiotemporal density, and feature analysis is performed to extract frequency distribution and spatial accuracy parameters. After entering the extraction layer, distributed feature extraction is started: the global color histogram of the image is calculated in parallel on the GPU cluster: HSV three channels, local texture LBP operator: 128-dimensional descriptor of the wheel area, point cloud shape features: minimum circumscribed rectangle main axis. Synchronous feature extraction: multi-scale feature extraction, identification of macro road network topology and micro vehicle corner entities, and the construction of a "vehicle and road" attribution relationship chain through semantic association and automatic registration: spatial inclusion verification and trajectory direction logical matching. The spatiotemporal hybrid feature alignment and registration framework, which is pushed to the alignment and registration layer, is processed in a separate path. The temporal alignment module compensates for sensor time differences using UTC as the reference: linear interpolation of camera and radar timestamps. The spatial alignment module performs a WGS84 to UTM projection transformation and iterative ICP registration of feature points. Finally, the joint alignment and optimization module integrates spatiotemporal parameters using the loss function L = αL_time + βL_space. Dynamic weighting (α = 0.7, β = 0.3) is used to output a centimeter-level accurate fusion model of vehicle trajectories and road networks.
[0144] By introducing evaluation indicators such as error rate, alignment accuracy, response time, etc., the alignment and registration effects are quantitatively analyzed, thereby comprehensively and objectively evaluating the performance of alignment and registration, and improving the accuracy of multi-source data fusion.
[0145] Based on any of the above embodiments, in the sixth embodiment of the present application, refer to Figure 7 , Figure 7This is a flow chart of the sixth embodiment of the method for cross-scale spatiotemporal registration and alignment of the present application. Step D13 includes steps E11 to E13:
[0146] Step E11: Based on the evaluation result, analyze the dynamic level and spatial resolution level corresponding to the target result to generate feedback information.
[0147] In this embodiment, the dynamic level refers to the grading of the temporal frequency of data updates. The spatial resolution level refers to the grading of the spatial accuracy of the data. Feedback information refers to the evaluation of the target result and the feedback of information that needs to be adjusted and optimized.
[0148] As an optional implementation, the standard deviation of the time interval sequence of trajectory points in the target result is calculated, and according to the dynamic classification rule, the minimum value of the grid vertex spacing is scanned at the same time to match the spatial resolution level and generate feedback information.
[0149] As another optional implementation, the proportion of moving entities and the average speed in the target result are detected to determine the level of dynamics, the entity boundary positioning error corresponds to the spatial resolution level, and feedback information is generated accordingly.
[0150] Step E12: generating a feedback strategy for adjusting the time alignment frequency and spatial registration mode of the fusion result according to the feedback information.
[0151] As an optional implementation, a linear adjustment strategy is triggered according to high-dynamic and centimeter-level feedback information, the time alignment frequency is increased according to the linear adjustment strategy, and the spatial registration method is switched from affine transformation to ICP registration to generate a feedback strategy.
[0152] As another optional implementation, a hybrid strategy is called based on the feedback information at the medium dynamic and centimeter levels. The time alignment frequency is maintained according to the hybrid strategy, the interpolation algorithm is changed from linear to cubic spline, the spatial registration method retains ICP, the number of iterations is increased, and the feedback strategy is output.
[0153] Step E13: The feedback strategy and the target strategy are integrated and adjusted through a learning algorithm to generate target parameters.
[0154] In this embodiment, fusion adjustment refers to the process of integrating feedback instructions with existing strategy parameters.
[0155] As an optional implementation, the feedback strategy is combined with the target strategy ICP for a preset number of iterations, and the target parameters are output by constructing the target function and parameter space and performing Gaussian process regression convergence.
[0156] For example, in a smart city traffic dispatch scenario, millimeter-wave radar and cameras at intersections collect real-time ambulance trajectory data. A spatial registration error of 1.2 meters in the target result triggers the evaluation result "spatial offset error." Analysis of the target result reveals a standard deviation of 0.05 seconds between trajectory points, determining the dynamics level as high. A point cloud matching residual of 0.08 meters corresponds to a centimeter-level spatial resolution, generating feedback information: "High dynamics + centimeter-level + spatial error of 1.2 meters." Based on this feedback information, a feedback strategy is generated: "Increase time alignment frequency from 10Hz to 15Hz + switch spatial registration to ICP." This feedback strategy is then fused with the current target strategy, a Kalman filter algorithm. The state is: error 1.2m, frequency 10Hz, registration method: SIFT, action space: Q±0.005, frequency ±5Hz, switch to ICP, and reward function R=1 / spatial error. Execute the action: Q + 0.005, frequency → 15 Hz. After switching to ICP, the measured error is reduced to 0.3 meters. The target parameters are generated: Q_new = 0.015, frequency_new = 15 Hz, and registration method = ICP.
[0157] Since the results are comprehensively optimized through the joint alignment and optimization module, accurate registration and efficient alignment of data in the two key dimensions of time and space are achieved, providing a more accurate and reliable data foundation for subsequent data analysis and application, and significantly improving the effect of multi-source data fusion.
[0158] Based on any of the above embodiments, in the seventh embodiment of the present application, step S10 includes steps F11 to F12:
[0159] Step F11 : performing global feature extraction and distributed encoding on the multimodal data to generate parent attributes, and performing local feature extraction and distributed encoding on the multimodal data to generate child attributes.
[0160] In this embodiment, the data source refers to the raw data input provided by external sensors or databases. Distributed encoding refers to compressing feature dimensions in a multi-node parallel computing architecture. Parent attributes refer to abstract semantic encodings at the entity category level. Child attributes refer to quantitative parameters at the entity instance level. Global feature extraction refers to extracting detailed features from the target region. Local feature extraction refers to extracting detailed features within the target subregion. Dimensionality reduction refers to compressing high-dimensional features using algorithms such as principal component analysis, preserving information while reducing computational effort.
[0161] As an optional implementation method, a road monitoring camera is called to collect multimodal data of a video stream and a millimeter-wave radar point cloud sequence.
[0162] As another optional implementation, drones equipped with multispectral cameras are deployed to acquire multimodal data including visible light images, near-infrared vegetation index, and thermal infrared temperature matrix.
[0163] As another optional implementation, a public data source is connected through an interface to obtain multimodal data in the public data source.
[0164] As an optional implementation, cluster distributed computing is performed on multimodal data. The global color histogram of the entire image and the overall normal distribution of the point cloud are extracted, compressed, and encoded into parent attribute labels. Keypoint detection is performed in parallel in the building outline area, and the gray-level co-occurrence matrix energy value is calculated in the road crack area. The linear discriminant analysis is then used to reduce the dimension and generate the child attribute vector.
[0165] As another optional implementation, cluster distributed computing is performed on multimodal data to globally extract color, texture, and shape features, which are then compressed and encoded to generate parent attribute labels. Local color, texture, and shape features are then extracted from the corresponding local area of the global data, and distributedly encoded to generate child attribute labels.
[0166] Step F12: Based on the parent attribute and the child attribute, time alignment and redundant information compression are performed to generate the feature vector.
[0167] In this embodiment, time alignment refers to unifying the timestamps of multi-source data to a reference clock. Compressing redundant information refers to removing low-value features through a dimensionality reduction algorithm.
[0168] As an optional implementation, the parent attribute and the child attribute are input, the timestamps are aligned based on the preset pulse, the parent attribute and the child attribute are concatenated to form a high-dimensional vector, which is compressed to the preset dimension through principal component analysis and the feature vector is output.
[0169] As another optional implementation, the parent attribute and the child attribute synchronize the time axis through a preset protocol, the merged vector is reduced in dimension through factor analysis, and a feature vector is generated and output.
[0170] For example, in a smart city traffic dispatch scenario, visible light cameras and millimeter-wave radars deployed at arterial intersections collect real-time vehicle video streams and point cloud trajectories. These are then processed in parallel within a cluster of edge computing nodes. RGB third-order color moments are extracted from the full-width video frames to generate a traffic flow density heat map. Distributed PCA encoding is then used to output parent attribute labels. Motion vectors are extracted from the radar point cloud for the emergency vehicle target area: speed = 72 km / h, acceleration = 1.2 m / s². SIFT keypoint spatial descriptors are then calculated and compressed using a distributed autoencoder to generate sub-attributes: "Vehicle Type = Ambulance" and "Location = Northbound Southbound Lane." Sub-attribute timestamps are aligned based on the parent attribute's time base. The parent and sub-attributes are then concatenated to form a high-dimensional vector: "Main Road," 0.85 (congestion index), "Ambulance." Dimensionality reduction is then performed using t-SNE to output a 16-dimensional feature vector: 0.72, -1.3, 0.05... This is then sent to the traffic light control system, causing the target intersection to extend the green light by 12 seconds in real time, reducing emergency vehicle delays by 40%.
[0171] By extracting global and local features of multimodal data, feature data in global scenarios and feature data in local scenarios are obtained simultaneously, thereby improving the accuracy and reliability of fusion result analysis.
[0172] Based on any of the above embodiments, in the eighth embodiment of the present application, step F11 includes steps G11 to G12:
[0173] Step G11: extracting global features of the multimodal data, encoding the global features through distributed clustering, and generating parent attributes corresponding to the global features.
[0174] In this embodiment, global features refer to macroscopic statistical properties of the complete data set. Distributed cluster refers to a multi-node parallel computing architecture.
[0175] As an optional implementation, multimodal data is input, global features corresponding to the multimodal data are extracted, and distributed cluster computing is performed on the global features. The first node group calculates the global color histogram of the image in blocks, and the second node group fits the normal distribution histogram. The global color histogram and normal distribution histogram of the image are concatenated and encoded as the label of the parent attribute.
[0176] Step G12: extracting local features of the target area corresponding to the multimodal data, performing gradient calculation and encoding on the local features through computing nodes, and generating sub-attributes corresponding to the local features.
[0177] In this embodiment, the target region refers to a specific spatial range to be analyzed. A computing node refers to a hardware unit that performs operations. Gradient calculation refers to a mathematical operation that quantifies the rate of change of a feature.
[0178] As an optional implementation, local features corresponding to the target region of the multimodal data are extracted and assigned to computational nodes. Key points of the local features are extracted, and the gradient magnitude and direction are calculated in the vicinity of the key points. These gradient magnitudes and directions are concatenated into gradient features, and principal component analysis is performed on the gradient features to obtain analysis results. The analysis results are compressed into sub-attributes of a preset dimension and output.
[0179] As another optional implementation, local features corresponding to the target area of the multimodal data are extracted and processed by computational nodes. The normal vector angle distribution of these local features is extracted, and the point cloud curvature field gradient is calculated. The features of the normal vector angle distribution and the point cloud curvature field gradient are then concatenated and subjected to linear discriminant analysis to reduce the dimensionality to a sub-attribute of a preset dimension.
[0180] For example, in the smart city security monitoring scenario, multimodal data is constructed by satellite images and drone point clouds, and global features are extracted in a distributed cluster: node group A calculates the HSV color moment of the full image in blocks, node group B fits the normal vector distribution histogram, and after AllReduce aggregation, PCA reduces the dimension to 16 dimensions, and encodes and generates parent attribute labels: "high-rise building complex" and "main road". At the same time, the target areas are delineated: building facades (contour polygons mapped by BIM models) and road cracks (morphological detection area > 50cm² area) and assigned to edge computing nodes: SURF key points are extracted in the building facade area: 128-dimensional descriptors, and the gradient direction histogram is calculated: 8-directional amplitude, which is compressed to 8 dimensions by PCA; the gray-level co-occurrence matrix is used to calculate the energy value in the road crack area, and the gradient modulus peak is obtained in combination with the Roberts operator (threshold > 0.6), and LDA reduces the dimension to 2 dimensions; the final output sub-attribute vectors: "facade key point density = 4.2 points / m²" and "crack energy index = 0.78".
[0181] By generating parent attributes and child attributes, semantics and relationships are matched, integrated and optimized, which reduces the data resource occupancy rate and improves the efficiency of fusion result analysis.
[0182] Based on any of the above embodiments, in Embodiment 9 of the present application, step F12 includes steps H11 to H12:
[0183] In step H11 , after synchronizing the time axes of the parent attribute and the child attribute according to a preset protocol, the synchronized parent attribute and child attribute are concatenated to generate a high-dimensional vector.
[0184] In this embodiment, the preset protocol refers to predefined spatiotemporal synchronization rules. Timeline synchronization refers to aligning the time tags of data from different sources to a common reference. Splicing refers to the operation of connecting multiple feature vectors by dimension.
[0185] As an optional implementation, the parent attribute and the child attribute are aligned on a time axis through a preset protocol, the attribute value is intercepted at a preset time, and the parent attribute and the child attribute are serially concatenated to generate a high-dimensional vector.
[0186] As another optional implementation, the parent attribute and the child attribute are synchronized using an atomic clock, the parent attribute is compressed to a preset dimension, the child attribute is linearly interpolated to the same dimension of the compressed parent attribute, and the corresponding dimensions are added to generate a high-dimensional vector.
[0187] Step H12: perform principal component analysis on the high-dimensional vector and perform dimensionality compression and dimensionality reduction according to a preset dimension to generate the feature vector.
[0188] In this embodiment, the preset dimension refers to the pre-set output vector dimension. Compression dimensionality reduction refers to the computational process of reducing feature dimensions while retaining core information. Principal component analysis refers to a dimensionality reduction method that converts related features into linearly independent principal components through orthogonal transformation.
[0189] As an optional implementation, a high-dimensional vector is input, and its mean is normalized to zero and its standard deviation is normalized to obtain the original data. The sample matrix is decomposed by singular value, and a preset number of left singular vectors are taken as the basis. The original data is projected into the subspace to generate the eigenvector.
[0190] As another optional implementation, the high-dimensional vector is subjected to correlation coefficient matrix analysis, the principal component direction is solved using the power iteration method, the maximum variance direction is intercepted according to the preset dimension, the matrix is diagonalized by the rotation method, and the eigenvector is generated after projection.
[0191] For example, in a smart city traffic dispatch scenario, data is collected from multiple sensors at intersections. Parent attributes are generated by global feature extraction, encoding the entire road network topology as a 16-dimensional vector: "main road," "emergency priority lane." Child attributes are generated by local features, reducing the 128-dimensional key points of the emergency vehicle's outline to 8 dimensions: "Speed = 62 km / h," "Acceleration = 1.3 m / s²." Time is synchronized using the PTP precision clock protocol, at a UTC second, such as t=20230716T14:30:00Z. Parent and child attribute values are truncated and serially concatenated to generate a 24-dimensional high-dimensional vector: 0.83, -0.12, ..., 62.0, 1.3. Principal component analysis is used to calculate the eigenvalues of the covariance matrix. The first 16 principal components are truncated according to the preset dimension of 16, and then multiplied by the projection matrix to reduce the dimensionality and generate the eigenvector: 0.72, -1.3, 0.05.
[0192] Because distributed feature extraction covers both global and local feature extraction, it can accurately extract basic attributes such as color, texture, and shape, and then generate feature vectors, thereby improving the accuracy and reliability of the fusion result analysis.
[0193] Based on any of the above embodiments, in the tenth embodiment of the present application, step S10 includes steps I11 to I13:
[0194] Step I11: After performing target analysis on the multimodal data, identify the entity corresponding to the feature vector.
[0195] In this embodiment, recognition refers to the technical process of mapping feature vectors to entity identifiers. Entities refer to physical object instances identified by features. Relational networks refer to entity association models with graph structures.
[0196] As an optional implementation, the feature vector is fed into a pre-trained classification model, which then outputs a probability distribution of categories: first probability, second probability, and third probability. The category with the highest probability is selected as the entity corresponding to the feature vector.
[0197] As another optional implementation, based on a reinforcement learning decision tree model, the feature vector is input into the model, and the entity corresponding to the leaf node is finally identified according to the state space partitioning rule.
[0198] Step I12: determining the relationship network between the entities based on the spatial topological association and the event logical association in the semantic association.
[0199] In this embodiment, semantic association refers to a predefined logical relationship framework between entities. Spatial topological association refers to a geometric relationship formed based on the spatial location of entities. Event logical association refers to a causal or conditional relationship derived from a sequence of events. A relational network refers to a graph structure with entities as nodes and associations as edges.
[0200] As an optional implementation, the spatial topological associations of the input entity set are calculated to generate inclusion edges. The state change times in the event log are parsed to derive event logical association edges. The topological associations and event logical association edges are combined to construct a relationship network.
[0201] As another optional implementation, spatial topological associations are calculated using entity sets. When the distance between line segment endpoints is less than a threshold, an adjacent edge is generated. When the response delay in an event sequence is less than a preset delay, an event logical association edge is established. The adjacent edges and event logical association edges are fused to generate a relationship network.
[0202] Step I13: extracting rules and elements from the relationship network, and associating and aggregating the rules and elements to generate the knowledge elements.
[0203] In this embodiment, rule extraction refers to extracting frequent patterns from a relational network. Element extraction refers to extracting key parameters from relationships. Association aggregation refers to binding rules and elements into executable units. Knowledge element database refers to a structured, computable rule base.
[0204] As an optional implementation, a relational network is input and rules are extracted through frequent subgraph mining. Spatial thresholds and event delays are parsed from edge attributes as elements. Rules are then associated with elements to generate knowledge elements.
[0205] As another alternative, association rule learning is performed on the relationship network to mine high-frequency item sets. Key elements are extracted: the maximum distance between adjacent entities and the minimum interval between causal events. Association rules and key elements are aggregated into knowledge elements.
[0206] For example, in a smart city traffic dispatch scenario, intersection cameras and radar capture vehicle video streams and point cloud data. Target analysis identifies the entities corresponding to the feature vectors: E01 (ambulance), E02 (traffic light), and E03 (main road). Spatial topological associations are used to calculate the inclusion relationship between E01's trajectory point and E03's road polygon, achieving an inclusion rate of >98%. Combined with event logic, the green light extension event for E02, when E01 approaches the intersection, is analyzed, with a delay of <0.5 seconds. A relationship network is constructed: {edges: E01 to E03 "path attribution", E01 to E02 "trigger extension"}. The rule "extend green light by 15 seconds when ambulance is less than 200 meters from the intersection" and the element "distance threshold 200 meters, extension time 15 seconds" are extracted from the network. These knowledge elements are then generated and pushed to the signal control system.
[0207] Because the generation of knowledge elements makes the relationship between each data clearer, the fusion result will also reflect the relationship between the data, thereby improving the reliability of the fusion result.
[0208] Based on any of the above embodiments, in the eleventh embodiment of the present application, step S30 includes steps J11 and J12:
[0209] Step J11 , based on the data type, screening the association between the data type and the alignment strategy to obtain an initial strategy corresponding to the multimodal data.
[0210] In this embodiment, the association between data types and alignment strategies refers to a predefined mapping rule library. Screening refers to the process of matching association rules based on the input data type. The initial strategy refers to the basic alignment method and parameter set for matching the current data type.
[0211] As an optional implementation, a data type is input, rules corresponding to the data type in a database of association relationships between data types and alignment strategies are retrieved, and an initial strategy corresponding to the multimodal data is extracted and output.
[0212] Step J12: According to the feedback rule library, a feedback compensation strategy is added to the initial strategy to generate the target strategy.
[0213] In this embodiment, the feedback rule base refers to a database that stores configuration error response logic. The feedback compensation strategy refers to an incremental optimization solution for errors. The target strategy refers to the optimization instruction set that combines the initial strategy and the compensation strategy.
[0214] As an optional implementation, the initial policy is run with a feedback rule base to monitor abnormal operation and trigger a compensation policy. Based on the initial policy and the compensation policy, a target policy is generated.
[0215] For example, in a smart city traffic dispatch scenario, when the data type is determined to be dynamic traffic flow data with an update frequency ≥ 10Hz, the initial strategy ID_D004 is selected based on the association relationship "Dynamic and Kalman Filter": the state vector has 6 dimensions and the noise covariance Q = 0.01I. During execution, the feedback rule base detects that the trajectory offset is greater than 2 meters due to multipath interference at the overpass. This triggers the compensation strategy FBC_D004: expanding the state vector to 9 dimensions, adding angular velocity and curvature radius, and dynamically increasing the process noise covariance to Q = 0.05I, thus generating the target strategy.
[0216] Due to the setting of the feedback mechanism, problems that arise during the test fusion data can be promptly addressed through the addition of compensation strategies through the feedback mechanism to ensure the reliability of the fusion results.
[0217] The present application provides a device for spatiotemporal registration and alignment, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method for cross-scale spatiotemporal registration and alignment in the above-mentioned embodiment one.
[0218] Reference below Figure 8 , which shows a schematic diagram of the structure of a device suitable for implementing spatiotemporal registration and alignment in an embodiment of the present application. The spatiotemporal registration and alignment devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, smart city control terminals, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), embedded processing terminals, and fixed terminals such as mobile digital platforms and desktop computers. Figure 8 The device for spatiotemporal registration and alignment shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0219] like Figure 8 As shown, the spatiotemporal registration and alignment device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the spatiotemporal registration and alignment device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007, including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008, including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003, including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. The communication devices 1009 can allow the spatiotemporal registration and alignment device to communicate with other devices wirelessly or wired to exchange data. Although the spatiotemporal registration and alignment device is shown with various systems, it should be understood that implementation or presence of all the illustrated systems is not required. More or fewer systems may alternatively be implemented or present.
[0220] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0221] The spatiotemporal registration and alignment device provided by this application adopts the cross-scale spatiotemporal registration and alignment method in the above-mentioned embodiment, which can solve the technical problem that the fusion of multimodal data in the related art requires the mandatory unification of the coordinate system and time base, solidifies the fusion method of multimodal data, and thus leads to poor multi-source data fusion effect. Compared with the existing technology, the beneficial effects of the spatiotemporal registration and alignment device provided by this application are the same as the beneficial effects of the cross-scale spatiotemporal registration and alignment method provided by the above-mentioned embodiment, and the other technical features of the spatiotemporal registration and alignment device are the same as the features disclosed in the method of the previous embodiment, and will not be repeated here.
[0222] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0223] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0224] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the method for cross-scale spatiotemporal registration and alignment in the above-mentioned embodiment.
[0225] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0226] The computer-readable storage medium may be included in the device for spatiotemporal registration and alignment; or may exist independently without being incorporated into the device for spatiotemporal registration and alignment.
[0227] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the device for spatiotemporal registration and alignment, the device for spatiotemporal registration and alignment: performs global feature extraction and local feature extraction on at least one multimodal data obtained through the data source to generate a feature vector after dimensionality reduction, and constructs a relationship network between the entities according to the entities corresponding to the multimodal data to generate knowledge elements; determines the data type of the multimodal data according to the feature vector and the knowledge elements; determines the target strategy corresponding to the multimodal data based on the data type and the association between the data type and the alignment strategy; fuses the feature vector and the corresponding knowledge element to generate a fusion result, and based on the target strategy, spatiotemporally aligns the fusion result to generate a target result.
[0228] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0229] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0230] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0231] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned method for cross-scale spatiotemporal registration and alignment. This computer-readable storage medium can address the technical issue in related technologies where multimodal data fusion requires a unified coordinate system and time base, solidifying the multimodal data fusion method and leading to poor multi-source data fusion results. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the method for cross-scale spatiotemporal registration and alignment provided in the aforementioned embodiments, and are not further elaborated here.
[0232] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A method for cross-scale spatiotemporal registration and alignment, characterized in that: The method comprises: Preprocess the acquired multimodal data to generate feature vectors and knowledge elements; Analyzing time update characteristics and spatial position characteristics based on the feature vector and the corresponding knowledge element to determine the data type of the multimodal data; According to the data type, screening in the association relationship between the data type and the alignment and registration strategy to determine the target strategy corresponding to the multimodal data; Transform and project the feature vector into a three-dimensional hierarchical coordinate system, and use the knowledge element as a constraint function to fuse the projected feature vector and its constraint function to generate a fusion result; Calibrate the timestamps of each data in the fusion result, and determine the missing parts of the timestamps corresponding to the fusion result; According to the time interpolation algorithm, reasonable data is added to the time points corresponding to the missing parts of the timestamp to generate an initial result; Performing alignment operations on the initial results at different time scales to generate time-aligned results; Performing initial alignment and fine alignment of the time alignment results according to the geometric transformation and precision algorithm of the target strategy, and generating a target result by aligning the time alignment results at various spatial scales; Obtaining at least one evaluation indicator based on monitoring the alignment and registration effect of the target result; Quantitatively analyzing the evaluation indicators, evaluating the alignment and registration effect of the target result, and generating an evaluation result; Based on the evaluation result, analyzing the dynamic level and spatial resolution level corresponding to the target result to generate feedback information; generating a feedback strategy for adjusting the time alignment frequency and spatial registration mode of the fusion result according to the feedback information; The feedback strategy and the target strategy are fused and adjusted through a learning algorithm to generate target parameters.
2. The method according to claim 1, wherein The step of analyzing the time update characteristics and spatial position characteristics based on the feature vector and the corresponding knowledge element to determine the data type of the multimodal data includes: Extracting the time identification field and the space identification field corresponding to the feature vector according to preset index bits from the time series field obtained by parsing the feature vector; Associating the time identification field and the space identification field according to the knowledge elements, and determining the spatiotemporal label corresponding to the feature vector based on the associated field relationship; Based on the data type judgment rule, the spatiotemporal label matching is judged to determine the data type of the multimodal data.
3. The method according to claim 1, wherein The step of preprocessing the acquired multimodal data to generate feature vectors and knowledge elements includes: Performing global feature extraction and distributed encoding on the multimodal data to generate a parent attribute, and performing local feature extraction and distributed encoding on the multimodal data to generate a child attribute; Based on the parent attribute and the child attribute, time alignment and redundant information compression are performed to generate the feature vector.
4. The method according to claim 3, wherein The steps of performing global feature extraction and distributed encoding on the multimodal data to generate a parent attribute, and performing local feature extraction and distributed encoding on the multimodal data to generate a child attribute include: Extracting global features of the multimodal data, encoding the global features through distributed clustering, and generating parent attributes corresponding to the global features; Extract local features of the target area corresponding to the multimodal data, perform gradient calculation and encoding on the local features through computing nodes, and generate sub-attributes corresponding to the local features.
5. The method according to claim 3, wherein The step of performing time alignment and compressing redundant information based on the parent attribute and the child attribute to generate the feature vector includes: After synchronizing the time axes of the parent attribute and the child attribute according to a preset protocol, splicing the synchronized parent attribute and child attribute to generate a high-dimensional vector; The high-dimensional vector is subjected to principal component analysis, and is compressed and reduced in dimension according to a preset dimension to generate the feature vector.
6. The method according to claim 1, wherein The step of preprocessing the acquired multimodal data to generate feature vectors and knowledge elements includes: After performing target analysis on the multimodal data, identifying the entity corresponding to the feature vector; Determine a relationship network between the entities based on the spatial topological association and the event logical association in the semantic association; Rules and elements are extracted from the relationship network, and the rules and elements are associated and aggregated to generate the knowledge elements.
Citation Information
Patent Citations
Road condition monitoring identification method, device and equipment based on machine vision and medium
CN119445503A
Small and medium-sized reservoir operation safety evaluation system based on sky-ground water conservancy project
CN120338508A