A spatiotemporal contract-driven multi-source heterogeneous data analysis method and system
Patent Information
- Application Number
- CN202611327585.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]针对现有技术的不足,本发明的目的在于提供一种时空契约驱动的多源异构数据分析方法及系统,旨在解决现有技术中多源时空数据分析存在的时空需求契约化缺失、数据时空契约适配性差、任务拆解无统一协议支撑、调度逻辑缺乏契约感知的技术问题
[0016]与现有技术相比,本发明的有益效果在于:通过基于所述分析需求及所述多源时空数据生成所述数据契约文档,为后续的数据分析提供了时空契约规范,为时空需求契约化奠定了基础;通过所述数据契约文档规范所述多源时空数据,完成了异构数据间的自动化的时空属性对齐,提升了所述数据契约文档与所述标准化数据集之间的适配度;通过将所述数据契约文档拆分为所述原子化时空计算单元,其数据契约、执行算子及标准化制品构成的DOP协议为任务的拆解提供了统一的协议支撑,能够构建结构一致、约束可追溯的多源时空数据分析流程,有效缓解了多源数据格式不统一、时空约束混乱、调度协同性差的问题,具备良好的泛化与可扩展性。
Smart Images

Figure CN122840035A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a spatiotemporal contract-driven method and system for analyzing multi-source heterogeneous data. Background Technology
[0002] Multi-source spatiotemporal data analysis requires the integration of spatial location, time series, and attribute information, which is a core requirement in fields such as environment, transportation, and land resources. Existing technologies rely on traditional GIS tools (such as ArcGIS and QGIS) or general data analysis platforms, which have the following shortcomings in intelligent analysis processes: 1. Lack of explicit spatiotemporal requirement constraints: Traditional tools lack a dedicated mechanism to transform vague requirements (such as "analyze regional ecological changes") into explicit spatiotemporal constraints (coordinate system, time granularity, spatial resolution). Users must manually define parameters, which can easily lead to the omission of key constraints, causing results to deviate from actual needs. Furthermore, manually breaking down complex requirements with spatiotemporal dimensions is time-consuming and has low accuracy; 2. Poor adaptability of multi-source data spatiotemporal contracts: shapefile vectors, GPS... 1. Heterogeneous data such as time series (including timestamps) have large differences in spatiotemporal attributes. Existing platforms require users to manually handle contract consistency (such as transforming coordinate systems and aligning timestamps), resulting in a high error rate related to spatiotemporal attributes; 2. Task decomposition lacks unified protocol support: Subtask decomposition only focuses on logical flow (such as "preprocessing → analysis → visualization"), and spatiotemporal constraints are not rigidly passed, leading to contract conflicts when executing subsequent tasks (such as using low-resolution data for refined plot-level analysis); 3. Scheduling logic lacks contract awareness: Task dependencies only consider data flow and do not verify spatiotemporal contract consistency, resulting in a high proportion of invalid calculations. Summary of the Invention
[0003] To address the shortcomings of existing technologies, the present invention aims to provide a spatiotemporal contract-driven method and system for multi-source heterogeneous data analysis, which seeks to solve the technical problems in existing multi-source spatiotemporal data analysis, such as the lack of spatiotemporal demand contractualization, poor data spatiotemporal contract adaptability, lack of unified protocol support for task decomposition, and lack of contract awareness in scheduling logic.
[0004] To achieve the above objectives, firstly, embodiments of this application provide a spatiotemporal contract-driven method for analyzing multi-source heterogeneous data, comprising the following steps: A data contract document, including spatiotemporal constraint parameters, is obtained by using multi-source spatiotemporal data and analysis requirements input by the user. Based on the data contract document, the multi-source spatiotemporal data is standardized to obtain a standardized dataset corresponding to the data contract document. The data contract document is decomposed into several atomic spatiotemporal computing units with interconnected relationships. Each atomic spatiotemporal computing unit includes a data contract, an execution operator, and a standardized artifact. The data flow dependencies between each atomic spatiotemporal computing unit are determined. Using the atomic spatiotemporal computing unit as the topology node and the data flow dependencies as the topology edge, a spatiotemporal computing unit topology structure is constructed. A consistency check is performed on the topology of the spatiotemporal computing unit. If the consistency check passes, the input constraints of the atomized spatiotemporal computing unit are derived based on the reverse topological order of the spatiotemporal computing unit topology, and all the input constraints are combined into a spatiotemporal computing contract. Based on the data contract document, the spatiotemporal computing contract, and the topology of the spatiotemporal computing unit, executable analysis code carrying contract verification breakpoints is generated. The executable analysis code is executed to perform full-process contract verification through the contract verification breakpoints and output the analysis results.
[0005] Furthermore, the step of obtaining a data contract document including spatiotemporal constraint parameters based on user-input multi-source spatiotemporal data and analysis requirements includes: Extract the core fields and spatiotemporal attributes of the multi-source spatiotemporal data, and compare the core fields and spatiotemporal attributes with the analysis requirements to determine whether data adjustment is needed for the multi-source spatiotemporal data. If no data adjustment is required for the multi-source spatiotemporal data, then the presence of temporal keywords and spatial keywords in the analysis requirements is identified based on a preset spatiotemporal dictionary. If the analysis requirements contain temporal and spatial keywords, then a spatiotemporal fusion contract template including spatiotemporal constraints and custom constraints is invoked. The spatiotemporal constraints and custom constraints are filled with preset parameter values and the analysis requirements to obtain a data contract document including spatiotemporal constraint parameters.
[0006] Furthermore, the preset spatiotemporal dictionary includes several preset keywords, and the step of identifying whether temporal and spatial keywords exist in the analysis requirements based on the preset spatiotemporal dictionary includes: The analysis requirements are segmented to obtain the keywords to be identified. Set keyword weights for preset keywords and obtain the edit distance between the keyword to be identified and the preset keywords. Based on the keyword weights and edit distances, obtain the word similarity between the keyword to be identified and the preset keywords. The word similarity is compared with the similarity threshold. If the word similarity is greater than the similarity threshold, the keyword to be identified is selected as a temporal keyword or a spatial keyword.
[0007] Furthermore, the spatiotemporal constraint parameters include coordinate system parameters, spatial range parameters, time interval parameters, and time granularity parameters; the multi-source spatiotemporal data includes initial vector data and initial time-series data; and the step of standardizing the multi-source spatiotemporal data based on the data contract document to obtain a standardized dataset corresponding to the data contract document includes: The initial vector data is transformed using the coordinate system parameters to obtain stage vector data, and the initial time series data is aggregated using the time granularity parameters to obtain stage time series data. Based on the spatial range parameter, the stage vector data is spatially clipped to obtain standard vector data. Based on the time interval parameter, the stage time series data is time-sliced to obtain standard time series data. The standard vector data and the standard time series data are combined to form a standardized dataset corresponding to the data contract document.
[0008] Furthermore, the step of decomposing the data contract document into several interconnected atomic spatiotemporal computation units, each atomic spatiotemporal computation unit comprising a data contract, an execution operator, and a standardized artifact, includes: Based on the analysis requirements, several initial units with interconnected relationships are constructed, and it is determined whether an upstream unit exists for each initial unit based on the interconnected relationships. If the initial unit does not have an upstream unit, a data contract is set for the initial unit based on the data contract document; if the initial unit has an upstream unit, the output of the upstream unit is used as the data contract of the initial unit. An execution operator is assigned to the initial unit through the data contract, and a standardized artifact is obtained based on the data contract and the execution operator. The data contract, the execution operator, and the standardized artifact constitute the atomized spatiotemporal computing unit.
[0009] Furthermore, the step of performing consistency verification on the spatiotemporal computing unit topology includes: Perform global loop dependency detection on the spatiotemporal computing unit topology; Obtain the rationality score of the spatiotemporal computing unit topology and obtain the contract compatibility score between two atomic spatiotemporal computing units with data flow dependencies; If no circular dependency structure is detected, the reasonableness score is greater than or equal to the first score threshold, and the contract compatibility score is greater than or equal to the second score threshold, then the consistency check is considered to have passed.
[0010] Furthermore, the formula for obtaining the reasonableness score is as follows: , in, Indicates the reasonableness score. Indicates the number of effective dependencies. This represents the total number of data flow dependencies. This indicates the number of redundant dependencies. This represents the redundancy penalty coefficient.
[0011] Furthermore, the steps of generating executable analysis code carrying contract verification breakpoints based on the data contract document, spatiotemporal computing contract, and spatiotemporal computing unit topology, executing the executable analysis code to perform full-process contract verification through the contract verification breakpoints, and outputting analysis results include: Based on the contract-based programming paradigm, constraint parameters in the data contract document and the spatiotemporal computing unit topology are extracted to obtain the pre-constraint conditions, and the product specifications of the standardized products in the atomic spatiotemporal computing unit are extracted to obtain the post-constraint conditions. The pre-constraints and post-constraints are converted into assertion verification statements, and the assertion verification statements are embedded as contract verification breakpoints into the execution function to generate executable analysis code carrying contract verification breakpoints. Based on the execution order determined by the spatiotemporal computing unit topology, the executable analysis code is executed sequentially. If all contract verification breakpoints are passed, the full-process contract verification is deemed successful, and the analysis results are output.
[0012] Furthermore, after the steps of generating executable analysis code carrying contract verification breakpoints based on the data contract document, spatiotemporal computing contract, and spatiotemporal computing unit topology, executing the executable analysis code to perform full-process contract verification through the contract verification breakpoints, and outputting the analysis results, the method further includes: Determine if the user needs to modify the contract. If so, regenerate the analysis results.
[0013] Secondly, embodiments of this application provide a spatiotemporal contract-driven multi-source heterogeneous data analysis system, applied to the spatiotemporal contract-driven multi-source heterogeneous data analysis method as described in the first aspect above, the system comprising: The parsing module is used to obtain a data contract document including spatiotemporal constraint parameters based on the multi-source spatiotemporal data and analysis requirements input by the user, and to perform standardization processing on the multi-source spatiotemporal data based on the data contract document to obtain a standardized dataset corresponding to the data contract document. The encapsulation module is used to decompose the data contract document into several atomic spatiotemporal computing units with interconnected relationships. The atomic spatiotemporal computing unit includes a data contract, an execution operator, and a standardized artifact. The module determines the data flow dependency relationship between each atomic spatiotemporal computing unit and constructs a spatiotemporal computing unit topology structure with the atomic spatiotemporal computing unit as the topology node and the data flow dependency relationship as the topology edge. The verification module is used to perform consistency verification on the topology of the spatiotemporal computing unit. If the consistency verification passes, the input constraints of the atomic spatiotemporal computing unit are derived based on the reverse topological order of the spatiotemporal computing unit topology, and all the input constraints are combined into a spatiotemporal computing contract. The execution module is used to generate executable analysis code carrying contract verification breakpoints based on the data contract document, spatiotemporal computing contract, and spatiotemporal computing unit topology, execute the executable analysis code to perform full-process contract verification through the contract verification breakpoints, and output the analysis results.
[0014] Thirdly, embodiments of this application provide a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the spatiotemporal contract-driven multi-source heterogeneous data analysis method as described in the first aspect above.
[0015] Fourthly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the spatiotemporal contract-driven multi-source heterogeneous data analysis method as described in the first aspect above.
[0016] Compared with existing technologies, the beneficial effects of this invention are as follows: By generating the data contract document based on the analysis requirements and the multi-source spatiotemporal data, a spatiotemporal contract specification is provided for subsequent data analysis, laying the foundation for the contractualization of spatiotemporal requirements; by standardizing the multi-source spatiotemporal data through the data contract document, automated spatiotemporal attribute alignment between heterogeneous data is achieved, improving the adaptability between the data contract document and the standardized dataset; by splitting the data contract document into atomic spatiotemporal computing units, the DOP protocol composed of its data contract, execution operators, and standardized artifacts provides unified protocol support for task decomposition, enabling the construction of a structurally consistent and constraint-traceable multi-source spatiotemporal data analysis process, effectively alleviating the problems of inconsistent multi-source data formats, chaotic spatiotemporal constraints, and poor scheduling coordination, and possessing good generalization and scalability. Attached Figure Description
[0017] Figure 1 This is a flowchart of the spatiotemporal contract-driven multi-source heterogeneous data analysis method in the first embodiment of the present invention; Figure 2 This is a structural block diagram of the spatiotemporal contract-driven multi-source heterogeneous data analysis system in the second embodiment of the present invention; The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0018] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0019] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0021] Please see Figure 1 The spatiotemporal contract-driven multi-source heterogeneous data analysis method provided in the first embodiment of the present invention includes the following steps: S10: Obtain a data contract document including spatiotemporal constraint parameters based on the multi-source spatiotemporal data and analysis requirements input by the user; perform standardization processing on the multi-source spatiotemporal data based on the data contract document to obtain a standardized dataset corresponding to the data contract document. In this embodiment, the analysis requirement submitted by the user is: to analyze the distribution of congested road sections during the morning rush hour in a certain city, identify the congestion time periods and core congested road sections, and the multi-source spatiotemporal data input by the user are: road network data in shapehp format and GPS vehicle speed data in CSV format.
[0022] Step S10 includes: S110: Extract the core fields and spatiotemporal attributes of the multi-source spatiotemporal data, and compare the core fields and spatiotemporal attributes with the analysis requirements to determine whether data adjustment is needed for the multi-source spatiotemporal data; The uploaded shapefile format road network data and CSV format GPS vehicle speed data are parsed to extract core fields and spatiotemporal attributes: the shapefile format road network data contains core fields such as road segment ID, road segment length, and spatial coordinates, and its spatial coverage is confirmed to be the core urban area of a target city; the CSV format GPS vehicle speed data contains core fields such as vehicle ID, latitude and longitude, vehicle speed, and timestamp, with a time range covering 6:00-10:00, which covers the morning rush hour. The analysis requirement (analyzing the distribution of congested road segments during the morning rush hour in a city, identifying congestion periods and core congested road segments) explicitly requires data on morning rush hour and congested road segments. Since the uploaded data already covers these requirements in terms of region, vehicle speed, and time range, step S120 can be directly performed. If data adjustment is needed for the multi-source spatiotemporal data, a corresponding data supplementation request is sent to the user for data supplementation.
[0023] S120: If no data adjustment is required for the multi-source spatiotemporal data, then identify whether there are temporal keywords and spatial keywords in the analysis requirements based on the preset spatiotemporal dictionary; S130: If the analysis requirements contain temporal keywords and spatial keywords, then call the spatiotemporal fusion contract template that includes spatiotemporal constraints and custom constraints, and fill the spatiotemporal constraints and custom constraints with preset parameter values and the analysis requirements to obtain a data contract document that includes spatiotemporal constraint parameters. Understandably, the preset spatiotemporal dictionary includes several preset keywords, which are preset time prompts or preset space prompts. Specifically, word segmentation is performed on the analysis requirements to obtain the keywords to be identified. Set keyword weights for preset keywords and obtain the edit distance between the keyword to be identified and the preset keywords. Based on the keyword weights and edit distances, obtain the word similarity between the keyword to be identified and the preset keywords. Keyword weights range from 0.1 to 1.0, and are set based on the importance of preset keywords in spatiotemporal analysis, combined with differences in spatiotemporal dimensions. The formula for obtaining word similarity is: , in, This indicates that the keyword to be identified is related to the first... Word similarity between preset keywords Indicates the first Keyword weight of a preset keyword, This indicates that the keyword to be identified is related to the first... The edit distance between preset keywords, as can be understood, represents the minimum number of character operations required to convert between the keyword to be identified and the preset keywords. This indicates the text length of the keyword to be identified and the length of the first key. The maximum value among the text lengths of the preset keywords.
[0024] The word similarity is compared with the similarity threshold. If the word similarity is greater than the similarity threshold, the keyword to be identified is selected as a time-series keyword or a spatial keyword. Understandably, if the preset keyword is a preset time prompt word, the keyword to be identified is a time-series keyword; if the preset keyword is a preset spatial prompt word, the keyword to be identified is a spatial keyword.
[0025] In this embodiment, the analysis requirements include morning peak (temporal keywords) and road segment (spatial keywords). Therefore, it is determined to call a spatiotemporal fusion contract template that includes spatiotemporal constraints and custom constraints. Understandably, if only temporal keywords or spatial keywords exist, a single-type template will be called to fill them in. This application only performs corresponding analysis on multi-source spatiotemporal data.
[0026] It should be noted that the spatiotemporal fusion contract template also includes: requirement type, analysis target and data type. The spatiotemporal constraints include time interval, time granularity, timestamp accuracy, coordinate system, spatial range, road network topology requirements, and GPS resolution requirements. The custom constraints include logical items. In this embodiment, the logical item is the vehicle speed threshold.
[0027] Correspondingly, the spatiotemporal constraint parameters include coordinate system parameters, spatial range parameters, time interval parameters, and time granularity parameters. In this embodiment, the data contract document is: { "Requirement Type": "Spatiotemporal Fusion", "Analysis Objective": "Analysis of Congested Road Sections During Morning Rush Hour in a Certain City", "Data Type": ["shp Road Network Data", "GPS Vehicle Speed Data"], "Spatiotemporal Constraints": {"Time Interval": "2025-05-20 7:00-9:00" (Analysis Requirement), "Time Granularity": "5 minutes" (Preset Parameter Value), "Timestamp Accuracy": "Minute Level" (Preset Parameter Value), "Coordinate System": "EPSG:4326" (Preset Parameter Value), "Spatial Range": "116.3°-116.5°E, 39.9°-40.1°N" (Analysis Requirement), "Road Network Topology Requirements": "No Repeating Surfaces, No Suspended Points", "GPS Resolution Requirements": "Latitude and Longitude Error ≤ 0.0001° (Preset Parameter Value)"}, "Custom Constraints": {"Vehicle Speed Threshold":} "Vehicle speed < 20km / h" (Analysis requirements)}.
[0028] S140: Perform coordinate system transformation on the initial vector data using the coordinate system parameters to obtain stage vector data; perform data aggregation on the initial time series data using the time granularity parameters to obtain stage time series data. S150: Spatial clipping is performed on the stage vector data based on the spatial range parameter to obtain standard vector data; time slicing is performed on the stage time series data based on the time interval parameter to obtain standard time series data; and the standard vector data and the standard time series data are combined into a standardized dataset corresponding to the data contract document. The multi-source spatiotemporal data includes initial vector data and initial time-series data. In this embodiment, the road network data in shp format is the initial vector data, which has a native coordinate system. The GDAL tool is used to convert the native coordinate system to the contract-specified coordinate system. Specifically, the initial vector data undergoes a coordinate system normalization affine transformation, and the transformation formula is as follows: , in, Represents the planar coordinates of the initial vector data. The normalized coordinates in the target coordinate system represent the planar coordinates of the stage vector data, and M represents the affine transformation matrix obtained based on the source and target coordinate systems.
[0029] The data is then cropped, specifically based on spatial boundaries. Perform spatial clipping, preserving all areas that fall within the clipping area. Vector elements within the specified range are used to generate standard vector data. The initial time-series GPS vehicle speed data in CSV format contains raw timestamps at the second level. The Pandas library is used to aggregate the data, converting the second-level data into 5-minute-level data (extracting the average vehicle speed within each 5-minute interval). Morning rush hour data is extracted according to the contractually specified time. Specifically, the aggregation formula is: , in, This represents the aggregated value of the i-th time-granularity window in the phased time series data. This represents the observation at time t in the initial time series data. Indicates the start time of the initial time series data. Indicates the time granularity parameter. This represents the aggregation function. After aggregation, time slicing is performed using the time interval parameters. It should be noted that both the standard vector data and the standard time series data have corresponding metadata, which refers to the process records of processing the initial vector data and initial time series data into the standard vector data and standard time series data.
[0030] S20: Decompose the data contract document into several atomic spatiotemporal computing units with interconnected relationships. Each atomic spatiotemporal computing unit includes a data contract, an execution operator, and a standardized artifact. Determine the data flow dependency relationship between each atomic spatiotemporal computing unit. Using the atomic spatiotemporal computing unit as the topology node and the data flow dependency relationship as the topology edge, construct the spatiotemporal computing unit topology structure. Step S20 includes: S210: Construct several initial units with interconnected relationships based on the analysis requirements, and determine whether an upstream unit exists for each initial unit based on the interconnected relationships; It should be noted that if the analysis requirements only contain time-series keywords or spatial keywords, after generating the data contract document by calling the single-type template, several progressive initial units are generated based on the data contract document. Taking spatial dominance as an example, it is decomposed according to the progressive logic of "spatial scope definition → spatial operation → result filtering". Taking time dominance as an example, it is decomposed according to the progressive logic of "time slicing → time series calculation → trend fitting". For spatiotemporal dominance, the processing logic between different processing units needs to be clarified to form a connection relationship.
[0031] S220: If the initial unit does not have an upstream unit, then a data contract is set for the initial unit based on the data contract document; if the initial unit has an upstream unit, then the output of the upstream unit is used as the data contract of the initial unit. The upstream unit refers to the previous initial unit that has a connection relationship with the current initial unit.
[0032] S230: Assign an execution operator to the initial unit through the data contract, obtain a standardized artifact based on the data contract and the execution operator, wherein the data contract, the execution operator, and the standardized artifact constitute the atomized spatiotemporal computing unit; In this embodiment, five initial units (D1~D5) with interconnected relationships are generated based on the analysis requirements. These units correspond to five sub-problems generated based on the analysis requirements: road network data preprocessing, GPS data preprocessing, spatial association between road network and GPS, congested road segment identification, and generation of congestion analysis reports. Therefore, the connection relationship of the five initial units is ((D1, D2) → D3 → D4 → D5). For initial units D1 and D2, they are parallel processing units, and their data contracts are generated based on the data contract document (D1 corresponds to at least coordinate system parameters and spatial range parameters, and D2 corresponds to at least time interval parameters and time granularity parameters). After the allocation of execution operators is completed, both D1 and D2 generate standardized artifacts. It can be understood that the standardized artifacts of D1 and D2 are retrieved from the standardized dataset. For D3, it uses the standardized artifacts generated by D1 and D2 as its data contract. For D4, it uses the standardized artifacts generated by D3 and the vehicle speed threshold parameters corresponding to the vehicle speed threshold item as its data contract. For D5, it uses the standardized artifacts generated by D4 as its data contract. It should be noted that each of the standardized products has corresponding metadata.
[0033] Understandably, the standardized artifact is the intersection of the data contract and the processing capabilities of the execution operator, ultimately generating the complete atomic spatiotemporal computing unit and data flow dependencies.
[0034] The allocation principle for the execution operator is as follows: The unit types of the atomized spatiotemporal computing unit are analyzed, including normalized units, spatiotemporal correlation units, and visualization output units; If the unit type of the atomic spatiotemporal computation unit is a standardized unit, then a code node is bound to it; if the unit type of the atomic spatiotemporal computation unit is a spatiotemporal correlation unit, then an analysis node is bound to it, and GPU resources are allocated preferentially; if the unit type of the atomic spatiotemporal computation unit is a visualization output unit, then a report node is bound to it. It should be noted that the data contract, the execution operator, and the standardized artifact constitute the spatiotemporal contractual task protocol (DOP protocol). The encapsulation structure of the atomic spatiotemporal computation unit is shown in Table 1 below: Table 1 .
[0035] It should be noted that the atomized spatiotemporal computation unit also includes metadata, and the quadruple structure of the atomized spatiotemporal computation unit is defined as follows: , in, This represents the j-th atomized spatiotemporal computational unit. This represents the data contract of the j-th atomized spatiotemporal computational unit. This represents the execution operator of the j-th atomized spatiotemporal computation unit. This represents the standardized artifact of the j-th atomized spatiotemporal computational unit. This represents the metadata of the j-th atomized spatiotemporal computation unit; The core rule of the DOP protocol is the contract transitivity consistency rule, which can be expressed mathematically as follows: , in, This represents the data contract of the (j+1)th atomized spatiotemporal computation unit. That is, the data contract of the subsequent initial unit with a connection relationship must be completely covered by the standardized artifact of the previous initial unit to ensure that spatiotemporal constraints are rigidly transmitted throughout the process.
[0036] S30: Perform consistency verification on the topology of the spatiotemporal computing unit. If the consistency verification passes, derive the input constraints of the atomized spatiotemporal computing unit based on the reverse topological order of the spatiotemporal computing unit topology, and combine all the input constraints into a spatiotemporal computing contract. The data flow dependencies include contractual dependencies and logical dependencies. Contractual dependencies refer to the requirement that the standardized artifacts of the first atomic spacetime computing unit must fully match the data contract of the second atomic spacetime computing unit in two atomic spacetime computing units that have a connection relationship. Logical dependencies refer to the order of the connection relationships between different atomic spacetime computing units.
[0037] Step S30 includes: S310: Perform global loop dependency detection on the spatiotemporal computing unit topology; Specifically, nodes with an in-degree of 0 are added to the queue, and the nodes in the queue are traversed sequentially, their outgoing edges are deleted, and the in-degree of the downstream nodes is updated. If the final number of traversed nodes is less than the total number of nodes N, it is determined that there is a cyclic dependency structure, and the verification fails.
[0038] S320: Obtain the rationality score of the spatiotemporal computing unit topology and obtain the contract compatibility score between two atomized spatiotemporal computing units with data flow dependencies; The formula for obtaining the rationality score is as follows: , in, Indicates the reasonableness score. Indicates the number of effective dependencies. This represents the total number of data flow dependencies. This indicates the number of redundant dependencies. This represents the redundancy penalty coefficient.
[0039] The formula for obtaining the contract compatibility score is: , in, This represents the contract compatibility score between upstream topology node a and downstream topology node b. This represents the data contract corresponding to the downstream topology node b. This represents the standardized artifact corresponding to the upstream topological node a. Indicates the number of elements in the set; S330: If no circular dependency structure is detected, the rationality score is greater than or equal to the first score threshold, and the contract compatibility score is greater than or equal to the second score threshold, then the consistency check is deemed to have passed. In summary, in this embodiment, an improved Kahn algorithm is constructed through steps S310 to S330. Consistency verification is achieved through the improved Kahn algorithm. To address the limitation of the traditional Kahn algorithm, which relies solely on in-degree and cannot perceive the dependencies between spatiotemporal data, a spatiotemporal contract awareness mechanism (reasonableness scoring) and a spatiotemporal priority judgment strategy (global loop dependency detection + reasonableness score + contract compatibility score) are introduced. This upgrades the single topology verification of the traditional Kahn algorithm to an integrated capability of topology verification + priority scheduling + contract verification.
[0040] S40: Based on the data contract document, spatiotemporal computing contract, and spatiotemporal computing unit topology, generate executable analysis code carrying contract verification breakpoints, execute the executable analysis code to perform full-process contract verification through the contract verification breakpoints, and output the analysis results; Step S40 includes: S410: Based on the contract-based programming paradigm, extract the constraint parameters in the data contract document and the topology of the spatiotemporal computing unit to obtain the pre-constraint conditions, and extract the product specifications of the standardized products in the atomic spatiotemporal computing unit to obtain the post-constraint conditions. S420: Convert the pre-constraints and post-constraints into assertion verification statements, and embed the assertion verification statements as contract verification breakpoints into the execution function to generate executable analysis code carrying contract verification breakpoints; Understandably, contract verification breakpoints are used to perform compliance comparisons between inputs and outputs at runtime to identify abnormal runtime states that violate constraints.
[0041] S430: Based on the execution order determined by the spatiotemporal computing unit topology, the executable analysis code is executed sequentially. If all contract verification breakpoints are passed, the full-process contract verification is deemed successful, and the analysis results are output. If a contract verification breakpoint fails, execution will be interrupted and an abnormal running status will be reported.
[0042] In this embodiment, the scheduling execution adopts a "parallel triggering + priority scheduling" strategy, which allocates computing resources based on the contract attributes and resource requirements of each unit. The specific execution is as follows: ① 0-8 minutes: D1 and D2 have no dependencies and no contract conflicts, and are executed in parallel; high CPU resources (CPU utilization 70%) are allocated to D1 to adapt to the needs of vector data topology processing; high IO resources (CPU utilization 65%) are allocated to D2 to adapt to time series data slicing, and both units complete preprocessing according to the contract requirements.
[0043] ② 8-15 minutes: D3 depends on the execution results of D1 and D2, triggers execution; allocates GPU resources to it (GPU utilization 60%), accelerates GeoPandas spatial connection operations, ensures that the spatiotemporal association of road network and GPS data is completed as required by the contract, and generates road network artifacts containing vehicle speed.
[0044] ③ 15-22 minutes: D4 depends on the execution result of D3 and triggers execution; continue to allocate GPU resources (GPU utilization 50%), and based on the contractually agreed vehicle speed threshold parameters (vehicle speed <20km / h), use NetworkX tools to filter and mark congested road sections to complete congestion identification.
[0045] ④ 22-30 minutes: D5 depends on the execution result of D4 and triggers execution; allocates regular CPU resources (CPU utilization 40%), draws interactive maps and statistical charts using Matplotlib according to the contract visualization requirements, and generates a PDF analysis report in combination with ReportLab, clearly marking all spatiotemporal contract attributes in the report.
[0046] Output: Execution logs for each atomic unit (including execution time, resource utilization, and contract verification results). Through the consistency verification, efficient scheduling and fault-tolerant execution with contract awareness are achieved, ensuring that tasks are completed in an orderly manner according to contract requirements. It should be noted that if the consistency verification fails, a retry mechanism is triggered. Specifically, the node with the error is selected as the node to be checked, and the dependent nodes of the node to be checked (the preceding nodes with data flow dependencies on the node to be checked) are obtained. It is determined whether the standardized artifacts of the dependent nodes match the data contract of the node to be checked. If they do not match, the data contract of the node to be checked is adjusted, or the standardized artifacts of the dependent nodes are re-obtained. Furthermore, by combining full-process contract verification, a closed loop from "static design" to "dynamic verification" is achieved. By implementing penetrating monitoring of inputs, intermediate results, and outputs during execution, it has the ability to identify abnormal operating states in real time, effectively ensuring data consistency and result reliability throughout the spatiotemporal analysis process, and significantly reducing the cost of manual error checking.
[0047] The data contract document is generated based on the analytical requirements and the multi-source spatiotemporal data, providing a spatiotemporal contract specification for subsequent data analysis and laying the foundation for the contractualization of spatiotemporal requirements. By standardizing the multi-source spatiotemporal data through the data contract document, automated spatiotemporal attribute alignment between heterogeneous data is achieved, improving the adaptability between the data contract document and the standardized dataset. By decomposing the data contract document into atomic spatiotemporal computing units, the DOP protocol, composed of its data contract, execution operators, and standardized artifacts, provides unified protocol support for task decomposition. This enables the construction of a structurally consistent and constraint-traceable multi-source spatiotemporal data analysis process, effectively alleviating the problems of inconsistent multi-source data formats, chaotic spatiotemporal constraints, and poor scheduling coordination, and possesses good generalization and scalability.
[0048] Preferably, the method further includes: S50: Determine if the user has a contract modification requirement. If a contract modification requirement exists, regenerate the analysis results. After users view the analysis report, they can change the spatiotemporal constraint parameters based on their needs, such as expanding the time range to 6:00~10:00, which forms a contract modification requirement.
[0049] Step S50 includes: S510: Based on the spatiotemporal computing unit topology, the atomic spatiotemporal computing units corresponding to the contract modification requirements are selected as marked units, and the remaining atomic spatiotemporal computing units are selected as reserved units; Taking the modification of the time range as an example, the affected cells (marked cells) are D2~D4, and the unaffected cells (reserved cells) are D1.
[0050] S520: Add the marked unit to the backtracking execution queue to generate a replacement unit, and regenerate the analysis results based on the replacement unit and the retained unit; By acquiring the contract modification requirements and individually modifying the marked units to form replacement units, efficient iteration based on the contract modification requirements is ensured, reducing iteration costs and improving analysis efficiency. Understandably, during the generation of replacement units, standardized artifacts are regenerated through backtracking and recalculation. Furthermore, the replacement units and retained units can reconstruct the spatiotemporal computing unit topology, and through the same processing steps, executable analysis code carrying contract verification breakpoints is generated, thereby outputting the analysis results.
[0051] Furthermore, the method also includes: S60: Add the process for obtaining the analysis results to the historical task execution library, and dynamically optimize the acquisition strategy of the data contract document and the atomic spatiotemporal computing unit based on the historical task execution library.
[0052] Dynamic optimization revolves around three core objects: contract extraction rules, operator matching strategies, and scheduling priorities. It drives incremental updates and batch iterations of the mapping relationship library through the historical task execution library, as detailed below: Historical task feature collection: Structured logs are collected for each completed spatiotemporal analysis task. The collection dimensions include: user requirement type, data contract items, decomposition results of atomic spatiotemporal computing units, binding relationship between execution operators and data contracts, actual scheduling time, contract matching success rate, unit execution success rate, user retrospective modification records and contract adjustment parameters. The above information is uniformly quantified into requirement feature vectors, contract feature vectors, operator feature vectors, and effect evaluation vectors to form a trainable historical task sample library. Initialization and rule encoding of the mapping relationship library: Based on the initial task samples, a four-level mapping relationship library is constructed, consisting of "requirement type - unit contract constraint - optimal operator set - resource scheduling priority". Each mapping entry in the library is bound to the constraint threshold of the DOP protocol, including: coordinate system compatibility priority, resolution allowable error range, time granularity matching weight, operator execution time threshold, and GPU resource triggering conditions. Association rule mining based on historical performance: Multi-dimensional association statistics and rule scoring are performed on the sample library to calculate the success rate of operator execution for different combinations of contract constraints under the same type of requirement; the average scheduling time and resource consumption of different operators under different spatiotemporal data scales are calculated; the recall and precision of contract item extraction are calculated to identify high-frequency missing or misjudged constraints; and the optimal DOP protocol configuration template is automatically selected for each type of requirement through performance-weighted ranking, including which contract items to extract first, which operators to bind first, and which resources to allocate first. The dynamic update mechanism of the mapping relationship library is divided into two levels to achieve continuous self-evolution: First, incremental real-time update. After each task is completed, the new sample is immediately merged into the sample library, and the rule weights of the corresponding requirement type in the mapping relationship library are fine-tuned; Second, batch global optimization. When the cumulative number of tasks reaches a preset threshold, the entire sample is iterated as a whole. If the contract extraction accuracy is lower than 95%, the recognition weight of key spatiotemporal contract items is automatically increased, and the contract generation logic is optimized. Self-evolving closed-loop feedback: The updated mapping relationship library directly affects the entire process of subsequent atomic decomposition and operator scheduling. The new execution results flow back to the sample library, forming a closed-loop self-evolving system of "task execution → data collection → rule mining → mapping update → better execution". In continuous operation, it automatically improves the adaptation accuracy and overall execution efficiency of the DOP protocol.
[0053] By performing dynamic optimization, the analysis method can learn and evolve on its own. Based on a clear algorithm process, the mapping relationship library driven by historical task data is dynamically updated. Combined with standardized preprocessing and backtracking functions, the efficiency and accuracy of data analysis are greatly improved, the cost of manual intervention is reduced, and it has high interpretability and engineering application value.
[0054] The method in this embodiment is compared with the traditional GIS analysis tool (ArcGIS + Excel) in the task of "Analysis of Congested Road Sections during Morning Rush Hour in a Certain City" for core indicators. The comparison results are shown in Table 2 below: Table 2 , The comparison results show that, compared with traditional GIS tools, this embodiment has significantly improved in key indicators such as total task time, requirement decomposition efficiency, contract error rate, iteration time, and manual intervention rate, which fully demonstrates the effectiveness and superiority of the method in this embodiment. The core benefit is the full-process constraints of the DOP protocol, the collaborative work of the modular architecture, and the contract-aware scheduling mechanism, which effectively reduces manual intervention and invalid calculations, and improves analysis accuracy and execution efficiency.
[0055] The method in this embodiment has good universality and scalability. By adjusting the DOP protocol parameters and operator mapping rules, it can be adapted to different multi-source spatiotemporal data analysis scenarios without modifying the core architecture. The specific extended scenarios are as follows: (1) Environmental monitoring scenario: The demand type is "time-series dominant", and the core analysis target is "monthly NDVI index change trend analysis of a certain area"; the atomic unit decomposition is performed according to "time slice (monthly contract) → NDVI calculation (30m resolution contract) → trend fitting"; the execution operator uses R-sp package to process time series data, and the product metadata is labeled with monthly time range and 30m spatial resolution to adapt to the time series analysis requirements of environmental monitoring; (2) Land space scenario: The demand type is "space-dominant", and the core analysis target is "land space ownership conflict detection of a certain township"; the atomic unit decomposition is performed according to "spatial range definition (township contract) → ownership data overlay (same coordinate system contract) → conflict detection"; the execution operator uses ArcGIS. Engine processes topological relationships, and product metadata labels the spatial range of townships and the unified coordinate system (EPSG:4326) to adapt to the spatial constraint requirements of land space analysis; (3) Data type expansion: such as LiDAR point cloud data, POI data, etc.; For LiDAR point cloud data, the PDAL tool is added to process the spatial resolution contract to realize the standardized resampling of point cloud data; For POI data, the GeoPandas tool is called to associate the spatial range contract to realize the spatial matching of POI data and road network data. The DOP protocol adaptation of new data types can be completed through the configuration file, which has strong scalability.
[0056] Please see Figure 2 The second embodiment of the present invention provides a spatiotemporal contract-driven multi-source heterogeneous data analysis system. This system is applied to the spatiotemporal contract-driven multi-source heterogeneous data analysis method described in the above embodiments, and will not be repeated hereafter. As used below, the terms "module," "unit," "subunit," etc., can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0057] The system includes: The parsing module 10 is used to obtain a data contract document including spatiotemporal constraint parameters based on the multi-source spatiotemporal data and analysis requirements input by the user, and to perform standardization processing on the multi-source spatiotemporal data based on the data contract document to obtain a standardized dataset corresponding to the data contract document. The parsing module 10 includes: The first unit is used to extract the core fields and spatiotemporal attributes of the multi-source spatiotemporal data, and compare the core fields and spatiotemporal attributes with the analysis requirements to determine whether data adjustment is needed for the multi-source spatiotemporal data. The second unit is used to identify whether there are temporal keywords and spatial keywords in the analysis requirements based on a preset spatiotemporal dictionary if no data adjustment is required for the multi-source spatiotemporal data. The second unit is specifically used to perform word segmentation on the analysis requirements in order to obtain the keywords to be identified; Set keyword weights for preset keywords and obtain the edit distance between the keyword to be identified and the preset keywords. Based on the keyword weights and edit distances, obtain the word similarity between the keyword to be identified and the preset keywords. The word similarity is compared with the similarity threshold. If the word similarity is greater than the similarity threshold, the keyword to be identified is selected as a temporal keyword or a spatial keyword. The third unit is used to call a spatiotemporal fusion contract template that includes spatiotemporal constraints and custom constraints if there are temporal and spatial keywords in the analysis requirements. The spatiotemporal constraints and custom constraints are filled with preset parameter values and the analysis requirements to obtain a data contract document that includes spatiotemporal constraint parameters. The fourth unit is used to perform coordinate system transformation on the initial vector data through the coordinate system parameters to obtain stage vector data, and to perform data aggregation on the initial time series data through the time granularity parameters to obtain stage time series data. The fifth unit is used to spatially truncate the stage vector data based on the spatial range parameter to obtain standard vector data, to perform time slicing on the stage time series data based on the time interval parameter to obtain standard time series data, and to combine the standard vector data and the standard time series data into a standardized dataset corresponding to the data contract document. The encapsulation module 20 is used to decompose the data contract document into several atomic spatiotemporal computing units with interconnected relationships. The atomic spatiotemporal computing unit includes a data contract, an execution operator, and a standardized artifact. The module determines the data flow dependency relationship between each atomic spatiotemporal computing unit and constructs a spatiotemporal computing unit topology structure with the atomic spatiotemporal computing unit as the topology node and the data flow dependency relationship as the topology edge. The packaging module 20 includes: The sixth unit is used to construct several initial units with interconnected relationships based on the analysis requirements, and to determine whether an upstream unit exists in the initial unit based on the interconnected relationships; The seventh unit is used to set a data contract for the initial unit based on the data contract document if the initial unit does not have an upstream unit, and to use the output of the upstream unit as the data contract of the initial unit if the initial unit has an upstream unit. The eighth unit is used to allocate execution operators to the initial unit through the data contract, and to obtain standardized artifacts based on the data contract and the execution operators. The data contract, the execution operators, and the standardized artifacts constitute the atomized spatiotemporal computing unit. The verification module 30 is used to perform consistency verification on the topology of the spatiotemporal computing unit. If the consistency verification passes, the input constraints of the atomized spatiotemporal computing unit are derived based on the reverse topological order of the spatiotemporal computing unit topology, and all the input constraints are combined into a spatiotemporal computing contract. The verification module 30 includes: The ninth unit is used to perform global loop dependency detection on the spatiotemporal computing unit topology. The tenth unit is used to obtain the rationality score of the spatiotemporal computing unit topology and to obtain the contract compatibility score between two atomic spatiotemporal computing units that have data flow dependencies. Unit 11 is used to determine that the consistency check passes if no circular dependency structure is detected, the reasonableness score is greater than or equal to the first score threshold, and the contract compatibility score is greater than or equal to the second score threshold. The execution module 40 is used to generate executable analysis code carrying contract verification breakpoints based on the data contract document, the spatiotemporal computing contract, and the topology of the spatiotemporal computing unit, execute the executable analysis code to perform full-process contract verification through the contract verification breakpoints, and output the analysis results. The execution module 40 includes: Unit 12 is used to extract constraint parameters from data contract documents and spatiotemporal computing unit topology based on contract-based programming paradigm to obtain pre-constraint conditions, and to extract product specifications of standardized products in atomic spatiotemporal computing units to obtain post-constraint conditions. Unit 13 is used to convert pre-constraints and post-constraints into assertion verification statements, and embed the assertion verification statements as contract verification breakpoints into the execution function to generate executable analysis code carrying contract verification breakpoints. The fourteenth unit is used to execute the executable analysis code sequentially based on the execution order determined by the spatiotemporal computing unit topology. If all contract verification breakpoints are passed, the full-process contract verification is determined to be successful, and the analysis results are output. Preferably, the system further includes: The adjustment module 50 is used to determine whether the user has a contract modification requirement. If a contract modification requirement exists, the analysis results are regenerated. The adjustment module 50 includes: The fifteenth unit is used to select the atomic spatiotemporal computing units corresponding to the contract modification requirements as marked units based on the spatiotemporal computing unit topology, and to select the remaining atomic spatiotemporal computing units as reserved units. The sixteenth unit is used to add the marked unit to the backtracking execution queue to generate a replacement unit, and regenerate the analysis results based on the replacement unit and the retained unit; Furthermore, the system also includes: The optimization module 60 is used to add the process of obtaining the analysis results to the historical task execution library, and dynamically optimize the acquisition strategy of the data contract document and the atomic spatiotemporal computing unit based on the historical task execution library.
[0058] The present invention also provides a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the spatiotemporal contract-driven multi-source heterogeneous data analysis method as described in the above technical solutions.
[0059] The present invention also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the spatiotemporal contract-driven multi-source heterogeneous data analysis method as described in the above technical solution.
[0060] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0061] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A spatiotemporal contract-driven method for analyzing multi-source heterogeneous data, characterized in that, Includes the following steps: A data contract document, including spatiotemporal constraint parameters, is obtained by using multi-source spatiotemporal data and analysis requirements input by the user. Based on the data contract document, the multi-source spatiotemporal data is standardized to obtain a standardized dataset corresponding to the data contract document. The data contract document is decomposed into several atomic spatiotemporal computing units with interconnected relationships. Each atomic spatiotemporal computing unit includes a data contract, an execution operator, and a standardized artifact. The data flow dependencies between each atomic spatiotemporal computing unit are determined. Using the atomic spatiotemporal computing unit as the topology node and the data flow dependencies as the topology edge, a spatiotemporal computing unit topology structure is constructed. A consistency check is performed on the topology of the spatiotemporal computing unit. If the consistency check passes, the input constraints of the atomized spatiotemporal computing unit are derived based on the reverse topological order of the spatiotemporal computing unit topology, and all the input constraints are combined into a spatiotemporal computing contract. Based on the data contract document, the spatiotemporal computing contract, and the topology of the spatiotemporal computing unit, executable analysis code carrying contract verification breakpoints is generated. The executable analysis code is executed to perform full-process contract verification through the contract verification breakpoints and output the analysis results.
2. The spatiotemporal contract-driven multi-source heterogeneous data analysis method according to claim 1, characterized in that, The steps of obtaining a data contract document including spatiotemporal constraint parameters based on multi-source spatiotemporal data and analysis requirements input by the user include: Extract the core fields and spatiotemporal attributes of the multi-source spatiotemporal data, and compare the core fields and spatiotemporal attributes with the analysis requirements to determine whether data adjustment is needed for the multi-source spatiotemporal data. If no data adjustment is required for the multi-source spatiotemporal data, then the presence of temporal keywords and spatial keywords in the analysis requirements is identified based on a preset spatiotemporal dictionary. If the analysis requirements contain temporal and spatial keywords, then a spatiotemporal fusion contract template including spatiotemporal constraints and custom constraints is invoked. The spatiotemporal constraints and custom constraints are filled with preset parameter values and the analysis requirements to obtain a data contract document including spatiotemporal constraint parameters.
3. The spatiotemporal contract-driven multi-source heterogeneous data analysis method according to claim 2, characterized in that, The preset spatiotemporal dictionary includes several preset keywords. The step of identifying whether temporal and spatial keywords exist in the analysis requirements based on the preset spatiotemporal dictionary includes: The analysis requirements are segmented to obtain the keywords to be identified. Set keyword weights for preset keywords and obtain the edit distance between the keyword to be identified and the preset keywords. Based on the keyword weights and edit distances, obtain the word similarity between the keyword to be identified and the preset keywords. The word similarity is compared with the similarity threshold. If the word similarity is greater than the similarity threshold, the keyword to be identified is selected as a temporal keyword or a spatial keyword.
4. The spatiotemporal contract-driven multi-source heterogeneous data analysis method according to claim 1, characterized in that, The spatiotemporal constraint parameters include coordinate system parameters, spatial range parameters, time interval parameters, and time granularity parameters. The multi-source spatiotemporal data includes initial vector data and initial time-series data. The step of standardizing the multi-source spatiotemporal data based on the data contract document to obtain a standardized dataset corresponding to the data contract document includes: The initial vector data is transformed using the coordinate system parameters to obtain stage vector data, and the initial time series data is aggregated using the time granularity parameters to obtain stage time series data. Based on the spatial range parameter, the stage vector data is spatially clipped to obtain standard vector data. Based on the time interval parameter, the stage time series data is time-sliced to obtain standard time series data. The standard vector data and the standard time series data are combined to form a standardized dataset corresponding to the data contract document.
5. The spatiotemporal contract-driven multi-source heterogeneous data analysis method according to claim 4, characterized in that, The step of decomposing the data contract document into several interconnected atomic spatiotemporal computation units, wherein each atomic spatiotemporal computation unit includes a data contract, an execution operator, and a standardized artifact, includes: Based on the analysis requirements, several initial units with interconnected relationships are constructed, and it is determined whether an upstream unit exists for each initial unit based on the interconnected relationships. If the initial unit does not have an upstream unit, a data contract is set for the initial unit based on the data contract document; if the initial unit has an upstream unit, the output of the upstream unit is used as the data contract of the initial unit. An execution operator is assigned to the initial unit through the data contract, and a standardized artifact is obtained based on the data contract and the execution operator. The data contract, the execution operator, and the standardized artifact constitute the atomized spatiotemporal computing unit.
6. The spatiotemporal contract-driven multi-source heterogeneous data analysis method according to claim 1, characterized in that, The steps for performing consistency verification on the spatiotemporal computing unit topology include: Perform global loop dependency detection on the spatiotemporal computing unit topology; Obtain the rationality score of the spatiotemporal computing unit topology and obtain the contract compatibility score between two atomic spatiotemporal computing units with data flow dependencies; If no circular dependency structure is detected, the reasonableness score is greater than or equal to the first score threshold, and the contract compatibility score is greater than or equal to the second score threshold, then the consistency check is considered to have passed.
7. The spatiotemporal contract-driven multi-source heterogeneous data analysis method according to claim 6, characterized in that, The formula for obtaining the rationality score is as follows: , in, Indicates the reasonableness score. Indicates the number of effective dependencies. This represents the total number of data flow dependencies. This indicates the number of redundant dependencies. This represents the redundancy penalty coefficient.
8. The spatiotemporal contract-driven multi-source heterogeneous data analysis method according to claim 1, characterized in that, The steps of generating executable analysis code carrying contract verification breakpoints based on the data contract document, spatiotemporal computing contract, and spatiotemporal computing unit topology, executing the executable analysis code to perform full-process contract verification through the contract verification breakpoints, and outputting analysis results include: Based on the contract-based programming paradigm, constraint parameters in the data contract document and the spatiotemporal computing unit topology are extracted to obtain the pre-constraint conditions, and the product specifications of the standardized products in the atomic spatiotemporal computing unit are extracted to obtain the post-constraint conditions. The pre-constraints and post-constraints are converted into assertion verification statements, and the assertion verification statements are embedded as contract verification breakpoints into the execution function to generate executable analysis code carrying contract verification breakpoints. Based on the execution order determined by the spatiotemporal computing unit topology, the executable analysis code is executed sequentially. If all contract verification breakpoints are passed, the full-process contract verification is deemed successful, and the analysis results are output.
9. The spatiotemporal contract-driven multi-source heterogeneous data analysis method according to claim 1, characterized in that, After the steps of generating executable analysis code carrying contract verification breakpoints based on the data contract document, spatiotemporal computing contract, and spatiotemporal computing unit topology, executing the executable analysis code to perform full-process contract verification through the contract verification breakpoints, and outputting the analysis results, the method further includes: Determine if the user needs to modify the contract. If so, regenerate the analysis results.
10. A spatiotemporal contract-driven multi-source heterogeneous data analysis system, applied to the spatiotemporal contract-driven multi-source heterogeneous data analysis method as described in any one of claims 1 to 9, characterized in that, The system includes: The parsing module is used to obtain a data contract document including spatiotemporal constraint parameters based on the multi-source spatiotemporal data and analysis requirements input by the user, and to perform standardization processing on the multi-source spatiotemporal data based on the data contract document to obtain a standardized dataset corresponding to the data contract document. The encapsulation module is used to decompose the data contract document into several atomic spatiotemporal computing units with interconnected relationships. The atomic spatiotemporal computing unit includes a data contract, an execution operator, and a standardized artifact. The module determines the data flow dependency relationship between each atomic spatiotemporal computing unit and constructs a spatiotemporal computing unit topology structure with the atomic spatiotemporal computing unit as the topology node and the data flow dependency relationship as the topology edge. The verification module is used to perform consistency verification on the topology of the spatiotemporal computing unit. If the consistency verification passes, the input constraints of the atomic spatiotemporal computing unit are derived based on the reverse topological order of the spatiotemporal computing unit topology, and all the input constraints are combined into a spatiotemporal computing contract. The execution module is used to generate executable analysis code carrying contract verification breakpoints based on the data contract document, spatiotemporal computing contract, and spatiotemporal computing unit topology, execute the executable analysis code to perform full-process contract verification through the contract verification breakpoints, and output the analysis results.