Operator arrangement method and device for spatio-temporal data analysis and calculation and storage medium

By analyzing the operator orchestration graph and partitioning the distributed data set, dividing it into multiple subtasks and allocating it to the computing nodes in the distributed cluster for parallel execution, the problem that a single operator node's computing task cannot be distributed to multiple servers is solved, and the efficiency of large-scale data volume calculation is improved.

CN119938283AActive Publication Date: 2025-05-06SHENZHEN SMARTCITY TECH DEV GRP CO LTD

Patent Information

Application Number
CN202510438083.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art cannot distribute the computing tasks of a single operator node to multiple servers, resulting in limited performance of a single server and inability to support large-scale data calculations.

Method used

By analyzing the operator arrangement graph, determining the operator to be executed and its dependencies are generated, and executable files are divided into multiple subtasks according to the partition of the distributed data set of input data, and the tasks are allocated to the computing nodes in the distributed cluster for parallel execution.

Benefits of technology

It realizes that a single operator node runs calculations on multiple servers, improves the spatial data calculation efficiency under large data volume, and supports large-scale data calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938283A_ABST
    Figure CN119938283A_ABST
Patent Text Reader

Abstract

The invention discloses an operator arrangement method and device for spatio-temporal data analysis and calculation and a storage medium, and relates to the technical field of data processing.The method comprises the steps that after an operator arrangement diagram is generated, the operator arrangement diagram is analyzed, and to-be-executed operators and the dependency relationship between the to-be-executed operators are determined; generating an executable file according to the code snippets corresponding to the operator to be executed and the dependency relationship; dividing a running task corresponding to the executable file into a plurality of sub-tasks according to partitions of a distributed data set of the input data; and distributing the sub-tasks to respective corresponding computing nodes in the distributed cluster, performing parallel execution, and obtaining an execution result after the execution of each sub-task is completed. According to the method, the subtasks which run in parallel are created and allocated to the computing nodes for computing according to the partition condition of the distributed data set of the input data, so that a single operator node runs and computes on a plurality of servers, and the computing efficiency of spatial data under a large data volume is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to an operator arrangement method, device and storage medium for spatiotemporal data analysis and calculation. Background Art

[0002] In the operator orchestration system, running operators involves the process orchestration and scheduling execution of operator nodes. Operator nodes represent specific operator operations, while process orchestration defines the execution order and relationship between these operator nodes. The scheduling engine allocates operator nodes to different servers or resources for operation according to the orchestration process.

[0003] Usually, after the process design and arrangement of the operator nodes are completed and the operation phase is entered, the scheduling engine will assign each operator node to the server one by one according to the arranged process. One or more operator nodes can run on a server at the same time. However, it is not possible to distribute the computing tasks of a single operator node to multiple servers. Due to the limited overall performance of a single server, a single operator node can only run on one server, which limits the operating efficiency of the operator node and cannot support large-scale data volume calculations.

[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The main purpose of this application is to provide an operator orchestration method, device and storage medium for spatiotemporal data analysis and calculation, aiming to solve the technical problem of how to improve the operating efficiency of a single operator node to support large-scale data calculation.

[0006] To achieve the above purpose, the present application proposes an operator arrangement method for spatiotemporal data analysis and calculation, and the operator arrangement method for spatiotemporal data analysis and calculation includes: After generating the operator orchestration graph, the operator orchestration graph is parsed to determine the operators to be executed and the dependencies between the operators to be executed; Generate an executable file according to the code snippet corresponding to the operator to be executed and the dependency relationship; Dividing the running task corresponding to the executable file into a plurality of subtasks according to the partition of the distributed data set of the input data; The subtasks are assigned to the corresponding computing nodes in the distributed cluster for parallel execution, and the execution results are obtained after the execution of each subtask is completed.

[0007] In one embodiment, the operator orchestration graph is a directed acyclic graph, and the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators to be executed includes: Determine a starting node in the operator orchestration graph; Based on the depth-first search algorithm, traverse all successor nodes of the starting node; The dependency relationship is obtained according to the traversal order and the direction of the edges in the directed graph.

[0008] In one embodiment, after generating the operator orchestration graph, before the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators to be executed, the method further includes: If it is detected that the user drags a first operator onto the canvas, an identifier of the first operator is obtained; Acquire an associated operator of the first operator according to the identifier; Obtaining a comprehensive prediction score of the association operator according to the usage frequency, association degree and corresponding weight of the association operator; Arrange the associated operators in descending order according to the comprehensive prediction scores, and store them in a recommended operator list; The recommended operator list is displayed in the recommendation area.

[0009] In one embodiment, before the step of obtaining an identifier of the first operator if it is detected that the user drags the first operator onto the canvas, the step further includes: Establishing an operator association rule base to store successor operators of each of the first operators and corresponding association degrees; When it is detected that the user selects the prediction operator of the recommendation area, the association degree between the prediction operator and the first operator is updated.

[0010] In one embodiment, the step of dividing the running task corresponding to the executable file into a plurality of subtasks according to the partition of the distributed data set of the input data comprises: Based on the master node of the distributed cluster, the executable file is divided into task stages to obtain each task stage and the execution order of the task stages; According to the execution order and the data source path of the executable file, the distributed data set of the input data is created and partitioned, and the partitioning strategy is the space dimension and / or the time dimension; The subtasks to be executed in parallel are generated based on the partitioning result and the cluster configuration.

[0011] In one embodiment, before the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators, the following step is also included: Determine a deployment mode, and build the distributed computing cluster according to the deployment mode; Create and encapsulate operators, and save the encapsulated operators to the operator library.

[0012] In one embodiment, before the step of assigning the subtask to the corresponding computing node for execution and obtaining the execution result, the method further includes: Create a monitoring chain list, and store the running tasks in the monitoring chain list; Performing a traversal operation on the monitoring linked list to obtain the running status of the running task; When the running status is running completed, task completion information is generated, and the corresponding running task is deleted from the monitoring linked list.

[0013] In one embodiment, the step of allocating the subtasks to the corresponding computing nodes in the distributed cluster for parallel execution, and obtaining the execution results after the execution of each subtask is completed, includes: Sending the subtask to the corresponding computing node according to a task distribution algorithm; After receiving the task, the computing node calls the operator to be executed to process the data and outputs the execution result.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes an operator orchestration device for spatiotemporal data analysis and calculation, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the operator orchestration method for spatiotemporal data analysis and calculation as described above.

[0015] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the operator arrangement method for spatiotemporal data analysis and calculation are implemented as described above.

[0016] The present application provides an operator orchestration method for spatiotemporal data analysis and calculation, which parses the operator orchestration graph to determine the operators to be executed and the dependencies between them, and generates an executable file. By encapsulating the operator orchestration logic into a whole, it is convenient for subsequent unified management and scheduling. According to the partitioning of the distributed data set of the input data, subtasks running in parallel are created and assigned to computing nodes for calculation. By splitting the operator task into multiple subtasks and assigning the subtasks to multiple computing nodes in the cluster, a single operator node can be used to run calculations on multiple servers, thereby improving the computing efficiency of spatial data under large data volumes. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 A flowchart diagram of the first embodiment of the operator arrangement method for spatiotemporal data analysis and calculation provided in this application; Figure 2 A schematic diagram of the overall process provided for the operator arrangement method for spatiotemporal data analysis and calculation in this application; Figure 3 Functional architecture diagram provided for the operator arrangement method of spatiotemporal data analysis and calculation in this application; Figure 4 A flowchart diagram of the operator arrangement method for spatiotemporal data analysis and calculation provided in Example 3 of the present application; Figure 5 A schematic diagram of operator arrangement provided for the operator arrangement method for spatiotemporal data analysis and calculation in this application; Figure 6 Flow chart of the operator arrangement method for spatiotemporal data analysis and calculation provided in Embodiments 4 and 5 of the present application; Figure 7 A flowchart diagram of the sixth embodiment of the operator arrangement method for spatiotemporal data analysis and calculation provided in this application; Figure 8 A timing diagram provided for the operator arrangement method for spatiotemporal data analysis and calculation in this application; Fig. 9 Schematic diagram of the structure of the hardware operating environment involved in the operator arrangement method for spatiotemporal data analysis and calculation in the embodiment of the present application.

[0020] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0022] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0023] The main solution of the embodiment of the present application is: after generating the operator orchestration graph, the operator orchestration graph is parsed to determine the operators to be executed and the dependencies between the operators to be executed; an executable file is generated according to the code snippets corresponding to the operators to be executed and the dependencies; according to the partition of the distributed data set of the input data, the running task corresponding to the executable file is divided into multiple subtasks; the subtasks are assigned to the corresponding computing nodes in the distributed cluster for parallel execution, and the execution results are obtained when the execution of each of the subtasks is completed.

[0024] In the operator orchestration system, running operators involves the process orchestration and scheduling execution of operator nodes. Operator nodes represent specific operator operations, while process orchestration defines the execution order and relationship between these operator nodes. The scheduling engine allocates operator nodes to different servers or resources for operation according to the orchestration process.

[0025] In the prior art, after completing the process design and orchestration of the operator node and entering the operation stage, the scheduling engine assigns each operator node to the server one by one according to the orchestrated process. One or more operator nodes can run on a server at the same time. However, it is not possible to distribute the computing tasks of a single operator node to multiple servers. Due to the limited overall performance of a single server, a single operator node can only run on one server, which limits the operating efficiency of the operator node and cannot support large-scale data calculations. In addition, in the existing operator orchestration system, only conventional string parameter passing is supported between operators, and file and two-dimensional and three-dimensional spatiotemporal data type parameter passing cannot be supported.

[0026] In order to solve the above problems, the present application provides an operator orchestration method for spatiotemporal data analysis and calculation, which parses the operator orchestration graph to determine the operators to be executed and the dependencies between them, and generates an executable file. By encapsulating the operator orchestration logic into a whole, it is convenient for subsequent unified management and scheduling. According to the partitioning of the distributed data set of the input data, subtasks running in parallel are created and assigned to computing nodes for calculation. By splitting the operator task into multiple subtasks and assigning the subtasks to multiple computing nodes in the cluster, a single operator node can be used to run calculations on multiple servers, thereby improving the computing efficiency of spatial data under large data volumes.

[0027] It should be noted that the execution subject of this embodiment can be a computing service device with network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a device, etc. that can realize the above functions. The following takes the operator arrangement device of spatiotemporal data analysis calculation as an example to illustrate this embodiment and the following embodiments.

[0028] Based on this, the embodiment of the present application provides an operator arrangement method for spatiotemporal data analysis and calculation, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the operator arrangement method for spatiotemporal data analysis and calculation in the present application.

[0029] In this embodiment, the operator arrangement method for spatiotemporal data analysis and calculation is applied to an operator arrangement device for spatiotemporal data analysis and calculation, and the method includes steps S100 to S400: Step S100: after generating the operator orchestration graph, the operator orchestration graph is parsed to determine the operators to be executed and the dependencies between the operators to be executed.

[0030] It should be noted that an operator can be regarded as an operation or function that receives one or more inputs, processes these inputs according to certain rules or algorithms, and finally outputs one or more results. Operator orchestration refers to the reasonable arrangement of the execution order of operators, the allocation of computing resources, and the optimization of data transmission between storage and computing units.

[0031] In this embodiment, when performing an operator orchestration task, a series of operators are defined to process data. And the dependencies between operators are determined, that is, which operators' outputs are the inputs of other operators. Use graph data structures (such as adjacency lists, adjacency matrices, etc.) or graph libraries to generate operator orchestration graphs. Traverse the operator orchestration graph using graph traversal algorithms such as depth-first search (DFS) or breadth-first search (BFS). During the traversal process, record all operators that need to be executed. Create a dependency table for each operator, listing all its predecessor operators. Determine the dependencies between operators based on the order of traversal and the direction of the edges.

[0032] In a feasible implementation, step S100 may include the following steps: Determine a starting node in the operator orchestration graph.

[0033] Based on the depth-first search algorithm, all successor nodes of the starting node are traversed.

[0034] The dependency relationship is obtained according to the traversal order and the direction of the edges in the directed graph.

[0035] In this implementation, the operator orchestration graph is a directed acyclic graph. A directed acyclic graph (DAG) consists of vertices and edges. A directed acyclic graph is a directed graph in which each edge has a clear direction and the entire graph is acyclic. That is, there is no path in the graph that can start from a vertex and return to the vertex after passing through a series of edges. Each edge points from one vertex to another, indicating a unidirectional relationship or dependency.

[0036] In this embodiment, after the operator arrangement is completed according to the directed acyclic graph specification, the scheduling engine parses the operator arrangement based on the depth-first search algorithm to parse out the key information such as the dependency relationship and execution order of the operator arrangement. In a directed acyclic graph, nodes represent operators and directed edges represent the dependency relationship between operators. When parsing according to the depth-first search algorithm, first, create a stack to store the nodes to be visited. Create a set or array to record the nodes that have been visited to prevent repeated visits. Secondly, traverse and find all operators without predecessor nodes (i.e., nodes with an in-degree of 0) from the directed acyclic graph, and use these operators as the starting nodes of the depth-first search. Push the starting node into the stack and mark it as visited. When the stack is not empty, execute the following loop: pop a node from the top of the stack and record it as the current node. Process the current node and record its information as a certain operator (such as operator ID, type, etc.). Traverse all successor nodes of the current node, push each unvisited successor node into the stack, and mark the current node as its predecessor node. Mark the current node as fully visited (i.e., all its successor nodes have been visited by depth-first search or are determined not to need to be visited). During the depth-first search traversal, a data structure (such as a hash table, adjacency list, or graph database) is used to record the dependencies between operators. When the stack is empty and all nodes have been visited, the depth-first search traversal ends. At this point, the dependency table or graph contains the dependencies between all operators.

[0037] In another feasible implementation, the nodes can be traversed layer by layer according to breadth-first search to obtain the dependency relationship between nodes. When obtaining the operator dependency, start from a starting operator and expand outward layer by layer until all reachable operators are traversed. The dependency relationship between operators is determined according to the traversal order and the direction of the edge.

[0038] Exemplarily, in the process of analyzing whether two or more spatial objects intersect and finding the operator arrangement of their intersecting parts, the following steps are specifically included: Step (1), enumerating the starting point of the DAG operator arrangement, that is, the operator without a parent node, and obtaining the starting point operators: "Read Shapefile Operator" and "Read PostGIS Operator". Step (2), using the DFS algorithm to traverse from the starting point operator "Read Shapefile Operator", whose child node is the "Spatial Intersection Operator", and after a complete traversal, the result is: "Read Shapefile Operator" -> "Spatial Intersection Operator" -> "Write Shapefile Operator". Step (3), using the DFS algorithm to traverse from another starting point operator "Read PostGIS Operator", whose child node is the "Spatial Intersection Operator", and after a complete traversal, the result is: "Read PostGIS Operator" -> "Spatial Intersection Operator" -> "Write Shapefile Operator". Step (4): According to the above traversal results, the execution order and dependency of the DAG operator arrangement are parsed. The two starting operators "Read Shapefile Operator" and "Read PostGIS Operator" are executed first and can be executed in parallel. The "Spatial Intersection Operator" is a common child node of the two starting operators and is executed one step later. The "Write Shapefile Operator" is a child node of the "Spatial Intersection Operator" and is executed last. According to the dependency parsing logic, the "Spatial Intersection Operator" is a common child node of the two starting nodes, and its operation depends on the two starting nodes. According to the mapping relationship between the output parameters of the two starting nodes and the "Spatial Intersection Operator", its parameter dependency can be obtained. Similarly, the "Write Shapefile Operator" depends on the "Spatial Intersection Operator".

[0039] Step S200: Generate an executable file according to the code snippet corresponding to the operator to be executed and the dependency relationship.

[0040] In this embodiment, the defined operators are compiled into executable code (such as Java bytecode, Python script, etc.), or they are packaged into executable modules. According to the dependency table, a linear execution plan is constructed to ensure that the operators are executed in the correct order. According to the parsed relationship and code information, the code snippets are combined into Java applications, compiled and packaged into Jar packages, and the operator scheduling (i.e., Jar package) is submitted to the master node of the Spark cluster, and the Application ID and Job ID of the Spark job are obtained and associated with the task ID. When submitting the task Job, the spark.executor.instances, spark.executor.cores, and spark.executor.memory parameters can be used to specify the number of executors (computing nodes), the number of tasks running in parallel in each executor, and the maximum memory size used by each executor.

[0041] Step S300, dividing the running task corresponding to the executable file into a plurality of subtasks according to the partition of the distributed data set of the input data; Step S400, allocating the subtasks to the corresponding computing nodes in the distributed cluster for parallel execution, and obtaining execution results after the execution of each subtask is completed.

[0042] Please refer to Figure 2In this embodiment, after receiving the Jar package, the master node conducts an in-depth analysis of the running tasks corresponding to the executable file. The order of each stage is clarified. Some stages must be started after other stages are completed, such as the data preprocessing stage must be performed after the data reading stage. According to the results of the stage division, the master node further subdivides each stage into multiple subtasks. For example, in the data reading stage, if the amount of data is large, the data can be divided according to certain rules (such as by the size of the data block, by the partition where the data is located, etc.), and the reading operation of each data block becomes a subtask. Then, the master node allocates these subtasks to different computing nodes according to the resource conditions of the computing nodes and the characteristics of the subtasks. When allocating subtasks, the master node will only allocate the subtask when the pre-subtasks of a subtask have been successfully completed on the corresponding computing node. For example, if a data computing subtask depends on the result of the data preprocessing subtask, the master node will allocate the data computing subtask to the appropriate computing node only after the data preprocessing subtask is completed and the result is correctly passed to the master node. Through the above stage division and task allocation process, a corresponding list of subtasks and computing nodes is obtained. This list records which computing node each subtask is assigned to, as well as the execution order and dependencies between subtasks. Through this correspondence, the progress of task execution can be effectively monitored, and possible problems can be discovered and handled in a timely manner. For example, when a computing node fails, the subtasks on the node can be reallocated to other available nodes to ensure that the entire running task can be completed smoothly.

[0043] Please refer to Figure 3 In this embodiment, it includes an operator library, operator arrangement, scheduling engine and status monitoring module, and Spark distributed computing cluster. The operator library includes management functions such as creation, editing, publishing, and deletion, which can manage the operator throughout its life cycle; the operator arrangement module can create operator arrangement in a visual drag-and-drop manner based on the operator library, and perform routine management operations on the operator arrangement; the scheduling engine includes functions such as operator arrangement analysis, scheduling, submission operation, and operation status maintenance, which can manage and maintain the operation of the operator; the status monitoring module can continuously monitor the operation status of the operator, and feed back the operation status to the scheduling engine, so that the scheduling engine can complete the state maintenance operation of the operator; the distributed computing cluster is built based on the Spark framework and built in Standalone mode, which is used for the operation environment of the operator. By partitioning the distributed data set of the input data, subtasks running in parallel are created and assigned to computing nodes for calculation, so that a single operator node can run calculations on multiple servers, which improves the computing efficiency of spatial data under large data volumes.

[0044] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. On this basis, the following steps can also be included before step S100: Determine a deployment mode, and build the distributed computing cluster according to the deployment mode.

[0045] In this embodiment, a distributed computing cluster can be built based on the Spark framework to provide an operating environment for operators. Spark cluster configuration can be performed based on the Standalone mode and the YARN mode. After successfully configuring and starting the Spark cluster, data processing and analysis can be performed by submitting Spark applications to the cluster.

[0046] In this embodiment, the distributed computing framework Spark can also be replaced with Apache Tez, and the scheduling engine adapts to the task submission mode of the Tez Java API (using TezClient). The scheduling engine parses the operator orchestration and converts it into Tez's DAG object to define the data processing flow.

[0047] Create and encapsulate operators, and save the encapsulated operators to the operator library.

[0048] In this embodiment, operators are defined for processing data. The development specifications of the operator meet the development specification requirements of the Spark framework, and can be developed based on Java or Scala language. When creating an operator, paste the operator code into the code editing area to save and publish the operator, and encapsulate the operator to read and write two or three-dimensional spatiotemporal data and file type data as a general data read and write operator, and then put it on the operator library and publish it. The introduction of the spark-streaming module to process real-time streaming data is not supported here, because the spark-streaming module will not automatically exit after running under normal circumstances. It is a resident memory program and is not suitable as an operator node in the operator arrangement. Otherwise, the operator arrangement will be blocked in the running state for a long time and cannot complete the operation. After creating the operator, the operator can be edited, deleted, and published, and it can be used in the operator arrangement layout after publishing.

[0049] In this embodiment, for the spatiotemporal data reading and writing operators, their reading and writing logic is extracted and encapsulated as general data reading and writing operators, which support reading Shapefile, FileGDB, GeoJSON, PostGIS, files and other types. The string type is implemented in the form of startup parameters for reading input, and there is no need to encapsulate it separately as a general data reading and writing operator; the operators that read Shapefile, FileGDB, and GeoJSON files use the GeoTools tool to read and parse, convert them into Spark RDD (ResilientDistributed Dataset), and output the RDD as a parameter, where the spatial data is parsed into WKB format; the spatiotemporal data in PostGIS is read using the Spark SQL module and the GeoTools tool to read and convert them into Spark RDD, and output the RDD as a parameter, where the spatial data is parsed into WKB format, and the analysis and calculation of spatial data in subsequent operators are based on Spark RDD.

[0050] In this embodiment, the general spatiotemporal data reading and writing operator is designed as follows: For the FileGDB file reading operator, input parameter: FileGDB file address (local directory / shared directory / remote address). Processing logic: Use tools such as GeoTools to parse the FileGDB file, read spatial data and attribute collections, read spatial fields in WKB mode, and convert the data collection into Spark RDD. Output parameter: Spark RDD (distributed dataset).

[0051] For the Shapefile read operator, input parameter: Shapefile file address (local directory / shared directory / remote address). Processing logic: Use tools such as GeoTools to parse the Shapefile file, read spatial data and attribute collections, read spatial fields in WKB mode, and convert the data collection into Spark RDD. Output parameter: Spark RDD.

[0052] For the operator that reads GeoJSON files, input parameters: GeoJSON file address (local directory / shared directory / remote address). Processing logic: Use tools such as GeoTools to parse GeoJSON files, read spatial data and attribute collections, read spatial fields in WKB mode, and convert data collections into Spark RDD. Output parameters: Spark RDD (distributed dataset).

[0053] For the operator to read PostGIS, input parameters: A. PostGIS database data source information B. Name of the database spatial table to be read. Processing logic: Use the Spark SQL module to access the PostGIS library through JDBC to read the specified spatial table data and attribute collection. The spatial field is read in WKB mode, and the data collection is converted to Spark RDD. Output parameter: Spark RDD (distributed dataset).

[0054] For the file reading operator, input parameter: file address (local directory / shared directory / remote address). Processing logic: parse the file in text mode and convert the parsed text data into Spark RDD. Output parameter: SparkRDD (distributed dataset).

[0055] For the operator that writes out a FileGDB file, input parameters are: Spark RDD (including spatiotemporal data), file save address, and file name. Processing logic: Use tools such as GeoTools to convert the RDD dataset and generate a FileGDB file. Output parameter: FileGDB file For the Shapefile operator, input parameters are: Spark RDD (including spatiotemporal data), file save address, and file name. Use tools such as GeoTools to convert the RDD dataset and generate a Shapefile file. Output parameter: Shapefile file For the operator that writes GeoJSON files, input parameters are: Spark RDD (including spatiotemporal data), file save address, and file name. Processing logic: Use tools such as GeoTools to convert the RDD dataset and generate a GeoJSON file. Output parameter: GeoJSON file.

[0056] For the Write PostGIS operator, input parameters: A. Spark RDD (including spatiotemporal data) B. PostGIS data source information C. Saved table name. Processing logic: Use the Spark SQL module to save the RDD dataset to the specified PostGIS spatiotemporal database table. Output parameter: PostGIS spatiotemporal database table.

[0057] For the write file operator, input parameters: A. Spark RDD (including spatiotemporal data) B. File save address C. File name. Processing logic: Convert the RDD dataset into text format data and save it to the specified file. Output parameter: file.

[0058] In this embodiment, the operator embeds a standardized data read and write layer to uniformly perform data read and write operations to support the transfer of parameter types such as files, strings, and 2D and 3D data, and the parameter type transfer support is richer. This solves the problem that operators only support conventional string parameter transfer, but cannot support file and 2D and 3D spatiotemporal data type parameter transfer.

[0059] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 4 , before step S100, steps B100 to B500 may also be included: Step B100: If it is detected that the user drags a first operator onto the canvas, an identifier of the first operator is obtained.

[0060] In this embodiment, in the front-end code of the operator arrangement page, the event detection mechanism of JavaScript is used to add a drag event handler to each operator element in the operator library. For example, addEventListener('dragstart', function(event) { / * Record the start information of the drag* / }) is used to detect the drag start event, and addEventListener('drop', function(event) { / * Process the drag end event* / }) is used to detect the drag end event. When it is detected that the user drags an operator from the operator library to the canvas, that is, when the drop event is triggered, the unique identifier of the dragged operator (such as the name and ID of the operator) and other related attributes (such as the category to which it belongs, etc.) are obtained through the event object. For example, assuming that each operator element has a data-operator-type attribute to indicate its type, the type information of the dragged operator is obtained through event.target.dataset.operatorType.

[0061] Step B200: Acquire an associated operator of the first operator according to the identifier.

[0062] In this embodiment, the obtained operator type information is used as a query condition to search in the established operator association rule library. For example, if the "Read Shapefile Operator" is dragged out, the record with "Read Shapefile Operator" as the key is searched in the adjacency table. Based on the search results, the information of all subsequent operators associated with the operator is extracted. This information may include the name of the operator, function description, input and output data types, etc. For example, the "spatial intersection operator", "buffer operator", "spatial query operator" and their related attribute information associated with the "Read Shapefile Operator" are obtained from the adjacency table.

[0063] Step B300, obtaining a comprehensive prediction score of the association operator according to the usage frequency, association degree and corresponding weight of the association operator.

[0064] In this embodiment, the prediction score is determined by comprehensively considering the usage frequency of the operator and the correlation of the current operator. The usage frequency can be measured by counting the total number of times all users use the operator; the correlation can be determined according to the correlation weight defined in the association rule library.

[0065] For each association operator, calculate the score corresponding to each factor. For example, for the "spatial intersection operator", assuming that its frequency of use among all users accounts for 40% of the total score, the score is calculated as 0.3 based on the proportion of its use times in the total number of times all operators are used; the association score with the "read shapefile operator" accounts for 30% of the total score, and the score is calculated as 0.25 based on the weight in the association rule library. The scores of each factor are weighted and summed according to the set weights to obtain the comprehensive prediction score of each association operator.

[0066] Step B400: Arrange the associated operators in descending order according to the comprehensive prediction scores and store them in a recommended operator list.

[0067] Step B500: display the recommended operator list in the recommendation area.

[0068] In this embodiment, the associated operator list is sorted in descending order according to the calculated comprehensive prediction score, and the comprehensive prediction score is sorted from high to low, so that the operator with the highest score is ranked first in the list.

[0069] Design a prediction operator recommendation area in a suitable location of the operator arrangement canvas (such as a sidebar, pop-up window, etc.). Display the sorted list of associated operators in the designed recommendation area. For each operator, display its name, a brief functional description, and possible operation buttons (such as a button to drag directly to the canvas). Continue to monitor the user's operation behavior after displaying the prediction operator in real time through the front-end event detection mechanism. For example, detect whether the user clicks the drag button of the prediction operator, or whether other operators are dragged out from the operator library.

[0070] In this embodiment, the dependency between operators is determined according to the operator orchestration task. The dependency between operators is represented by a directed graph, where nodes represent operators and edges represent data flow directions. The operator orchestration graph is generated using a graph data structure (such as an adjacency list, adjacency matrix, etc.) or a graph library. Different operators are combined and connected through a visual operator orchestration page, and the complete process of the output is output to construct a directed acyclic graph. For example, please refer to Figure 5, the specific operator arrangement process is as follows: Step (1), in the operator arrangement page, drag out the "Read Shapefile Operator" and "Read PostGIS Operator" from the operator library, which are used to read the two spatiotemporal data for spatial intersection calculation. Step (2), drag out the "Spatial Intersection Operator" from the operator library to the canvas, and connect it with the above two operators for reading spatiotemporal data as the input parameters of the spatial intersection operator. Step (3), drag out the "Write Shapefile Operator" from the operator library, connect it with the "Spatial Intersection Operator", use the output result of the "Spatial Intersection Operator" as the input parameter of the operator, and fill in the file save address and file name, and the DAG operator arrangement process for spatial intersection calculation is completed.

[0071] In a feasible implementation manner, before step B100, the following steps may also be included: An operator association rule base is established to store the successor operators of each of the first operators and the corresponding association degrees.

[0072] When it is detected that the user selects the prediction operator of the recommendation area, the association degree between the prediction operator and the first operator is updated.

[0073] In this implementation, the degree of association between operators is determined based on the function of each operator, the input and output data formats, and the applicable business scenarios. For example, the "read Shapefile file operator" is used to read spatial data from a Shapefile file, and its output is a spatial data object in a specific format. Through analysis, it can be seen that operators that process spatial data may be connected later, such as spatial analysis, conversion, and storage-related operators. The association rules are stored in the form of an adjacency list. The adjacency list is a storage structure of a graph. For each operator, its possible successor operators are recorded in the table. For example, for the "read Shapefile file operator", the "spatial intersection operator", "buffer operator", "spatial query operator", etc. associated with it are recorded in the adjacency list.

[0074] In this embodiment, based on the operator currently dragged by the user, the operator that may be used later is quickly recommended, which reduces the time for the user to search and filter in the operator library, allowing the user to complete the operator arrangement task more quickly and improve the overall operation efficiency. Based on the intrinsic association and actual usage between operators, reasonable operation suggestions are provided to users, reducing the user's learning cost and operation difficulty. In addition, by considering the frequency of use, correlation and corresponding weights of the associated operators to calculate the comprehensive prediction score, it is possible to guide users to use operators more reasonably, avoid excessive use of some uncommon or unnecessary operators, and thus optimize the allocation and utilization of system resources.

[0075] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 6 , step S300 may include steps S310 to S330: Step S310, based on the master node of the distributed cluster, dividing the executable file into task stages to obtain each task stage and the execution order of the task stages; In this embodiment, after the JAR package containing the application is passed to the master node of the Spark cluster, the master node decompresses and checks the contents of the JAR package. Find the class containing the main method, which is usually the entry point of the application and defines the main logic of the task. For example, in a Java application, the master node will find the main class based on the --class parameter specified in the spark-submit command or the manifest file of the JAR package (if configured). This main class contains the overall task description of the application, such as which operator orchestration operations to perform, which data processing logic to use, and the final result output.

[0076] Step S320, creating and partitioning the distributed data set of the input data according to the execution order and the data source path of the executable file, wherein the partitioning strategy is a spatial dimension and / or a temporal dimension; Step S330: Generate the subtasks to be executed in parallel based on the partitioning result and the cluster configuration.

[0077] In this embodiment, the master node determines the stage of the task based on the dependency and execution order information of the operator arrangement. Such as the data reading stage, the data processing stage, the result storage stage, etc. To ensure that the tasks are carried out in the correct logical order and avoid data dependency conflicts. The master node creates subtasks (Tasks) that run in parallel according to the partitioning of the input data RDD. First, according to the code and data source path in the JAR package, the Spark cluster uses the corresponding data reading logic (such as reading data from the file system, database or other data source) to create the input data RDD. The number of partitions is dynamically selected according to the size of the input data and the configuration of the cluster. For example, for a large text file, Spark may divide it into multiple data blocks, each data block as a partition, to form a distributed data set (RDD). Each partition will become the input of a subtask. For example, if the application uses sc.textFile("path / to / input / file") to read a file, the master node will use this code snippet to read the file data into the RDD according to the location of the file and Spark's file reading mechanism. At the same time, the number of partitions is automatically determined based on the file size and cluster configuration (such as parameters such as spark.default.parallelism) to create a distributed data set. After the input data RDD is created, subsequent operators (such as data processing operators) will read the RDD for operation. For example, a map or filter operator will read each partition of the RDD and perform corresponding calculation operations on the data in the partition. This reading process is distributed, and each computing node (Executor) will read the RDD partition assigned to it and perform the calculation task.

[0078] In this embodiment, partitioning is performed according to a partitioning strategy of a spatial dimension and / or a temporal dimension. In a partitioning strategy based on a spatial range, the geographic space is divided into multiple sub-areas, and each sub-area corresponds to a partition. For example, for global map data, it can be divided according to the latitude and longitude range, such as a partition for every 10 degrees of longitude and latitude. It can also be divided according to administrative regions, such as partitioning by country, province, city, etc. When the query is mainly based on a specific spatial area, the relevant partition can be quickly located to reduce unnecessary data scanning. For example, in an urban planning project, if you want to analyze the changes in population density in different areas, the city map can be partitioned according to blocks or administrative divisions. Each partition stores the population data and related geographic information in the area. When you need to query the population density of a certain block, directly obtain data from the corresponding partition for calculation, without traversing the data of the entire city.

[0079] In this embodiment, in the partition strategy based on time range, data is divided according to the time dimension, such as allocating spatiotemporal data to different partitions according to time intervals such as days, months, and years. For example, for meteorological monitoring data, data is stored in one partition per day. When it is necessary to analyze the meteorological data of a certain week, only the data of the corresponding 7 partitions need to be read, which improves the query efficiency.

[0080] In this embodiment, in a hybrid partitioning strategy based on spatial grid and time slice, the spatial and temporal dimensions are combined for partitioning. The space is first divided into grid-shaped sub-areas, and then each sub-area is sliced ​​and partitioned according to time. The hybrid strategy can take into account the query requirements of both spatial and temporal dimensions, and improve the efficiency of query and calculation. In the marine ecological monitoring project, the marine area is divided into several grids, and each grid stores data such as ocean temperature and salinity by month. When studying the ecological changes in a certain sea area within a specific time period, the corresponding spatial grid and time slice partition can be quickly located to obtain the required data.

[0081] Based on the first embodiment of the present application, in the fifth embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description, and will not be described in detail later. Figure 6 , step S400 may include steps S410 to S420: Step S410, sending the subtask to the corresponding computing node according to the task distribution algorithm; Step S420: After receiving the task, the computing node calls the operator to be executed to process the data and outputs the execution result.

[0082] In this embodiment, the master node first serializes the subtasks into a format suitable for network transmission, such as JSON, XML or a custom binary format. The serialized tasks are sent to the corresponding computing nodes through the communication mechanism provided by the distributed computing framework. After receiving the task, the computing node will send a confirmation message to the master node, indicating that the task has been successfully received and is ready to be executed. Then, the task is deserialized from the serialized format into an internal representation for execution. According to the operator information in the task, the computing node calls the corresponding operator function or module to process the data. During the execution process, the computing node can store the intermediate results in the local disk or memory for subsequent processing or exchange data with other nodes. When the operator is executed, the computing node generates output data and prepares it to be sent to the master node or for the next step of processing. When all subtasks are executed, the master node sends a request to collect output data to the computing node. The computing node sends the locally stored output data to the master node, and returns it to the user through the user interface or storage interface or stores it in a specified location.

[0083] In a feasible implementation, tasks can be distributed according to a load balancing algorithm, and tasks can be assigned to nodes with lower loads to ensure the overall performance of the system. The load balancing algorithm can measure the load of nodes according to different indicators, such as CPU usage, memory occupancy, network bandwidth, etc. By regularly monitoring these indicators and adjusting the allocation strategy, it is ensured that each computing node maintains a relatively balanced load level.

[0084] In a feasible implementation, tasks can be distributed according to a distributed hash table (DHT) algorithm, where tasks are regarded as data items, and a hash function is used to calculate the hash value of each task, and then the tasks are assigned to corresponding nodes according to the hash value.

[0085] Based on the first embodiment of the present application, in the sixth embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description, and will not be described in detail later. Figure 7 , before step S400, steps A100 to A300 may also be included: Step A100, creating a monitoring linked list, and storing the running task in the monitoring linked list.

[0086] Step A200, performing a traversal operation on the monitoring linked list to obtain the running status of the running task.

[0087] Step A300, when the running status is running completed, task completion information is generated, and the corresponding running task is deleted from the monitoring linked list.

[0088] Please refer to Figure 8 In this implementation, the operating status of the operator can be continuously monitored based on the status monitoring module, and the operating status can be fed back to the scheduling engine, so that the scheduling engine can complete the state maintenance operation of the operator. Create a monitoring linked list in the status monitoring module, put the running tasks at the end of the linked list, and associate the Spark Application ID and Job ID. At regular intervals, traverse the monitoring linked list and request the Spark REST API interface / api / v1 / applications / <application-id> / jobs / <job-id>Get the running status of the task job. When the operation is completed, push the running result to the scheduling engine and remove the task job from the monitoring list. The scheduling engine marks the operator as completed and starts the scheduling of the next operator until the entire operator is completed.

[0089] In this implementation, a monitoring linked list is created to store currently running tasks. The Application ID and Job ID of the task are used as parameters, and a new task instance is created and then added to the linked list. The tasks in the monitoring linked list are traversed regularly. For each task, the Spark REST API is used to query its status. When the task status is checked, corresponding actions are taken according to the returned status value (such as "SUCCEEDED", "RUNNING", "FAILED", etc.). If the task is successfully completed ("SUCCEEDED"), the interface of the scheduling engine is called to mark the task as completed and remove the task from the monitoring linked list. If the task fails or is still running, monitoring continues.

[0090] Through the above steps, a status monitoring module can be implemented to track the running status of Spark tasks and push the results to the scheduling engine when the tasks are completed, thereby supporting complex job scheduling and orchestration.

[0091] In this embodiment, the state monitoring module can also be removed, and the state callback can be used as the callback method after the Spark task is completed. The SparkListener class is used to implement state detection at the beginning of the task node. When the task is completed, the state feedback interface of the scheduling engine is called to modify the state of the task node in the scheduling engine. The disadvantage is that this implementation is invasive to the task node, and callback logic code needs to be added when the task node is developed; if the task node reports an abnormal error, the program terminates to avoid the inability to feed back the running state to the scheduling engine, resulting in the task node state in the scheduling engine being in a running state and unable to end.

[0092] The present application provides an operator orchestration device for spatiotemporal data analysis and calculation, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the operator orchestration method for spatiotemporal data analysis and calculation in the above-mentioned embodiment one.

[0093] Reference below Fig. 9 , which shows a schematic diagram of the structure of an operator arrangement device suitable for implementing the spatiotemporal data analysis and calculation in the embodiment of the present application. The operator arrangement device for spatiotemporal data analysis and calculation in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDA, Personal Digital Assistant), tablet computers (PAD, portable android device), portable multimedia players (PMP, Portable Media Player), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig. 9 The operator orchestration device for spatiotemporal data analysis calculation shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0094] like Fig. 9 As shown, the operator arrangement device for spatiotemporal data analysis and calculation may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in a read-only memory (ROM) 1002 or the program loaded from a storage device 1003 to a random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the operator arrangement device for spatiotemporal data analysis and calculation are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the operator orchestration device for spatiotemporal data analysis and calculation to communicate wirelessly or wired with other devices to exchange data. Although the figure shows an operator orchestration device for spatiotemporal data analysis and calculation with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.

[0095] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0096] The operator orchestration device for spatiotemporal data analysis and calculation provided by the present application adopts the operator orchestration method for spatiotemporal data analysis and calculation in the above-mentioned embodiment, which can solve the technical problem of how to improve the operating efficiency of a single operator node to support large-scale data volume calculation. Compared with the prior art, the beneficial effects of the operator orchestration device for spatiotemporal data analysis and calculation provided by the present application are the same as the beneficial effects of the operator orchestration method for spatiotemporal data analysis and calculation provided by the above-mentioned embodiment, and the other technical features in the operator orchestration device for spatiotemporal data analysis and calculation are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0097] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0098] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0099] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, wherein the computer-readable program instructions are used to execute the operator arrangement method for spatiotemporal data analysis calculations in the above-mentioned embodiments.

[0100] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequencies (RF, Radio Frequency), etc., or any suitable combination of the above.

[0101] The above-mentioned computer-readable storage medium may be included in the operator orchestration device of spatiotemporal data analysis and calculation; or it may exist independently and not be assembled into the operator orchestration device of spatiotemporal data analysis and calculation. The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the operator orchestration device of spatiotemporal data analysis and calculation, the operator orchestration device of spatiotemporal data analysis and calculation: after generating the operator orchestration graph, parse the operator orchestration graph to determine the operators to be executed and the dependencies between the operators to be executed; generate an executable file according to the code snippets corresponding to the operators to be executed and the dependencies; divide the running tasks corresponding to the executable file into multiple subtasks according to the partitioning of the distributed data set of the input data; assign the subtasks to the corresponding computing nodes in the distributed cluster for parallel execution, and obtain the execution results after the execution of each of the subtasks is completed.

[0102] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0103] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0104] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0105] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the operator arrangement method for the above-mentioned spatiotemporal data analysis calculation, and can solve the technical problem of how to improve the operating efficiency of a single operator node to support large-scale data volume calculation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the operator arrangement method for spatiotemporal data analysis calculation provided in the above-mentioned embodiment, and will not be repeated here.

[0106] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for arranging operators for spatiotemporal data analysis and calculation, characterized in that: The method includes: After generating the operator orchestration graph, the operator orchestration graph is parsed to determine the operators to be executed and the dependencies between the operators to be executed; Generate an executable file according to the code snippet corresponding to the operator to be executed and the dependency relationship; Dividing the running task corresponding to the executable file into a plurality of subtasks according to the partition of the distributed data set of the input data; The subtasks are assigned to the corresponding computing nodes in the distributed cluster for parallel execution, and the execution results are obtained after the execution of each subtask is completed.

2. The operator arrangement method for spatiotemporal data analysis and calculation according to claim 1, characterized in that: The operator orchestration graph is a directed acyclic graph, and the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators to be executed includes: Determine a starting node in the operator orchestration graph; Based on the depth-first search algorithm, traverse all successor nodes of the starting node; The dependency relationship is obtained according to the traversal order and the direction of the edges in the directed graph.

3. The operator arrangement method for spatiotemporal data analysis and calculation according to claim 1, characterized in that: After generating the operator orchestration graph, before the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators to be executed, the method further includes: If it is detected that the user drags a first operator onto the canvas, an identifier of the first operator is obtained; Acquire an associated operator of the first operator according to the identifier; Obtaining a comprehensive prediction score of the association operator according to the usage frequency, association degree and corresponding weight of the association operator; Arrange the associated operators in descending order according to the comprehensive prediction scores, and store them in a recommended operator list; The recommended operator list is displayed in the recommendation area.

4. The operator arrangement method for spatiotemporal data analysis and calculation according to claim 3, characterized in that: Before the step of obtaining the identifier of the first operator if it is detected that the user drags the first operator onto the canvas, the method further includes: Establishing an operator association rule base to store successor operators of each of the first operators and corresponding association degrees; When it is detected that the user selects the prediction operator of the recommendation area, the association degree between the prediction operator and the first operator is updated.

5. The operator arrangement method for spatiotemporal data analysis and calculation according to claim 1, characterized in that: The step of dividing the running task corresponding to the executable file into a plurality of subtasks according to the partition of the distributed data set of the input data comprises: Based on the master node of the distributed cluster, the executable file is divided into task stages to obtain each task stage and the execution order of the task stages; According to the execution order and the data source path of the executable file, the distributed data set of the input data is created and partitioned, and the partitioning strategy is the space dimension and / or the time dimension; The subtasks to be executed in parallel are generated based on the partitioning result and the cluster configuration.

6. The operator arrangement method for spatiotemporal data analysis and calculation according to claim 1, characterized in that: Before the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators, the method further includes: Determine a deployment mode, and build the distributed computing cluster according to the deployment mode; Create and encapsulate operators, and save the encapsulated operators to the operator library.

7. The operator arrangement method for spatiotemporal data analysis and calculation according to claim 1, characterized in that: Before the step of assigning the subtask to the corresponding computing node for execution and obtaining the execution result, the step further includes: Create a monitoring chain list, and store the running tasks in the monitoring chain list; Performing a traversal operation on the monitoring linked list to obtain the running status of the running task; When the running status is running completed, task completion information is generated, and the corresponding running task is deleted from the monitoring linked list.

8. The operator arrangement method for spatiotemporal data analysis and calculation according to claim 1, characterized in that: The step of allocating the subtasks to the corresponding computing nodes in the distributed cluster for parallel execution, and obtaining the execution results after each of the subtasks is completed, comprises: Sending the subtask to the corresponding computing node according to a task distribution algorithm; After receiving the task, the computing node calls the operator to be executed to process the data and outputs the execution result.

9. An operator arrangement device for spatiotemporal data analysis and calculation, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the operator arrangement method for spatiotemporal data analysis calculation as described in any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the operator arrangement method for spatiotemporal data analysis and calculation are implemented as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Space-time data visualization task execution method based on cloud edge-end architecture

    CN113032132A

  • Task processing flow configuration method and device, electronic equipment and storage medium

    CN115794064A

  • Data processing method and device based on space-time big data engine

    CN117931436A

  • Model operator parallel splitting method and device, equipment and storage medium

    CN119294463A

  • Task management method, system, device and medium based on graph structure

    CN119781961A

Cited By

  • Graph database query execution engine and method

    CN120386897A