Operator Scheduling Method, Device, and Storage Medium for Spatiotemporal Data Analysis and Calculation
Through the analysis and partitioning of the operator orchestration graph, the operator task is split into multiple subtasks and executed in parallel in a distributed cluster, solving the problem of performance limitation of a single server, and achieving efficient large-scale data calculation and parameter transfer of multi-dimensional data types.
Patent Information
- Application Number
- CN202510438083.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the prior art, the performance of a single server limits the operation efficiency of the operator node and cannot support large-scale data calculations. Moreover, the operators only support conventional string parameter transfer, and cannot support file and two- and three-dimensional spatiotemporal data type parameter transfer.
By analyzing the operator arrangement diagram, determining the operator to be executed and its dependencies, generating an executable file, and splitting it into multiple subtasks assigned to the computing nodes in the distributed cluster for parallel execution, supporting file and two- and three-dimensional spatiotemporal data type parameter transfer.
It improves the computing efficiency of a single operator node on multiple servers, supports large-scale data calculation, and realizes parameter transfer of files and two- and three-dimensional spatiotemporal data types.
Smart Images

Figure CN119938283B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to an operator orchestration method, device, and storage medium for spatio-temporal data analysis and calculation. Background Art
[0002] In an operator orchestration system, running operators involves the process orchestration and scheduling execution of operator nodes. An operator node represents a specific operator operation, and the process orchestration defines the execution order and relationship between these operator nodes. The scheduling engine allocates operator nodes to different servers or resources for running according to the orchestrated process.
[0003] Generally, after completing the process design and orchestration of operator nodes and entering the running phase, the scheduling engine allocates each operator node to a server for running one by one according to the orchestrated process. One server can run one or more operator nodes simultaneously. However, it is impossible to disperse the computing tasks of a single operator node across multiple servers. Due to the limited overall performance of a single server, the fact that a single operator node can only run on one server limits the running efficiency of the operator node and cannot support large-scale data volume calculations.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of this application is to provide an operator orchestration method, device, and storage medium for spatio-temporal data analysis and calculation, aiming to solve the technical problem of how to improve the running efficiency of a single operator node to support large-scale data volume calculations.
[0006] To achieve the above objective, this application proposes an operator orchestration method for spatio-temporal data analysis and calculation. The operator orchestration method for spatio-temporal data analysis and calculation includes:
[0007] After generating an operator orchestration graph, parse the operator orchestration graph to determine the operators to be executed and the dependency relationships between the operators to be executed;
[0008] Generate an executable file according to the code snippet corresponding to the operator to be executed and the dependency relationship;
[0009] Divide the running tasks corresponding to the executable file into multiple subtasks according to the partitioning of the distributed dataset of the input data;
[0010] Allocate the subtasks to the corresponding computing nodes in the distributed cluster for parallel execution. When each of the subtasks is executed, an execution result is obtained.
[0011] In one embodiment, the operator orchestration graph is a directed acyclic graph. The step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators to be executed includes:
[0012] Determine the starting node in the operator orchestration graph;
[0013] Based on the depth - first search algorithm, traverse all successor nodes of the starting node;
[0014] According to the traversal order and the direction of the edges in the directed graph, obtain the dependencies.
[0015] In one embodiment, before the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators to be executed after generating the operator orchestration graph, it further includes:
[0016] If it is detected that the user drags out a first operator onto the canvas, obtain the identifier of the first operator;
[0017] According to the identifier, obtain the associated operators of the first operator;
[0018] According to the usage frequency, association degree, and corresponding weights of the associated operators, obtain the comprehensive prediction score of the associated operators;
[0019] According to the comprehensive prediction score, sort the associated operators in descending order and store them in the recommended operator list;
[0020] Display the recommended operator list in the recommendation area.
[0021] In one embodiment, before the step of if it is detected that the user drags out a first operator onto the canvas and obtain the identifier of the first operator, it further includes:
[0022] Establish an operator association rule library, store the successor operators of each first operator, and the corresponding association degrees;
[0023] When it is detected that the user selects a predicted operator in the recommendation area, update the association degree between the predicted operator and the first operator.
[0024] In one embodiment, the step of dividing the running tasks corresponding to the executable file into multiple subtasks according to the partition of the distributed dataset of the input data includes:
[0025] Based on the master node of the distributed cluster, perform task - phase division on the executable file to obtain each task phase and the execution order of the task phase;
[0026] Create and partition the distributed dataset of the input data according to the execution order and the data source path of the executable file, and the partitioning strategy is the spatial dimension and / or the temporal dimension;
[0027] Generate the subtasks for parallel execution based on the partitioning result and the cluster configuration.
[0028] In one embodiment, before the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators, the method further includes:
[0029] Determine the deployment mode and set up the distributed computing cluster according to the deployment mode;
[0030] Create and encapsulate operators, and save the encapsulated operators to the operator library.
[0031] In one embodiment, before the step of assigning the subtasks to the corresponding computing nodes for execution to obtain the execution result, the method further includes:
[0032] Create a monitoring linked list and store the running tasks in the monitoring linked list;
[0033] Perform a traversal operation on the monitoring linked list to obtain the running status of the running tasks;
[0034] When the running status is running completed, generate a task completion message and delete the corresponding running task from the monitoring linked list.
[0035] In one embodiment, the step of assigning the subtasks to the respective corresponding computing nodes in the distributed cluster for parallel execution and obtaining the execution result after each subtask is executed includes:
[0036] Send the subtasks to the corresponding computing nodes according to the task distribution algorithm;
[0037] After receiving the tasks, the computing nodes call the operators to be executed for data processing and output the execution result.
[0038] In addition, to achieve the above object, the present application also proposes an operator orchestration device for spatio-temporal data analysis and calculation, the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the operator orchestration method for spatio-temporal data analysis and calculation as described above.
[0039] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the operator scheduling method for spatio-temporal data analysis and calculation as described above are implemented.
[0040] The present application provides an operator scheduling method for spatio-temporal data analysis and calculation, which parses an operator scheduling graph to determine the operators to be executed and the dependencies between them, and generates an executable file. By encapsulating the operator scheduling logic into a whole, it is convenient for subsequent unified management and scheduling. According to the partitioning situation of the distributed data set of the input data, parallel-running subtasks are created and assigned to computing nodes for calculation. By splitting the operator tasks into multiple subtasks and assigning the subtasks to multiple computing nodes in the cluster, it is possible to realize the operation and calculation of a single operator node on multiple servers, improving the calculation efficiency of spatial data under a large amount of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0042] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0043] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the operator scheduling method for spatio-temporal data analysis and calculation of the present application;
[0044] Figure 2 It is a schematic overall flowchart provided for the operator scheduling method for spatio-temporal data analysis and calculation of the present application;
[0045] Figure 3 It is a functional architecture diagram provided for the operator scheduling method for spatio-temporal data analysis and calculation of the present application;
[0046] Figure 4 It is a schematic flowchart provided for Embodiment 3 of the operator scheduling method for spatio-temporal data analysis and calculation of the present application;
[0047] Figure 5 It is an operator scheduling schematic diagram provided for the operator scheduling method for spatio-temporal data analysis and calculation of the present application;
[0048] Figure 6 It is a schematic flowchart provided for Embodiments 4 and 5 of the operator scheduling method for spatio-temporal data analysis and calculation of the present application;
[0049] Figure 7 It is a schematic flowchart provided for the sixth embodiment of the operator scheduling method for spatio-temporal data analysis and calculation in this application;
[0050] Figure 8 It is a sequence diagram provided for the operator scheduling method for spatio-temporal data analysis and calculation in this application;
[0051] Figure 9 It is a schematic structural diagram of the hardware operating environment involved in the operator scheduling method for spatio-temporal data analysis and calculation in the embodiments of this application.
[0052] The implementation, functional features and advantages of the purpose of this application will be further described in combination with the embodiments with reference to the accompanying drawings. Specific implementation manners
[0053] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0054] In order to better understand the technical solutions of this application, the following will be described in detail in combination with the drawings in the specification and specific implementation manners.
[0055] The main solution of the embodiments of this application is: after generating an operator scheduling diagram, parsing the operator scheduling diagram to determine the operators to be executed and the dependency relationships between the operators to be executed; generating an executable file according to the code segments corresponding to the operators to be executed and the dependency relationships; dividing the running tasks corresponding to the executable file into multiple subtasks according to the partitions of the distributed dataset of the input data; allocating the subtasks to the corresponding computing nodes in the distributed cluster for parallel execution, and obtaining an execution result after each subtask is executed.
[0056] In an operator scheduling system, running an operator involves the process scheduling and scheduling execution of operator nodes. An operator node represents a specific operator operation, and the process scheduling defines the execution order and relationships between these operator nodes. The scheduling engine allocates operator nodes to different servers or resources for running according to the scheduling process.
[0057] In the prior art, after the process design and orchestration of operator nodes are completed and enter the running stage, the scheduling engine distributes each operator node to a server for running according to the orchestrated process. One or more operator nodes can run on a single server simultaneously. However, it is impossible to disperse the computing tasks of a single operator node across multiple servers. Due to the limited overall performance of a single server, the fact that a single operator node can only run on one server limits the running efficiency of the operator node and cannot support large-scale data volume calculations. In addition, in existing operator orchestration systems, only conventional string parameter passing is supported between operators, and parameter passing of file and two- and three-dimensional spatio-temporal data types is not supported.
[0058] To solve the above problems, the present application provides an operator orchestration method for spatio-temporal data analysis and calculation. The operator orchestration graph is parsed to determine the operators to be executed and the dependencies between them, and an executable file is generated. By encapsulating the operator orchestration logic into a whole, it is convenient for subsequent unified management and scheduling. According to the partitioning situation of the distributed data set of the input data, parallel-running subtasks are created and assigned to computing nodes for calculation. By splitting the operator task into multiple subtasks and distributing the subtasks to multiple computing nodes in the cluster, it is possible to realize the running and calculation of a single operator node on multiple servers, improving the calculation efficiency of spatio-temporal data under a large amount of data.
[0059] It should be noted that the execution subject of this embodiment can be a computing service device with network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, apparatus, etc. that can implement the above functions. Hereinafter, taking the operator orchestration device for spatio-temporal data analysis and calculation as an example, this embodiment and the following embodiments will be described.
[0060] Based on this, the embodiment of the present application provides an operator orchestration method for spatio-temporal data analysis and calculation, referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the operator orchestration method for spatio-temporal data analysis and calculation of the present application.
[0061] In this embodiment, the operator orchestration method for spatio-temporal data analysis and calculation is applied to an operator orchestration device for spatio-temporal data analysis and calculation, and the method includes steps S100 to S400:
[0062] Step S100, after generating the operator orchestration graph, parse the operator orchestration graph to determine the operators to be executed and the dependencies between the operators to be executed.
[0063] It should be noted that an operator can be regarded as an operation or function that receives one or more inputs and processes these inputs according to certain rules or algorithms, and finally outputs one or more results. Operator scheduling refers to operations such as reasonably arranging the execution order of operators, allocating computing resources, and optimizing the transmission of data between storage and computing units.
[0064] In this embodiment, when performing the operator scheduling task, a series of operators are defined to process data. And the dependency relationships between the operators are determined, that is, the outputs of which operators are the inputs of other operators. A graph data structure (such as an adjacency list, an adjacency matrix, etc.) or a graphics library is used to generate an operator scheduling graph. According to graph traversal algorithms such as depth-first search (DFS, Depth-First Search) or breadth-first search (BFS, Breadth-First Search), the operator scheduling graph is traversed. During the traversal process, all the operators that need to be executed are recorded. A dependency table is established for each operator, listing all its preceding operators. The dependency relationships between the operators are determined according to the traversal order and the direction of the edges.
[0065] In a feasible implementation manner, step S100 may include the following steps:
[0066] Determine the starting node in the operator scheduling graph.
[0067] Based on the depth-first search algorithm, traverse all the successor nodes of the starting node.
[0068] According to the traversal order and the direction of the edges in the directed graph, obtain the dependency relationship.
[0069] In this implementation manner, the operator scheduling graph is a directed acyclic graph. A directed acyclic graph (DAG, Directed Acyclic Graph) consists of vertices and edges. A directed acyclic graph is a directed graph in which each edge has a clear direction and the entire graph is acyclic. That is, there is no path in the graph that can start from a vertex, pass through a series of edges, and then return to the vertex. Each edge points from one vertex to another vertex, representing a one-way relationship or dependency.
[0070] In this embodiment, after the operator scheduling is completed according to the directed acyclic graph specification, the scheduling engine parses the operator scheduling based on the depth-first search algorithm, and parses key information such as the dependency relationship and execution order of the operator scheduling. In a directed acyclic graph, nodes represent operators, and directed edges represent the dependency relationships between operators. When parsing according to the depth-first search algorithm, first, create a stack to store the nodes to be visited. Create a set or array to record the nodes that have been visited to prevent repeated visits. Secondly, perform a traversal to find all operators in the directed acyclic graph that have no predecessor nodes (i.e., nodes with an in-degree of 0), and use these operators as the starting nodes for the depth-first search. Push the starting nodes onto the stack and mark them as visited. When the stack is not empty, execute the following loop: Pop a node from the top of the stack and denote it as the current node. Process the current node and record its information as a certain operator (such as operator ID, type, etc.). Traverse all successor nodes of the current node. For each unvisited successor node, push it onto the stack and mark the current node as its predecessor node. Mark the current node as fully visited (i.e., all its successor nodes have been visited or determined not to be visited through the depth-first search). During the depth-first search traversal, use a data structure (such as a hash table, adjacency list, or graph database) to record the dependency relationships between operators. When the stack is empty and all nodes have been visited, the depth-first search traversal ends. At this time, the dependency relationship table or graph contains all the dependency relationships between operators.
[0071] In another feasible embodiment, it is also possible to traverse the nodes layer by layer according to the breadth-first search to obtain the dependency relationships between the nodes. When obtaining the operator dependency relationships, start from a starting operator and expand outward layer by layer until all reachable operators are traversed. Determine the dependency relationships between operators according to the traversal order and the direction of the edges.
[0072] Exemplarily, in the process of parsing whether two or more spatial objects intersect and calculating their intersection part in the operator scheduling, the following steps are specifically included: Step (1), enumerate the starting points of the DAG operator scheduling, that is, the operators without parent nodes, and the obtained starting point operators are: "Read Shapefile File Operator", "Read PostGIS Operator". Step (2), traverse from the starting point operator "Read Shapefile File Operator" using the DFS algorithm. Its child node is the "Spatial Intersection Operator". After a complete traversal, the result is: "Read Shapefile File Operator" -> "Spatial Intersection Operator" -> "Write Shapefile File Operator". Step (3), traverse from the other starting point operator "Read PostGIS Operator" using the DFS algorithm. Its child node is the "Spatial Intersection Operator". After a complete traversal, the result is: "Read PostGIS Operator" -> "Spatial Intersection Operator" -> "Write Shapefile File Operator". Step (4), parse the execution order and dependencies of the DAG operator scheduling according to the above traversal results. The two starting point operators "Read Shapefile File Operator" and "Read PostGIS Operator" are executed first and can be executed in parallel. The "Spatial Intersection Operator", as the common child node of the two starting operators, is executed in the next step. The "Write Shapefile File Operator", which is the child node of the "Spatial Intersection Operator", is executed last; for the dependency parsing logic, the "Spatial Intersection Operator" is the common child node of the two starting nodes, and its operation depends on these two starting nodes. According to the mapping relationship between the output parameters of the two starting nodes and the "Spatial Intersection Operator", its parameter dependency relationship can be obtained. Similarly, the "Write Shapefile File Operator" depends on the "Spatial Intersection Operator".
[0073] Step S200, generate an executable file according to the code snippet corresponding to the operator to be executed and the dependency relationship.
[0074] In this embodiment, the defined operators are compiled into executable code (such as Java bytecode, Python scripts, etc.), or they are packaged into executable modules. According to the dependency table, a linear execution plan is constructed to ensure that the operators are executed in the correct order. According to the parsed relationships and code information, code snippets are combined into a Java application and compiled and packaged into a Jar package. The operator orchestration and scheduling (i.e., the Jar package) is submitted to the master node of the Spark cluster, and the Application ID and Job ID of the Spark job are obtained and associated with the task ID. When submitting the task Job, the parameters spark.executor.instances, spark.executor.cores, and spark.executor.memory can be used to specify the number of executors (computing nodes), the number of tasks running in parallel in each executor, and the maximum memory size used by each executor, respectively.
[0075] Step S300: According to the partitions of the distributed dataset of the input data, the running tasks corresponding to the executable files are divided into multiple subtasks.
[0076] Step S400: The subtasks are assigned to the corresponding computing nodes in the distributed cluster for parallel execution. When all the subtasks are executed, the execution results are obtained.
[0077] Please refer to Figure 2, in this embodiment, after receiving the Jar package, the master node deeply analyzes the running tasks corresponding to the executable files, clarifies the sequence of each stage, and determines that some stages must start after other stages are completed. For example, the data preprocessing stage must be carried out after the data reading stage. According to the result of stage division, the master node further divides each stage into multiple subtasks. For example, in the data reading stage, if the data volume is large, the data can be divided according to certain rules (such as by data block size, by data partition, etc.), and the reading operation of each data block becomes a subtask. Then, the master node distributes these subtasks to different computing nodes according to the resource conditions of the computing nodes and the characteristics of the subtasks. When distributing subtasks, the master node will only distribute a subtask when all its prerequisite subtasks have been successfully completed on the corresponding computing nodes. For example, if a data calculation subtask depends on the result of a data preprocessing subtask, after the data preprocessing subtask is completed and the result is correctly transmitted to the master node, the master node will distribute the data calculation subtask to a suitable computing node. Through the above stage division and task distribution process, a corresponding list of subtasks and computing nodes is obtained. This list records which computing node each subtask is assigned to for execution, as well as the execution sequence and dependency relationship between subtasks. Through this correspondence, the execution progress of the task can be effectively monitored, and possible problems can be discovered and processed in a timely manner. For example, when a computing node fails, the subtasks on that node can be reassigned to other available nodes to ensure the smooth completion of the entire running task.
[0078] Please refer to Figure 3 , in this embodiment, it includes an operator library, operator orchestration, a scheduling engine, a status monitoring module, and a Spark distributed computing cluster. The operator library includes management functions such as creation, editing, publishing, and deletion, enabling full lifecycle management of operators; the operator orchestration module can create operator orchestration in a visual drag-and-drop manner based on the operator library and perform routine management operations on the operator orchestration; the scheduling engine includes functions such as operator orchestration parsing, scheduling, submission for execution, and running status maintenance, enabling management and maintenance of operator execution; the status monitoring module can continuously monitor the running status of operators and feedback the running status to the scheduling engine to facilitate the scheduling engine to complete the status maintenance operation of operators; the distributed computing cluster is built based on the Spark framework and in the Standalone mode, providing an operating environment for operators. By creating parallel-running subtasks based on the partitioning of the distributed dataset of the input data and distributing them to computing nodes for calculation, it realizes the operation and calculation of a single operator node on multiple servers, improving the calculation efficiency of spatial data under large data volumes.
[0079] Based on the first embodiment of this application, in the second embodiment of this application, for the same or similar content as in the above-mentioned first embodiment, reference can be made to the above introduction and will not be elaborated hereinafter. On this basis, the following steps may also be included before step S100:
[0080] Determine the deployment mode and build the distributed computing cluster according to the deployment mode.
[0081] In this embodiment, a distributed computing cluster can be built based on the Spark framework to provide an operating environment for operators. The Spark cluster can be configured based on the Standalone mode and the YARN mode. After successfully configuring and starting the Spark cluster, data processing and analysis can be performed by submitting a Spark application to the cluster.
[0082] In this embodiment, the distributed computing framework Spark can also be replaced with Apache Tez, and the scheduling engine adapts to the task submission mode of the Tez Java API (using TezClient). The scheduling engine parses the operator orchestration and converts it into a Tez DAG object to define the data processing flow.
[0083] Create and encapsulate operators, and save the encapsulated operators to the operator library.
[0084] In this embodiment, operators are defined to process data. The development specifications of the operators comply with the development specification requirements of the Spark framework and can be developed based on the Java or Scala language. When creating an operator, paste the operator code into the code editing area for saving and publishing the operator. The way of encapsulating the operator for reading and writing two-dimensional, three-dimensional spatio-temporal data and file type data into a general data reading and writing operator is put on the operator library and published. Here, it is not supported to introduce the spark-streaming module to process real-time stream data because, normally, the spark-streaming module will not actively exit after running. It is a resident memory program and is not suitable as an operator node in operator orchestration. Otherwise, the operator orchestration will be blocked in the running state for a long time and cannot complete the operation. After creating an operator, operations such as editing, deleting, and publishing can be performed on the operator, and it can be used in the operator orchestration layout after publishing.
[0085] In this embodiment, for the spatio-temporal data reading and writing operator, its reading and writing logic is extracted and encapsulated into a general data reading and writing operator, which supports reading types such as Shapefile, FileGDB, GeoJSON, PostGIS, files, etc. For string types, they are implemented in the form of startup parameters for reading input and do not need to be separately encapsulated into a general data reading and writing operator; The operators for reading Shapefile, FileGDB, and GeoJSON files use GeoTools tools for reading and parsing, convert them into Spark RDD (Resilient Distributed Dataset), and output this RDD as a parameter, where the spatial data is parsed into the WKB format; For reading spatio-temporal data in PostGIS, the Spark SQL module and GeoTools tools are used for reading and converting it into a Spark RDD, and this RDD is output as a parameter, where the spatial data is parsed into the WKB format. In subsequent operators, the analysis and calculation of spatial data are based on the Spark RDD.
[0086] In this embodiment, the general spatio-temporal data reading and writing operator is designed as follows:
[0087] For the operator for reading FileGDB files, input parameter: FileGDB file address (local directory / shared directory / remote address). Processing logic: Use tools such as GeoTools to parse the FileGDB file, read the spatial data and attribute set, read the spatial field in the WKB manner, and convert the data set into a Spark RDD. Output parameter: Spark RDD (distributed data set).
[0088] For the operator for reading Shapefile files, input parameter: Shapefile file address (local directory / shared directory / remote address). Processing logic: Use tools such as GeoTools to parse the Shapefile file, read the spatial data and attribute set, read the spatial field in the WKB manner, and convert the data set into a Spark RDD. Output parameter: Spark RDD.
[0089] For the operator for reading GeoJSON files, input parameter: GeoJSON file address (local directory / shared directory / remote address). Processing logic: Use tools such as GeoTools to parse the GeoJSON file, read the spatial data and attribute set, read the spatial field in the WKB manner, and convert the data set into a Spark RDD. Output parameter: Spark RDD (distributed data set).
[0090] For the PostGIS reading operator, input parameters: A. PostGIS database data source information; B. Name of the database spatial table to be read. Processing logic: Use the Spark SQL module to access the PostGIS library via JDBC to read the data and property set of the specified spatial table. The spatial fields are read in the WKB format, and the data set is converted into a Spark RDD. Output parameter: Spark RDD (distributed data set).
[0091] For the file reading operator, input parameter: File address (local directory / shared directory / remote address). Processing logic: Parse the file in text format and convert the parsed text data into a Spark RDD. Output parameter: Spark RDD (distributed data set).
[0092] For the FileGDB file writing operator, input parameters: Spark RDD (including spatio-temporal data), file save address, file name. Processing logic: Use tools such as GeoTools to convert the RDD data set and generate a FileGDB file. Output parameter: FileGDB file
[0093] For the Shapefile writing operator, input parameters: Spark RDD (including spatio-temporal data), file save address, file name. Use tools such as GeoTools to convert the RDD data set and generate a Shapefile file. Output parameter: Shapefile file
[0094] For the GeoJSON file writing operator, input parameters: Spark RDD (including spatio-temporal data), file save address, file name. Processing logic: Use tools such as GeoTools to convert the RDD data set and generate a GeoJSON file. Output parameter: GeoJSON file.
[0095] For the PostGIS writing operator, input parameters: A. Spark RDD (including spatio-temporal data); B. PostGIS data source information; C. Name of the table to be saved. Processing logic: Use the Spark SQL module to save the RDD data set into the specified PostGIS spatio-temporal database table. Output parameter: PostGIS spatio-temporal database table.
[0096] For the file writing operator, input parameters: A. Spark RDD (including spatio-temporal data); B. File save address; C. File name. Processing logic: Convert the RDD data set into text format data and save it to the specified file. Output parameter: File.
[0097] In this embodiment, the operator uniformly performs data reading and writing operations through an embedded normalized data reading and writing layer to support the transfer of parameter type data such as files, strings, two-dimensional and three-dimensional data, etc., and the parameter type transfer is more supported. This solves the problem that only conventional string parameter transfer is supported between operators, and file and two-dimensional and three-dimensional spatio-temporal data type parameter transfer cannot be supported.
[0098] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 4 , before step S100, steps B100 to B500 may further be included:
[0099] Step B100, if it is detected that the user drags out the first operator onto the canvas, obtain the identifier of the first operator.
[0100] In this embodiment, in the front-end code of the operator orchestration page, using the event detection mechanism of JavaScript, a drag event handler is added to each operator element in the operator library. For example, addEventListener('dragstart', function(event) { / * Record the drag start information * / }) is used to detect the drag start event, and addEventListener('drop', function(event) { / * Process the drag end event * / }) is used to detect the drag end event. When it is detected that the user drags an operator from the operator library onto the canvas, that is, when the drop event is triggered, the unique identifier (such as the name, ID, etc.) of the dragged operator and other related attributes (such as the category it belongs to, etc.) are obtained through the event object. For example, assuming that each operator element has a data-operator-type attribute to represent its type, the type information of the dragged operator is obtained through event.target.dataset.operatorType.
[0101] Step B200, obtain the associated operator of the first operator according to the identifier.
[0102] In this embodiment, the obtained operator type information is used as a query condition to search in the established operator association rule library. For example, if the "Read Shapefile File Operator" is dragged out, then search for the record with the "Read Shapefile File Operator" as the key in the adjacency list. According to the retrieval result, extract the information of all subsequent operators associated with this operator. This information may include the name of the operator, function description, input and output data types, etc. For example, obtain the "Spatial Intersection Operator", "Buffer Operator", "Spatial Query Operator" and their related attribute information associated with the "Read Shapefile File Operator" from the adjacency list.
[0103] Step B300, obtain the comprehensive prediction score of the associated operator according to the usage frequency, association degree and corresponding weight of the associated operator.
[0104] In this embodiment, the prediction score is determined by comprehensively considering the usage frequency of the operator and the association degree with the current operator. The usage frequency can be measured by counting the total number of times all users use this operator; the association degree can be determined according to the association weight defined in the association rule library.
[0105] For each associated operator, calculate the scores corresponding to each factor respectively. For example, for the "Spatial Intersection Operator", assume that its usage frequency score among all users accounts for 40% of the total score, and the score calculated according to the proportion of its usage times in the total usage times of all operators is 0.3; the association degree score with the "Read Shapefile File Operator" accounts for 30% of the total score, and the score calculated according to the weight in the association rule library is 0.25. Weighted sum the scores of each factor according to the set weight to obtain the comprehensive prediction score of each associated operator.
[0106] Step B400, sort the associated operators in descending order according to the comprehensive prediction score and store them in the recommended operator list.
[0107] Step B500, display the recommended operator list in the recommendation area.
[0108] In this embodiment, according to the calculated comprehensive prediction score, sort the associated operator list in descending order, sort the comprehensive prediction scores from high to low, so that the operator with the highest score is ranked first in the list.
[0109] Design a prediction operator recommendation area at an appropriate position on the operator orchestration canvas (such as the sidebar, pop-up window, etc.). Display the sorted list of associated operators in the designed recommendation area. For each operator, display its name, brief function description, and possible operation buttons (such as a button to directly drag it to the canvas). Continue to use the front-end event detection mechanism to monitor the user's operation behavior in real time after the prediction operators are displayed. For example, detect whether the user clicks the drag button of the prediction operator, or whether other operators are dragged out from the operator library.
[0110] In this embodiment, determine the dependency relationship between operators according to the operator orchestration task. And represent the dependency relationship between operators by a directed graph, where nodes represent operators and edges represent the data flow direction. Use a graph data structure (such as an adjacency list, adjacency matrix, etc.) or a graphics library to generate an operator orchestration graph. Through the visual operator orchestration page, combine and connect different operators to output the complete process to construct a directed acyclic graph. Exemplarily, please refer to Figure 5 , the specific operator orchestration process is as follows: Step (1), on the operator orchestration page, drag out the "Read Shapefile File Operator" and "Read PostGIS Operator" from the operator library, which are respectively used to read two spatio-temporal data for spatial intersection calculation. Step (2), drag out the "Spatial Intersection Operator" from the operator library to the canvas and connect it to the above two spatio-temporal data reading operators as the input parameters of the spatial intersection operator. Step (3), drag out the "Write Shapefile File Operator" from the operator library and connect it to the "Spatial Intersection Operator", take the output result of the "Spatial Intersection Operator" as the input parameter of this operator, and fill in the file save address and file name, that is, complete the DAG operator orchestration process of spatial intersection calculation.
[0111] In a feasible implementation manner, before step B100, it may further include:
[0112] Establish an operator association rule library to store the successor operators of each of the first operators and the corresponding association degrees.
[0113] When it is detected that the user selects a prediction operator in the recommendation area, update the association degree between the prediction operator and the first operator.
[0114] In this embodiment, the correlation degree between operators is determined according to the function of each operator, the input / output data format, and the applicable business scenario. For example, the "Read Shapefile File Operator" is used to read spatial data from a Shapefile file, and its output is a spatial data object in a specific format. Through analysis, it can be seen that subsequent operators for processing spatial data may be connected, such as spatial analysis, transformation, and storage-related operators. The association rules are stored in the form of an adjacency list. An adjacency list is a storage structure of a graph. For each operator, its possible successor operators are recorded in the table. For example, for the "Read Shapefile File Operator", the "Spatial Intersection Operator", "Buffer Operator", "Spatial Query Operator", etc. associated with it are recorded in the adjacency list.
[0115] In this embodiment, based on the operators that the user has currently dragged out, the subsequent operators that may be used are quickly recommended, reducing the time for the user to search and filter in the operator library, enabling the user to complete the operator orchestration task more quickly, and improving the overall operation efficiency. Based on the internal association and actual usage of operators, reasonable operation suggestions are provided for the user, reducing the user's learning cost and operation difficulty. In addition, by considering the usage frequency, correlation degree, and corresponding weights of associated operators to calculate the comprehensive prediction score, it is possible to more reasonably guide the user to use operators, avoid overusing some infrequently used or unnecessary operators, and thus optimize the allocation and utilization of system resources.
[0116] Based on the first embodiment of this application, in the fourth embodiment of this application, the same or similar content as the above-mentioned Embodiment 1 can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 6 , step S300 may include steps S310 to S330:
[0117] Step S310, based on the master node of the distributed cluster, divide the executable file into task stages to obtain each task stage and the execution order of the task stage;
[0118] In this embodiment, after the JAR package containing the application program is passed to the master node of the Spark cluster, the master node decompresses and checks the content of the JAR package. Find the class containing the main method, which is usually the entry point of the application program and defines the main logic of the task. For example, in a Java application, the master node will find the main class according to the --class parameter specified in the spark-submit command or the manifest file of the JAR package (if configured). This main class contains the overall task description of the application program, such as which operator orchestration operations to execute, which data processing logics to use, and the final result output, etc.
[0119] Step S320: Create and partition the distributed dataset of the input data according to the execution order and the data source path of the executable file, and the partitioning strategy is the spatial dimension and / or the temporal dimension.
[0120] Step S330: Generate the subtasks to be executed in parallel based on the partitioning result and the cluster configuration.
[0121] In this embodiment, the master node determines the stages of the task based on the dependency relationship and execution order information of the operator orchestration, such as the data reading stage, the data processing stage, the result storage stage, etc., to ensure that the task proceeds in the correct logical order and avoid data dependency conflicts. The master node creates subtasks (Tasks) to run in parallel according to the partitioning of the input data RDD. First, according to the code in the JAR package and the data source path, the Spark cluster uses the corresponding data reading logic (such as reading data from a file system, a database, or other data sources) to create the input data RDD. The number of partitions is dynamically selected according to the size of the input data and the cluster configuration. For example, for a large text file, Spark may divide it into multiple data blocks, and each data block serves as a partition to form a distributed dataset (RDD). Each partition will be the input of a subtask. For example, if the application uses sc.textFile("path / to / input / file") to read a file, the master node will use this code snippet to read the file data into the RDD according to the file location and the Spark file reading mechanism. At the same time, the number of partitions is automatically determined according to the file size and the cluster configuration (such as parameters like spark.default.parallelism) to create a distributed dataset. After the input data RDD is created, subsequent operators (such as data processing operators) will read this RDD for operations. For example, a map or filter operator will read each partition of the RDD and perform corresponding calculation operations on the data within the partition. This reading process is distributed, and each computing node (Executor) will read the RDD partition assigned to it and execute the computing task.
[0122] In this embodiment, partitioning is performed according to the partitioning strategy based on spatial dimension and / or time dimension. In the partitioning strategy based on spatial range, the geographical space is divided into multiple sub-regions, and each sub-region corresponds to a partition. For example, for global map data, it can be divided according to the longitude and latitude range, such as every 10 degrees of longitude and latitude as a partition. It can also be divided according to administrative regions, such as partitioning by country, province, city, etc. When the query is mainly based on a specific spatial region, it can quickly locate the relevant partition and reduce unnecessary data scanning. For example, in an urban planning project, if you want to analyze the population density changes in different regions, the urban map can be partitioned according to blocks or administrative divisions. Each partition stores the population data and related geographical information within that region. When querying the population density of a certain block, directly obtain the data from the corresponding partition for calculation without traversing the data of the entire city.
[0123] In this embodiment, in the partitioning strategy based on time range, the data is divided according to the time dimension. For example, the spatio-temporal data is allocated to different partitions at time intervals such as days, months, years, etc. For example, for meteorological monitoring data, the data is stored with each day as a partition. When analyzing the meteorological data of a certain week, only the data of the corresponding 7 partitions needs to be read, improving the query efficiency.
[0124] In this embodiment, in the hybrid partitioning strategy based on spatial grid and time slice, the spatial and time dimensions are combined for partitioning. First, the space is divided into grid-like sub-regions, and then time slicing partitioning is performed within each sub-region. The hybrid strategy can take into account the query requirements of both spatial and time dimensions and improve the query and calculation efficiency. In a marine ecological monitoring project, the ocean area is divided into several grids, and each grid stores data such as ocean temperature and salinity by month. When studying the ecological changes in a certain sea area during a specific time period, the corresponding spatial grid and time slice partition can be quickly located to obtain the required data.
[0125] Based on the first embodiment of the present application, in the fifth embodiment of the present application, the same or similar content as in the above-mentioned embodiment one can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 6 , step S400 may include steps S410 to S420:
[0126] Step S410, sending the sub-task to the corresponding computing node according to the task distribution algorithm;
[0127] Step S420, after receiving the task, the computing node calls the operator to be executed for data processing and outputs the execution result.
[0128] In this embodiment, the master node first serializes the subtasks into a format suitable for network transmission, such as JSON, XML, or a custom binary format. The serialized tasks are sent to the corresponding computing nodes through the communication mechanism provided by the distributed computing framework. After receiving the tasks, the computing nodes will send an acknowledgment message to the master node, indicating that the tasks have been successfully received and are ready to be executed. Then, the tasks are deserialized from the serialized format into an internal representation for execution. According to the operator information in the tasks, the computing nodes call the corresponding operator functions or modules to process the data. During the execution process, the computing nodes can store the intermediate results on the local disk or in memory for subsequent processing or data exchange with other nodes. When the operator execution is completed, the computing nodes generate output data and prepare to send it to the master node or for further processing. When all subtasks are executed, the master node sends a request to the computing nodes to collect the output data. The computing nodes send the locally stored output data to the master node, which is returned to the user through the user interface or storage interface or stored at the specified location.
[0129] In a feasible implementation, task distribution can be performed according to the load balancing algorithm, and the tasks are assigned to the nodes with lower load to ensure the overall performance of the system. The load balancing algorithm can measure the load of the nodes according to different metrics, such as CPU usage, memory occupancy, network bandwidth, etc. By regularly monitoring these metrics and adjusting the allocation strategy, it is ensured that each computing node maintains a relatively balanced load level.
[0130] In a feasible implementation, task distribution can be performed according to the distributed hash table (DHT) algorithm. The tasks are regarded as data items, and a hash function is used to calculate the hash value of each task, and then the tasks are assigned to the corresponding nodes according to the hash value.
[0131] Based on the first embodiment of this application, in the sixth embodiment of this application, the same or similar content as in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 7 , steps A100 to A300 may also be included before step S400:
[0132] Step A100, create a monitoring linked list and store the running tasks in the monitoring linked list.
[0133] Step A200, perform a traversal operation on the monitoring linked list to obtain the running status of the running tasks.
[0134] Step A300, when the running status is running completed, generate task completion information and delete the corresponding running tasks from the monitoring linked list.
[0135] Please refer to Figure 8, in this embodiment, the running state of the operator can be continuously monitored based on the status monitoring module, and the running state is fed back to the scheduling engine to facilitate the scheduling engine to complete the status maintenance operation of the operator. Create a monitoring linked list in the status monitoring module, put the running tasks into the end of the linked list queue, and associate the Spark Application ID and Job ID. Traverse the monitoring linked list every certain period of time and request the Spark REST API interface / api / v1 / applications / <application-id> / jobs / <job-id>Obtain the running status of the task job Job. When the running is completed, push the running result to the scheduling engine, and remove the task Job from the monitoring linked list. The scheduling engine marks the operator as completed and starts scheduling and running the next operator until the entire operator orchestration is completed.
[0136] In this embodiment, a monitoring linked list is created to store the currently running tasks. Using the Application ID and Job ID of the task as parameters, a new task instance is created and then added to the linked list. Periodically traverse the tasks in the monitoring linked list. For each task, use the Spark REST API to query its status. When checking the task status, take corresponding actions according to the returned status values (such as "SUCCEEDED", "RUNNING", "FAILED", etc.). If the task is successfully completed ("SUCCEEDED"), call the interface of the scheduling engine to mark the task as completed and remove the task from the monitoring linked list. If the task fails or is still running, continue to monitor.
[0137] Through the above steps, a status monitoring module can be implemented to track the running status of Spark task jobs and push the results to the scheduling engine when the tasks are completed, thus supporting complex job scheduling and orchestration.
[0138] In this embodiment, the status monitoring module can also be removed, and the status callback can be used in the way of callback after the Spark task runs to completion. The status detection is implemented using the SparkListener class at the start of the task node. When the task runs to completion, call the status feedback interface of the scheduling engine to modify the status of the task node in the scheduling engine. The disadvantage is that this implementation is invasive to the task node and the callback logic code needs to be added during the development of the task node; if the task node reports an exception error, the program terminates to avoid the inability to feedback the running status to the scheduling engine, resulting in the status of the task node in the scheduling engine remaining in the running state and unable to end.
[0139] This application provides an operator orchestration device for spatio-temporal data analysis and calculation, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the spatio-temporal data analysis and calculation operator orchestration method in the first embodiment above.
[0140] Next, refer to Figure 9 , which shows a schematic structural diagram of an operator orchestration device suitable for implementing spatio-temporal data analysis and calculation according to embodiments of the present application. The operator orchestration device for spatio-temporal data analysis and calculation in embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The operator orchestration device for spatio-temporal data analysis and calculation shown is merely an example and should not impose any limitations on the functions and usage scope of embodiments of the present application.
[0141] As Figure 9 shown, the operator orchestration device for spatio-temporal data analysis and calculation may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the operator orchestration device for spatio-temporal data analysis and calculation are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the operator orchestration device for spatio-temporal data analysis and calculation to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an operator orchestration device for spatio-temporal data analysis and calculation with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.
[0142] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0143] The operator orchestration device for spatio-temporal data analysis and calculation provided by the present application adopts the operator orchestration method for spatio-temporal data analysis and calculation in the above embodiments, and can solve the technical problem of how to improve the operation efficiency of a single operator node to support large-scale data volume calculation. Compared with the prior art, the beneficial effects of the operator orchestration device for spatio-temporal data analysis and calculation provided by the present application are the same as those of the operator orchestration method for spatio-temporal data analysis and calculation provided by the above embodiments, and other technical features in the operator orchestration device for spatio-temporal data analysis and calculation are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.
[0144] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0145] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0146] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the operator orchestration method for spatio-temporal data analysis and calculation in the above embodiments.
[0147] The computer-readable storage medium provided by the present application may, for example, be a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.
[0148] The above computer-readable storage medium may be included in an operator orchestration device for spatio-temporal data analysis and calculation; or it may exist independently and not be assembled into the operator orchestration device for spatio-temporal data analysis and calculation. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed by the operator orchestration device for spatio-temporal data analysis and calculation, the operator orchestration device for spatio-temporal data analysis and calculation is caused to: after generating an operator orchestration graph, parse the operator orchestration graph to determine the operators to be executed and the dependencies between the operators to be executed; generate an executable file according to the code fragments corresponding to the operators to be executed and the dependencies; divide the running tasks corresponding to the executable file into multiple subtasks according to the partitions of the distributed dataset of the input data; allocate the subtasks to the respective corresponding computing nodes in the distributed cluster for parallel execution, and when each of the subtasks is completed, obtain an execution result.
[0149] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).
[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0151] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0152] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned operator orchestration method for spatio-temporal data analysis and calculation, and can solve the technical problem of how to improve the running efficiency of a single operator node to support large-scale data volume calculation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the operator orchestration method for spatio-temporal data analysis and calculation provided by the above embodiments, and will not be elaborated here.
[0153] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields is included in the patent protection scope of the present application.
Claims
1. An operator scheduling method for spatio-temporal data analysis and calculation, characterized in that, The method described above includes: After generating the operator orchestration graph, parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators; Generating an executable file according to the code snippet corresponding to the operator to be executed and the dependencies; According to the partitioning of the distributed dataset of the input data, dividing the running tasks corresponding to the executable file into multiple subtasks, including: based on the master node of the distributed cluster, dividing the executable file into task phases to obtain each task phase and the execution order of the task phases; according to the execution order and the data source path of the executable file, creating and partitioning the distributed dataset of the input data, and the partitioning strategy is the spatial dimension and / or the temporal dimension; generating parallel-executed subtasks based on the partitioning result and the cluster configuration, wherein the partitioning strategy based on the spatial dimension divides the geographical space into multiple sub-regions, each sub-region corresponds to a partition, the partitioning strategy based on the temporal dimension distributes the spatio-temporal data to different partitions according to time intervals, and the partitioning strategy based on the spatial dimension and the temporal dimension first divides the space into grid-like sub-regions, and then slices and partitions according to time within each sub-region; Assigning the subtasks to the corresponding computing nodes in the distributed cluster for parallel execution, and obtaining an execution result after each subtask is executed.
2. The operator scheduling method for spatio-temporal data analysis and calculation according to claim 1, wherein The operator orchestration graph is a directed acyclic graph, and the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators includes: Determining the starting node in the operator orchestration graph; Traversing all successor nodes of the starting node based on the depth-first search algorithm; Obtaining the dependencies according to the traversal order and the direction of the edges in the directed acyclic graph.
3. The operator scheduling method for spatio-temporal data analysis and calculation according to claim 1, characterized in that, Before the step of, after generating the operator orchestration graph, parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators, it further includes: If it is detected that the user drags an operator onto the canvas, obtaining the identifier of the operator; Obtaining the associated operators of the operator according to the identifier; Obtaining the comprehensive prediction score of the associated operator according to the usage frequency, association degree and corresponding weight of the associated operator; Sorting the associated operators in descending order according to the comprehensive prediction score and storing them in the recommended operator list; Displaying the recommended operator list in the recommendation area.
4. The operator scheduling method for spatio-temporal data analysis and calculation according to claim 3, wherein Before the step of, if it is detected that the user drags an operator onto the canvas, obtaining the identifier of the operator, it further includes: Establishing an operator association rule library, storing the successor operators of each operator and the corresponding association degrees; When it is detected that the user selects a predicted operator in the recommendation area, updating the association degree between the predicted operator and the operator.
5. The operator scheduling method for spatio-temporal data analysis and calculation according to claim 1, wherein Before the step of parsing the operator orchestration graph to determine the operators to be executed and the dependencies between the operators, it further includes: Determining the deployment mode and building the distributed cluster according to the deployment mode; Creating and encapsulating operators, and saving the encapsulated operators to the operator library.
6. The operator scheduling method for spatio-temporal data analysis and calculation according to claim 1, characterized in that Before the step of, assigning the subtasks to the corresponding computing nodes for execution and obtaining an execution result, it further includes: Create a monitoring linked list and store the running tasks in the monitoring linked list; Perform a traversal operation on the monitoring linked list to obtain the running status of the running tasks; When the running status is completed, generate task completion information and delete the corresponding running task from the monitoring linked list.
7. The operator scheduling method for spatio-temporal data analysis and calculation according to claim 1, wherein The step of allocating the subtasks to their respective corresponding computing nodes in the distributed cluster for parallel execution, and obtaining the execution results after each subtask is executed includes: Send the subtasks to the corresponding computing nodes according to the task distribution algorithm; After receiving the task, the computing node calls the operator to be executed for data processing and outputs the execution result.
8. An operator scheduling device for spatio-temporal data analysis and calculation, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the operator scheduling method for spatio-temporal data analysis and calculation according to any one of claims 1 to 7.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the operator scheduling method for spatio-temporal data analysis and calculation according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Model operator parallel splitting method and device, equipment and storage medium
CN119294463A