Data distribution method and device for distributed graph calculation

By adopting a unified Pull-based Shuffle mechanism in distributed graph computing, downstream nodes actively request data and upstream nodes distribute it on demand, solving the problem of high system complexity under different computing modes, and achieving flexible data distribution and efficient system adaptation.

CN120295802AActive Publication Date: 2025-07-11ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510787456.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In distributed graph computing, the existing technology needs to provide different data shuffle mechanisms for different computing modes, resulting in increased system complexity and excessive overhead, and it is impossible to flexibly adapt to different computing modes.

Method used

Provide a unified data distribution method. Through the Pull-based Shuffle mechanism, the downstream computing nodes actively send data acquisition requests, and the upstream computing nodes distribute data on demand, adapting to multiple computing modes.

Benefits of technology

It reduces the complexity of system development and maintenance, improves the flexibility of the system, and can efficiently distribute data in different computing modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295802A_ABST
    Figure CN120295802A_ABST
Patent Text Reader

Abstract

The invention discloses a data distribution method for distributed graph calculation. Computing nodes in a distributed system support execution of distributed graph calculation tasks according to multiple calculation modes. The distributed graph calculation task comprises a first subtask serving as an upstream task and a second subtask serving as a downstream task; the first subtask and the second subtask are both concurrently executed by a plurality of computing nodes; the method comprises the following steps: receiving a data acquisition request sent by any target downstream computing node in a plurality of downstream computing nodes for executing a second subtask; the data acquisition request comprises a data acquisition condition determined by the target downstream computing node on the basis of a computing mode adopted for executing the second subtask; different calculation modes adopted by the target downstream calculation node for executing the second subtask correspond to different data acquisition conditions respectively; acquiring target data meeting a data acquisition condition from data generated by executing the first subtask; and issuing the target data to the target downstream computing node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification belong to the field of graph computing, and in particular, relate to a data distribution method and apparatus for distributed graph computing. Background Art

[0002] In order to adapt to diverse business requirements, more and more graph computing engines begin to support executing graph computing tasks according to multiple different computing modes; for example, a graph computing engine can support executing graph computing tasks according to a stream processing mode, a batch processing mode, and a graph processing mode.

[0003] In a distributed graph computing scenario, no matter which mode the graph computing engine executes the graph computing in, it is inevitable to perform data shuffle between the upstream tasks and downstream tasks included in the distributed graph computing. The so-called data shuffle refers to the process in which the upstream tasks included in the distributed graph computing distribute data to the downstream tasks through the network.

[0004] However, in practical applications, when performing data shuffle between the upstream tasks and downstream tasks under different computing modes, the shuffle methods usually have certain differences, which will cause the graph computing engine to provide different data shuffle mechanisms for different computing modes, possibly resulting in an increase in the system complexity of the graph computing engine and an excessive system overhead. Summary of the Invention

[0005] This specification proposes a data distribution method for distributed graph computing. The computing nodes in the distributed system support executing distributed graph computing tasks according to multiple computing modes; the distributed graph computing tasks include a first subtask as an upstream task and a second subtask as a downstream task; wherein, both the first subtask and the second subtask are concurrently executed by multiple computing nodes in the distributed system; the method is applied to any target upstream computing node among the multiple upstream computing nodes executing the first subtask, and the method includes: Receive a data acquisition request sent by any target downstream computing node among multiple downstream computing nodes for executing the second subtask; wherein, the data acquisition request includes data acquisition conditions determined by the target downstream computing node based on the computing mode adopted for executing the second subtask; the data acquisition conditions are used to acquire input data required for the target downstream computing node to execute the second subtask from the data generated by the target upstream computing node executing the first subtask; different computing modes adopted by the target downstream computing node for executing the second subtask respectively correspond to different data acquisition conditions; In response to the data acquisition request, acquire target data that meets the data acquisition conditions from the data generated by executing the first subtask; Send the target data to the target downstream computing node, so that the target downstream computing node continues to execute the second subtask with the target data as input data.

[0006] Optionally, the data generated by executing the first subtask includes data shards respectively allocated to multiple downstream computing nodes for executing the second subtask; the data shards include at least one batch of data; any batch of data includes at least one data block; Wherein, the number of batches of data included in the data shard allocated to any downstream computing node is the number corresponding to the computing mode adopted by the downstream computing node for executing the second subtask.

[0007] Optionally, the data acquisition conditions include the shard identifier of the data shard allocated by the target upstream computing node to the target downstream computing node, the batch identifier of the starting data batch to be acquired, and the number of data batches to be acquired; In response to the data acquisition request, acquiring target data that meets the data acquisition conditions from the data generated by executing the first subtask includes: In response to the data acquisition request, determine a target data shard corresponding to the shard identifier from the data shards respectively allocated to the multiple downstream computing nodes; From at least one batch of data included in the target data shard, determine a target data batch corresponding to the batch identifier, and use this target data batch as the starting data batch, and sequentially acquire at least one batch of data corresponding to the value of the number of data batches as the target data.

[0008] Optionally, the multiple computing modes include a stream processing mode, a batch processing mode, and an iterative graph computing mode; If the computing mode adopted by the target downstream computing node to execute the second subtask is the stream processing mode, the target data slice contains data in multiple batches, the batch identifier of the starting data batch is the batch identifier corresponding to the data of the first batch among the data in the multiple batches, and the value of the number of batches is infinite; If the computing mode adopted by the target downstream computing node to execute the second subtask is the batch processing mode, the target data slice contains data in only one batch, the batch identifier of the starting data batch is the batch identifier corresponding to the data in the only one batch, and the value of the number of batches is 1; Among them, the batch identifiers of the data in the multiple batches increase monotonically.

[0009] Optionally, the iterative graph computing mode includes an iterative graph computing mode executed in the manner of stream processing; and, an iterative graph computing mode executed in the manner of batch processing; If the computing mode adopted by the target downstream computing node to execute the second subtask is the iterative graph computing mode executed in the manner of stream processing, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data slice contains data in multiple batches; the batch identifier of the starting data batch is the batch identifier corresponding to the data of the first batch among the data in the multiple batches; the value of the number of batches is infinite; If the computing mode adopted by the target downstream computing node to execute the second subtask is the iterative graph computing mode executed in the manner of batch processing, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data slice contains data in only one batch, the batch identifier of the starting data batch is the batch identifier corresponding to the data in the only one batch, and the value of the number of batches is 1; Among them, the batch identifiers of the data in the multiple batches increase monotonically.

[0010] Optionally, the data acquisition condition includes the slice identifier of the data slice allocated by the target upstream computing node to the target downstream computing node, and the maximum batch identifier of all data batches that need to be acquired; In response to the data acquisition request, acquiring target data that meets the data acquisition condition from the data generated by executing the first subtask includes: In response to the data acquisition request, determining a target data slice corresponding to the slice identifier from the data slices respectively allocated to the multiple downstream computing nodes; From at least one batch of data included in the target data shard, obtain at least one batch of data whose batch identifier is not greater than the maximum batch identifier as the target data.

[0011] Optionally, the multiple computing modes include a stream processing mode, a batch processing mode, and an iterative graph computing mode; If the computing mode adopted by the target downstream computing node to execute the second subtask is the stream processing mode, the target data shard contains multiple batches of data, and the value of the maximum batch identifier is infinity; If the computing mode adopted by the target downstream computing node to execute the second subtask is the batch processing mode, the target data shard contains only one batch of data, and the value of the maximum batch identifier is the batch identifier corresponding to the only batch of data; Among them, the batch identifiers of the multiple batches of data increase monotonically.

[0012] Optionally, the iterative graph computing mode includes an iterative graph computing mode executed in a stream processing manner; and an iterative graph computing mode executed in a batch processing manner; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in a stream processing manner, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data shard contains multiple batches of data, and the value of the maximum batch identifier is infinity; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in a batch processing manner, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data shard contains only one batch of data, and the value of the maximum batch identifier is the batch identifier corresponding to the only batch of data; Among them, the batch identifiers of the multiple batches of data increase monotonically.

[0013] Optionally, the data acquisition request is an asynchronous request; In response to the data acquisition request, obtaining target data that meets the data acquisition condition from the data generated by executing the first subtask includes: Asynchronously respond to the data acquisition request, and obtain target data that meets the data acquisition condition from the data generated by executing the first subtask.

[0014] Optionally, the target data includes at least one batch of data obtained from the data shards allocated for the target downstream computing node; Sending the target data to the target downstream computing node includes: Sending the data blocks included in the at least one batch of data obtained to the target downstream computing node in a streaming manner.

[0015] Optionally, multiple downstream computing nodes for executing the second subtask execute the second subtask using different computing modes.

[0016] This specification also proposes a data distribution method for distributed graph computing. The computing nodes in the distributed system support executing distributed graph computing tasks according to multiple computing modes; the distributed graph computing task includes a first subtask as an upstream task and a second subtask as a downstream task; wherein, both the first subtask and the second subtask are concurrently executed by multiple computing nodes in the distributed system; the method is applied to any target downstream computing node among the multiple downstream computing nodes for executing the second subtask, and the method includes: Determining a data acquisition condition based on the computing mode used to execute the second subtask; wherein, the data acquisition condition is used to acquire the input data required to execute the second subtask from any target upstream computing node among the multiple upstream computing nodes for executing the first subtask; Generating a data acquisition request including the data acquisition condition and sending the data acquisition request to the target upstream computing node, so that the target upstream computing node acquires target data that meets the data acquisition condition from the data generated by executing the first subtask; Receiving the target data sent by the target upstream computing node and using the target data as input data to continue executing the second subtask.

[0017] Optionally, the data generated by executing the first subtask includes data shards allocated for multiple downstream computing nodes for executing the second subtask respectively; wherein, the data shards include at least one batch of data; any one batch of data includes at least one data block; the number of batches of data included in the data shards allocated for any downstream computing node is the number corresponding to the computing mode used by this downstream computing node to execute the second subtask.

[0018] Optionally, the multiple computing modes include a stream processing mode, a batch processing mode, and an iterative graph computing mode; The data acquisition conditions include the shard identifier of the data shard allocated by the target upstream computing node to the target downstream computing node, the batch identifier of the starting data batch to be acquired, and the number of data batches to be acquired; If the computing mode adopted by the target downstream computing node to execute the second subtask is the stream processing mode, the target data shard contains data of multiple batches, the batch identifier of the starting data batch is the batch identifier corresponding to the data of the first batch among the multiple batches of data, and the value of the number of batches is infinite; If the computing mode adopted by the target downstream computing node to execute the second subtask is the batch processing mode, the target data shard contains data of only one batch, the batch identifier of the starting data batch is the batch identifier corresponding to the data of the only batch, and the value of the number of batches is 1; Among them, the batch identifiers of the multiple batches of data increase monotonically.

[0019] Optionally, the iterative graph computing mode includes an iterative graph computing mode executed in the manner of stream processing; and, an iterative graph computing mode executed in the manner of batch processing; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in the manner of stream processing, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data shard contains data of multiple batches; the batch identifier of the starting data batch is the batch identifier corresponding to the data of the first batch among the multiple batches of data; the value of the number of batches is infinite; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in the manner of batch processing, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data shard contains data of only one batch, the batch identifier of the starting data batch is the batch identifier corresponding to the data of the only batch, and the value of the number of batches is 1.

[0020] Among them, the batch identifiers of the multiple batches of data increase monotonically.

[0021] Optionally, the multiple computing modes include the stream processing mode, the batch processing mode, and the iterative graph computing mode; The data acquisition conditions include the shard identifier of the data shard allocated by the target upstream computing node for the target downstream computing node, and the maximum batch identifier of all data batches to be acquired; If the computing mode adopted by the target downstream computing node to execute the second subtask is a stream processing mode, the target data shard contains data of multiple batches, and the value of the maximum batch identifier is infinity; If the computing mode adopted by the target downstream computing node to execute the second subtask is a batch processing mode, the target data shard contains data of only one batch, and the value of the maximum batch identifier is the batch identifier corresponding to the data of the only one batch; Wherein, the batch identifiers of the data of the multiple batches increase monotonically.

[0022] Optionally, the iterative graph computing mode includes an iterative graph computing mode executed in a stream processing manner; and, an iterative graph computing mode executed in a batch processing manner; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in a stream processing manner, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data shard contains data of multiple batches, and the value of the maximum batch identifier is infinity; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in a batch processing manner, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data shard contains data of only one batch, and the value of the maximum batch identifier is the batch identifier corresponding to the data of the only one batch; Wherein, the batch identifiers of the data of the multiple batches increase monotonically.

[0023] Optionally, the data acquisition request is an asynchronous request.

[0024] Optionally, the target data includes at least one batch of data obtained from the data shard allocated by the target upstream computing node for the target downstream computing node; Receiving the target data sent by the target upstream computing node includes: Receiving the data blocks included in the at least one batch of data sent by the target upstream computing node in a streaming transmission manner.

[0025] Optionally, multiple downstream computing nodes for executing the second subtask execute the second subtask in different computing modes.

[0026] In the above embodiments, a unified data distribution method can be provided for multiple computing modes supported by a distributed system for executing distributed graph computing. For different computing modes, there is no need to separately provide different data distribution methods, thereby reducing the complexity of system development and maintenance and enabling the system to more flexibly adapt to different computing modes. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0028] Figure 1 It is a schematic diagram showing data distribution from an upstream task to a downstream task under different computing modes in an embodiment of this specification; Figure 2 It is a flowchart of a data distribution method for distributed graph computing shown in an embodiment of this specification; Figure 3 It is a schematic diagram of an upstream task and a downstream task shown in an embodiment of this specification; Figure 4 It is a schematic diagram showing concurrent execution of an upstream task and a downstream task in an embodiment of this specification; Figure 5 It is a schematic diagram showing concurrent execution of an upstream task and a downstream task in an embodiment of this specification; Figure 6 It is a data structure diagram of data generated by an upstream computing node executing a first subtask shown in an embodiment of this specification; Figure 7 It is a flowchart of another data distribution method for distributed graph computing shown in an embodiment of this specification; Figure 8 It is a schematic structural diagram of an electronic device shown in an embodiment of this specification; Figure 9 It is a block diagram of a data distribution device for distributed graph computing shown in an embodiment of this specification; Figure 10 It is a block diagram of a data distribution device for distributed graph computing shown in an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without making creative efforts shall fall within the scope of protection of this specification.

[0030] In a distributed graph computing scenario, for a graph computing engine that supports multiple computing modes, in different computing modes, if data distribution (shuffle) needs to be performed between the upstream tasks and downstream tasks included in the executed distributed graph computing, different data distribution methods are usually adopted.

[0031] Please refer to Figure 1 , Figure 1 a schematic diagram showing data distribution from an upstream task to a downstream task in different computing modes as shown in this specification.

[0032] As Figure 1 shown, if the computing mode adopted by the graph computing engine is the batch processing mode, since the batch processing mode processes a large amount of data and usually has a relatively low requirement for data latency, in this computing mode, after the upstream task finishes execution, the generated batch of data can usually be stored locally or in external storage. After the downstream task starts, it can send a request to the upstream task to actively pull the batch of data from the above-mentioned local or external storage as input data. This data distribution method is usually referred to as Pull-based Shuffle, that is, the downstream task actively pulls the required large amount of data from the upstream task.

[0033] As Figure 1 shown, if the computing mode adopted by the graph computing engine is the stream processing mode, since the stream processing mode has a high requirement for data latency and processes a small amount of data, in this computing mode, after the upstream task finishes execution, the generated data can usually be divided into several small batches of data, and then actively pushed to the downstream task batch by batch. This data distribution method is usually referred to as Push-based Shuffle, that is, the upstream task actively pushes small batches of data to the downstream task.

[0034] As Figure 1As shown, if the computing mode adopted by the graph computing engine is the graph computing mode, since the graph computing mode is relatively special, after the computing nodes complete the graph computing in one round of iteration, the generated data will be sent back to themselves as the input for the graph computing in the next round of iteration. Based on this characteristic, in the graph computing mode, data distribution can be carried out in the Pull-based Shuffle manner or in the Push-based Shuffle manner.

[0035] For example, please continue to refer to Figure 1 , for the graph computing scenario with a relatively small amount of data to be processed, the iterative graph computing can be executed in the streaming processing manner, and the data generated by the iterative computing in the previous round can be distributed to the iterative computing in the next round as input data in the Push-based Shuffle manner. For the graph computing scenario with a relatively large amount of data to be processed, the iterative graph computing can be executed in the batch processing manner, and the data generated by the iterative computing in the previous round can be distributed to the iterative computing in the next round as input data in the Pull-based Shuffle manner.

[0036] In different computing modes, when shuffling data between the upstream task and the downstream task, there are usually certain differences in the adopted shuffle methods, which will cause the graph computing engine to provide different data shuffle mechanisms for different computing modes.

[0037] It can be seen that when the distributed system executes distributed graph computing, although it supports executing graph computing according to multiple computing modes, since different data distribution methods are usually adopted when distributing data between the upstream task and the downstream task, this will cause the graph computing engine to provide different data shuffle mechanisms for different computing modes, which may lead to problems such as increased system complexity and excessive overhead. Moreover, it may also lead to the problem that the graph computing engine cannot flexibly adapt to different computing modes.

[0038] Based on this, this specification proposes a technical solution that provides a unified data distribution method for multiple computing modes supported by a distributed system for executing distributed graph computing.

[0039] Please refer to Figure 2 , Figure 2Flowchart of a data distribution method for distributed graph computing shown in this specification; computing nodes in a distributed system support executing distributed graph computing tasks according to multiple computing modes; the distributed graph computing task includes a first subtask as an upstream task and a second subtask as a downstream task; wherein, both the first subtask and the second subtask are concurrently executed by multiple computing nodes in the distributed system; the method is applied to any target upstream computing node among the multiple upstream computing nodes executing the first subtask; the method includes the following execution process: Step 202, receive a data acquisition request sent by any target downstream computing node among the multiple downstream computing nodes for executing the second subtask; wherein, the data acquisition request includes data acquisition conditions determined by the target downstream computing node based on the computing mode adopted for executing the second subtask; the data acquisition conditions are used to acquire the input data required for the target downstream computing node to execute the second subtask from the data generated by the target upstream computing node when executing the first subtask; different computing modes adopted by the target downstream computing node for executing the second subtask respectively correspond to different data acquisition conditions; The computing nodes in the above-mentioned distributed system can specifically support executing distributed graph computing tasks according to multiple computing modes.

[0040] For example, in practical applications, a graph computing engine that supports executing distributed graph computing tasks according to multiple computing modes can be deployed on the computing nodes in the distributed system, so that the computing nodes in the distributed system can also have the ability to execute distributed graph computing tasks according to multiple computing modes.

[0041] It should be noted that, in practical applications, the computing nodes in the above-mentioned distributed system can specifically be an independent node device in a computing cluster, or a thread running on the node device for executing graph computing tasks.

[0042] The above-mentioned distributed graph computing task can specifically include a first subtask as an upstream task and a second subtask as a downstream task; In practical applications, the above-mentioned distributed graph computing tasks usually include multiple subtasks with data dependencies. Among these subtasks, if the data generated after a certain target subtask is executed (i.e., the output data of this subtask) will continue to be used as the input data of another subtask, then this target subtask can usually be used as an upstream task, and this other subtask can usually be used as the downstream task corresponding to this upstream task.

[0043] Among them, it should be noted that the specific types of the subtasks that are the upstream tasks and downstream tasks included in the above-mentioned distributed graph computing tasks are not specifically limited in this specification. In actual applications, the types of the subtasks that are the upstream tasks and downstream tasks can be the same or different.

[0044] For example, please refer to Figure 3 , in one example, as a subtask of the upstream task, specifically it can be a map task included in the graph computing task; as a subtask of the downstream task, specifically it can be a sink task included in the graph computing task; in another example, as a subtask of the upstream task, specifically it can be a souce task included in the graph computing task, and as a subtask of the downstream task, specifically it can be an iterative graph task included in the graph computing task; or, as a subtask of the upstream task, specifically it can be an iterative graph task included in the graph computing task, and as a subtask of the downstream task, specifically it can be a sink task included in the graph computing task.

[0045] In some embodiments, in the application scenario of distributed graph computing, in order to make full use of the computing resources in the distributed system, both the above-mentioned first subtask and the above-mentioned second subtask can be concurrently executed by multiple computing nodes in the distributed system. That is to say, whether it is an upstream task or a downstream task, it can be a task with a certain degree of concurrency.

[0046] Among them, it should be noted that in actual applications, the degrees of concurrency of the upstream task and the downstream task can be the same or different.

[0047] For example, please refer to Figure 4 , Figure 4 is a schematic diagram showing the concurrent execution of the upstream task and the downstream task shown in this specification.

[0048] As Figure 4 shown, taking the example that the upstream task and the downstream task can have the same degree of concurrency, at this time the upstream task can be concurrently executed by a total of 3 upstream computing nodes, namely nodes 1-3 in the distributed system, and the downstream task can also be concurrently executed by a total of 3 downstream computing nodes, namely nodes A-C in the distributed system.

[0049] In some embodiments, the data generated by the multiple upstream computing nodes for concurrently executing the above-mentioned first subtask, when the first subtask is executed, will be used as input data and continue to be distributed to the multiple downstream computing nodes for concurrently executing the above-mentioned second subtask. The multiple downstream computing nodes will continue to execute the above-mentioned second subtask in parallel based on the data distributed by the upstream computing nodes.

[0050] Among them, in the scenario of distributed graph computing, for any one of the multiple upstream computing nodes above, specifically, the data generated by executing the first subtask can be partitioned according to the concurrency degree of the downstream computing nodes executing the second subtask, so as to further divide the data generated by executing the first subtask into multiple data shards corresponding one by one to the multiple downstream computing nodes executing the second subtask, and then distribute the multiple data shards to each of the multiple downstream computing nodes above.

[0051] For example, please refer to Figure 5 , Figure 5 which is another schematic diagram showing the concurrent execution of upstream tasks and downstream tasks shown in this specification.

[0052] As Figure 5 shown, still taking the upstream task being concurrently executed by a total of 3 upstream computing nodes, namely nodes 1-3 in the distributed system, and the downstream task being concurrently executed by a total of three downstream computing nodes, namely nodes A-C in the distributed system, as an example, the upstream node 1 can divide the data generated by executing the first subtask into slice1-A, slice1-B, and slice1-C according to the concurrency degree 3 of the downstream computing nodes, and distribute slice1-A to the downstream node A, distribute slice1-B to the downstream node B, and distribute slice1-C to the downstream node C. Similarly, the upstream node 2 can also divide the data generated by executing the first subtask into slice2-A, slice2-B, and slice2-C according to the concurrency degree 3 of the downstream computing nodes, and distribute slice2-A to the downstream node A, distribute slice2-B to the downstream node B, and distribute slice2-C to the downstream node C; the upstream node 2 can also divide the data generated by executing the first subtask into slice3-A, slice3-B, and slice3-C according to the concurrency degree 3 of the downstream computing nodes, and distribute slice3-A to the downstream node A, distribute slice3-B to the downstream node B, and distribute slice3-C to the downstream node C. Finally, the input data for the downstream node A when executing the second subtask is slice1-A, slice2-A, and slice3-A, the input data for the downstream node B when executing the second subtask is slice1-B, slice2-B, and slice3-B, and the input data for the downstream node C when executing the second subtask is slice1-C, slice2-C, and slice3-C.

[0053] In some embodiments, each upstream computing node that executes the above-mentioned first subtask, in addition to dividing the output data generated by executing the above-mentioned first subtask into multiple data shards corresponding one-to-one to the multiple downstream computing nodes that execute the above-mentioned second subtask according to the concurrency degree of the downstream computing nodes, can also use a unified data structure to abstract each data shard into the form of a data stream.

[0054] Among them, it should be noted that in one case, the upstream computing node can, after completing the execution of the above-mentioned first subtask, use a unified data structure to further abstract each data shard divided from the output data generated by executing the first subtask into the form of a data stream; in another case, the upstream computing node can also, after completing the execution of the above-mentioned first subtask, only divide the output data generated by executing the first subtask into multiple data shards corresponding one-to-one to the multiple downstream computing nodes that execute the above-mentioned second subtask, and instead wait until it receives a data acquisition request sent by the downstream computing node, and then use a unified data structure to further abstract each data shard divided from the output data generated by executing the first subtask into the form of a data stream.

[0055] In some embodiments, after abstracting each data shard into the form of a data stream based on the above-mentioned unified data structure, each data shard can specifically include at least one batch of data at this time, and any one batch of data can specifically include at least one data block. For example, please refer to Figure 6 , Figure 6 which is a data structure diagram of the data generated by an upstream computing node executing the first subtask shown in this specification.

[0056] As Figure 6 shown, in the data structure of the data generated by the upstream computing node executing the first subtask, it can specifically include Slice, batch, and Message.

[0057] Among them, the above-mentioned Slice is the basic management unit of the data generated by executing the first subtask. Usually, the upstream computing node will divide the data generated by executing the upstream task into N slices according to the concurrency degree N of the downstream computing node executing the above-mentioned second subtask. Each slice can also be further abstracted into the form of a data stream according to the unified data structure shown in Figure 5 .

[0058] The above-mentioned Batch refers to the basic management unit of the data contained within a Slice. A Batch can represent a batch of data, and within a Slice, there can be one or more batches of data. That is to say, a specific Slice can be a data stream composed of one or more batches of data. Within a Slice, different batches can be separated by a barrier, which can be understood as a separator between different batches, used to distinguish different batches and indicate the end position of the current batch. Some metadata of the current batch can be recorded within a barrier; for example, this metadata can include information such as batch ID and the length of the batch.

[0059] It should be noted that the number of batches of data contained within a Slice divided by an upstream computing node for any downstream computing node is usually a number corresponding to the computing mode adopted by the downstream computing node to execute downstream tasks. That is to say, the number of batches of data contained within a Slice divided by an upstream computing node for a downstream computing node usually depends on the computing mode adopted by the downstream computing node to execute downstream tasks.

[0060] For example, in a distributed graph computing scenario, the above-mentioned multiple computing modes can specifically include a stream processing mode, a batch processing mode, and an iterative graph computing mode.

[0061] Among them, in the stream processing mode, a Slice can contain multiple batches of data. In the batch processing mode, a Slice can contain only a single batch of data. For the iterative graph processing mode, in practical applications, it can be executed in the way of stream processing or in the way of batch processing. On the one hand, if the iterative graph computing mode is an iterative graph computing mode executed in the way of stream processing, at this time, a Slice can contain multiple batches of data; if the iterative graph computing mode is an iterative graph computing mode executed in the way of batch processing, at this time, a Slice can contain only a single batch of data.

[0062] In practical applications, after an upstream computing node divides a slice for any downstream computing node, it can further obtain the computing mode adopted by the downstream computing node when executing the second subtask. Then, based on the obtained computing mode, it can determine the number of batches of data that should be included within the slice divided for the downstream computing node, and further divide the data included within the slice into one or more batches of data based on the determined number. Furthermore, the data included in the slice divided for the downstream computing node is further abstracted into the form of a data stream.

[0063] For example, in one case, if the upstream computing node, after completing the execution of the above-mentioned first subtask, uses the above unified data structure to further abstract each data shard divided from the output data generated by executing the first subtask into the form of a data stream. At this time, it can obtain the computing mode adopted by the downstream computing node when executing the second subtask from the task scheduler (this computing mode is usually a computing mode planned in advance when specifying the execution plan of the graph computing task). Then, based on the obtained computing mode, it can determine the number of batches of data that should be included within the slice divided for the downstream computing node, and further divide the data included within the slice into one or more batches of data based on the determined number.

[0064] In another case, if the upstream computing node, after completing the execution of the above-mentioned first subtask, only divides the output data generated by executing the first subtask into multiple data shards corresponding one by one to the multiple downstream computing nodes that execute the above-mentioned second subtask. After receiving the data acquisition request sent by the downstream computing node, it then uses the unified data structure to further abstract each data shard divided from the output data generated by executing the first subtask into the form of a data stream. At this time, it can determine the computing mode adopted by the downstream computing node when executing the second subtask according to the data acquisition conditions included in the data acquisition request sent by the downstream computing node during the data acquisition phase. Then, based on the obtained computing mode, it can determine the number of batches of data that should be included within the slice divided for the downstream computing node, and further divide the data included within the slice into one or more batches of data based on the determined number.

[0065] The above Message refers to the basic management unit of the data contained within a batch. A Message can represent a smaller data block obtained by splitting the data contained in a batch; the data contained in a single batch can be further split into one or more Messages; that is to say, a batch can specifically also be a data stream composed of one or more Message data. Inside a Message, in addition to being able to record some of the data contained in a batch, in practical applications, some metadata of the current batch can also be recorded; for example, information such as batchID, the length of the batch, etc. It should be noted that the number of Messages finally split from the data contained within a batch generally depends on the total length of the data contained within a batch and the unit length set for each Message. In practical applications, based on specific requirements, the unit length of the Message can be flexibly set to split the data contained within a batch.

[0066] In some embodiments, the data distribution methods provided for the above-mentioned multiple computing modes in a distributed system can be unified into a Pull-based Shuffle data distribution method; that is to say, the downstream computing nodes actively send data acquisition requests to the upstream computing nodes, and the upstream computing nodes can process these data acquisition requests and send the data required by the downstream computing nodes to the downstream computing nodes as needed.

[0067] In this case, for any target downstream computing node among the multiple downstream computing nodes used to execute the second subtask, the data acquisition conditions corresponding to the computing mode can be determined first based on the computing mode adopted for executing the above-mentioned second subtask.

[0068] Among them, the data acquisition conditions are specifically used to obtain the input data required by the target downstream computing node to execute the second subtask from the data generated by the upstream computing nodes when executing the first subtask. Different computing modes adopted by the target downstream computing node to execute the second subtask can respectively correspond to different data acquisition conditions; that is to say, under different computing modes, the data acquisition conditions can have certain differences. The upstream computing nodes should be able to accurately determine the computing mode adopted by the downstream computing nodes to execute the second subtask based on the differences in the data acquisition conditions contained in the data acquisition requests sent by the downstream computing nodes.

[0069] After the target downstream computing node determines the data acquisition conditions corresponding to the computing mode adopted for executing the above-mentioned second subtask, it can construct data acquisition requests respectively for multiple upstream computing nodes used to execute the first subtask based on the determined data acquisition conditions; that is to say, construct multiple data acquisition requests based on the concurrency of the upstream computing nodes executing the first subtask; for example, assuming the concurrency of the upstream computing nodes executing the first subtask is N, then N data acquisition requests need to be constructed.

[0070] Then, the constructed data acquisition requests can be sent to the above-mentioned multiple upstream computing nodes respectively.

[0071] In some embodiments, in different computing modes, the upstream computing nodes executing the above-mentioned first subtask and the downstream computing nodes executing the above-mentioned second subtask can run simultaneously or successively.

[0072] On the one hand, if the upstream computing nodes executing the above-mentioned first subtask and the downstream computing nodes executing the above-mentioned second subtask run simultaneously, at this time, the downstream computing nodes executing the above-mentioned second subtask can send data acquisition requests to the upstream computing nodes immediately after the upstream and downstream computing nodes are started simultaneously. On the other hand, if the upstream computing nodes executing the above-mentioned first subtask run before the downstream computing nodes executing the above-mentioned second subtask, at this time, the downstream computing nodes executing the above-mentioned second subtask can also send data acquisition requests to the upstream computing nodes after the above-mentioned upstream computing nodes complete the execution of the above-mentioned second subtask.

[0073] For example, in one example, in the batch processing mode, since the downstream computing nodes are usually scheduled to continue executing the above-mentioned second subtask after the upstream computing nodes complete the execution of the above-mentioned first subtask; therefore, based on the characteristics of task execution in the batch processing mode, the upstream computing nodes usually start before the downstream computing nodes, and the downstream computing nodes start after the upstream computing nodes complete the execution of the above-mentioned first subtask; in this case, the downstream computing nodes can send data acquisition requests to the upstream computing nodes after the above-mentioned upstream computing nodes complete the execution of the above-mentioned first subtask and they themselves complete the startup.

[0074] In another example, in the stream processing mode, the upstream computing nodes and the downstream computing nodes usually start simultaneously, and in this case, the downstream computing nodes can send data acquisition requests to the upstream computing nodes immediately after the upstream and downstream computing nodes are started simultaneously.

[0075] In the third example, in the iterative graph computing mode, since the upstream computing node and the downstream computing node usually correspond to the same computing node, in this computing mode, the upstream computing node and the downstream computing node can be two different threads running on the same computing node, and these two threads can usually be started simultaneously. In this case, after the thread corresponding to the downstream computing node and the thread corresponding to the upstream computing node and the thread corresponding to the downstream computing node are started simultaneously, the thread corresponding to the downstream computing node can immediately send a data acquisition request to the thread corresponding to the above-mentioned upstream computing node.

[0076] Step 204, in response to the data acquisition request, obtain target data that meets the data acquisition conditions from the data generated by executing the first subtask; For any target upstream computing node among the multiple upstream computing nodes that execute the above first subtask, it can receive a data acquisition request sent by a downstream computing node. After receiving a data acquisition request sent by any target downstream computing node among the above multiple downstream computing nodes, at this time, it can further respond to the data acquisition request and adopt a unified Pull-based Shuffle data distribution method for different computing modes to distribute data to the target downstream computing node.

[0077] It should be noted that since a unified Pull-based Shuffle data distribution method can be adopted for different computing modes to distribute data to the downstream computing node, at this time, regardless of the computing mode adopted by the downstream computing node, the upstream computing node can adopt a unified data distribution method to distribute data to the downstream; therefore, in this case, the multiple downstream computing nodes used to execute the second subtask can adopt different computing modes to execute the second subtask in parallel.

[0078] In some embodiments, the unified Pull-based Shuffle data distribution method provided for the above multiple computing modes can specifically be a non-blocking pull mode.

[0079] In this non-blocking pull mode, the data acquisition request sent by the downstream computing node to the upstream computing node can specifically be an asynchronous request. Correspondingly, after the above target upstream computing node receives the data acquisition request sent by the above target downstream computing node, it can asynchronously respond to the data acquisition request and obtain target data that meets the data acquisition conditions from the data generated by executing the first subtask.

[0080] It should be noted that the specific form of the data acquisition conditions included in the above data acquisition request is not specifically limited in this specification. In practical applications, the shard identifier of the data shard allocated to the downstream computing node and the batch identifier of the data batch to be acquired or the number of batches to be acquired can be explicitly carried in the above data acquisition request as the above data acquisition conditions, or some implicit data acquisition rules can be carried in the above data acquisition request as the above data acquisition conditions, which will not be listed one by one in this specification.

[0081] In some embodiments, the data acquisition conditions included in the above data acquisition request may specifically include the shard identifier of the data shard allocated by the above target upstream computing node to the above target downstream computing node, the batch identifier of the starting data batch to be acquired, and the number of data batches to be acquired.

[0082] In this case, after receiving the data acquisition request sent by the above target downstream computing node, the above target upstream computing node can first obtain the above shard identifier, the above batch identifier, and the above number of batches included in the data acquisition request as the data acquisition conditions.

[0083] Secondly, it can also determine the target data shard corresponding to the above shard identifier from the data shards respectively allocated to the above multiple downstream computing nodes; For example, in practical applications, when the upstream computing node allocates data shards to the downstream computing node, by default, the node identifier of the downstream computing node can be used as the shard identifier of the data shard, or by default, the node identifier of the downstream computing node can be used as part of the shard identifier of the data shard; in this way, when constructing the data acquisition request, the target downstream computing node can default to carry its own node identifier as the shard identifier of the data shard allocated to it by the target upstream computing node in the data acquisition request. After receiving the data acquisition request, the target upstream computing node can directly match the shard identifier included in the data acquisition request with the shard identifiers of the data shards already allocated to each downstream computing node, and then determine the target data shard corresponding to the shard identifier included in the data acquisition request.

[0084] Then, it can further determine the target data batch corresponding to the above batch identifier from at least one batch of data included in the target data shard.

[0085] For example, in practical applications, the batch identifier of each batch of data included in the data shard allocated by the upstream computing node for the downstream computing node can specifically be an identifier that monotonically increases starting from a default value as the starting value (such as 0 or 1). In this case, the batch identifier of the starting data batch included in the data acquisition request sent by the target downstream computing node to the above-mentioned target upstream computing node can specifically be the default value as the starting value. At this time, this default value represents the batch identifier corresponding to the data of the first batch included in the target data shard. After receiving the data acquisition request, the target upstream computing node can defaultly determine the data of the first batch included in the above-mentioned target data shard as the target data batch corresponding to the above-mentioned batch identifier.

[0086] Finally, the target data batch can be used as the starting data batch, and at least one batch of data corresponding to the value of the above-mentioned batch quantity can be sequentially obtained as the target data. For example, if the value of the above-mentioned batch quantity is N, then N batches of data need to be sequentially obtained starting from the target data batch as the target data. Then, the obtained above-mentioned target data is sent to the above-mentioned target downstream computing node.

[0087] For example, still taking the above-mentioned multiple computing modes including the stream processing mode, batch processing mode, and iterative graph computing mode as an example, if the computing mode adopted by the above-mentioned target downstream computing node to execute the above-mentioned second subtask is the stream processing mode, at this time, the target data shard allocated by the above-mentioned target upstream computing node for the target downstream computing node can specifically include multiple batches of data. In this case, the batch identifier of the starting data batch included in the data acquisition request sent by the target downstream computing node to the above-mentioned target upstream computing node can specifically be the batch identifier corresponding to the data of the first batch among the multiple batches of data included in the target data shard. For example, the batch identifiers of the multiple batches of data included in the target data shard can specifically be an identifier that monotonically increases starting from a default value as the starting value. At this time, the identifier of the data of the first batch can be a default value as the starting value. The value of the above-mentioned batch quantity can be infinite (such as a character representing infinity), indicating that all batches of data included in the target data shard are to be obtained.

[0088] If the computing mode adopted by the above-mentioned target downstream computing node to execute the above-mentioned second subtask is the batch processing mode, at this time, the target data slice allocated by the above-mentioned target upstream computing node to the target downstream computing node may specifically include only the data of a single batch; in this case, the batch identifier of the above-mentioned starting data batch included in the data acquisition request sent by the target downstream computing node to the above-mentioned target upstream computing node may specifically be the batch identifier corresponding to the data of the single batch; for example, the batch identifier corresponding to the data of the single batch may be a default value as the starting value (such as 0 or 1); the value of the above-mentioned batch quantity may be 1, indicating that in the batch processing mode, only the data of the single batch included in the above-mentioned target data slice needs to be acquired.

[0089] If the computing mode adopted by the above-mentioned target downstream computing node to execute the above-mentioned second subtask is the iterative graph processing mode, that is to say, at this time, the second subtask itself is an iterative graph computing subtask. In this case, since the iterative graph processing mode usually includes an iterative graph computing mode executed in a stream processing manner and an iterative graph computing mode executed in a batch processing manner, it is necessary to discuss these two cases separately at this time: On the one hand, if the computing mode adopted by the above-mentioned target downstream computing node to execute the above-mentioned second subtask is the iterative graph computing mode executed in a stream processing manner, at this time, the above-mentioned first subtask may usually be the computing task of any round of iteration included in the iterative graph computing, and the above-mentioned second subtask may usually be the computing task of the next round of iteration of the any round of iteration. The target data slice allocated by the above-mentioned target upstream computing node to the target downstream computing node may specifically include the data of multiple batches. That is to say, in the iterative graph computing mode executed in a stream processing manner, the data included in the target data slice will be further divided into the data of multiple batches and sent to the above-mentioned target downstream computing node in multiple batches.

[0090] In this case, the batch identifier of the above-mentioned starting data batch included in the data acquisition request sent by the target downstream computing node to the above-mentioned target upstream computing node may specifically be the batch identifier corresponding to the data of the first batch among the data of multiple batches included in the target data slice; for example, the batch identifiers of the data of multiple batches included in the target data slice may specifically be an identifier that monotonically increases starting from a default value as the starting value. At this time, the identifier of the data of the first batch may be a default value as the starting value. The value of the above-mentioned batch quantity may be infinite.

[0091] On the other hand, if the computing mode adopted by the above-mentioned target downstream computing node to execute the above-mentioned second subtask is an iterative graph computing mode executed in a batch processing manner, at this time, the above-mentioned first subtask can still be the computing task of any round of iteration included in the iterative graph computing. The above-mentioned second subtask can usually be the computing task of the next round of iteration of the above-mentioned any round of iteration. The target data slice allocated by the above-mentioned target upstream computing node to the above-mentioned target downstream computing node can specifically contain only the data of a single batch. That is to say, in the iterative graph computing mode executed in a batch processing manner, the data contained in the target data slice will be further divided into the data of a single batch, and then batch-transmitted to the above-mentioned target downstream computing node.

[0092] In this case, the batch identifier of the above-mentioned starting data batch included in the data acquisition request sent by the above-mentioned target downstream computing node to the above-mentioned target upstream computing node can specifically be the batch identifier corresponding to the data of the single batch; for example, the batch identifier corresponding to the data of the single batch can be a default value (such as 0 or 1) as the starting value; the value of the above-mentioned batch quantity can be 1, indicating that in the iterative graph computing mode executed in a batch processing manner, only the data of the single batch contained in the above-mentioned target data slice needs to be acquired.

[0093] In some embodiments, the data acquisition conditions included in the above-mentioned data acquisition request can specifically further include the slice identifier of the data slice allocated by the above-mentioned target upstream computing node to the above-mentioned target downstream computing node, and the maximum batch identifier of all data batches to be acquired.

[0094] In this case, after receiving the data acquisition request sent by the above-mentioned target downstream computing node, the above-mentioned target upstream computing node can first acquire the above-mentioned slice identifier and the above-mentioned maximum batch identifier included in the data acquisition request as the data acquisition conditions.

[0095] Secondly, it can also determine the target data slice corresponding to the above-mentioned slice identifier from the data slices respectively allocated to the above-mentioned multiple downstream computing nodes, and then acquire at least one batch of data with a batch identifier not greater than the above-mentioned maximum batch identifier from at least one batch of data contained in the above-mentioned target data slice as the above-mentioned target data.

[0096] For example, still taking the above-mentioned multiple computing modes including the stream processing mode, batch processing mode, and iterative graph computing mode as an example, if the computing mode adopted by the above-mentioned target downstream computing node to execute the above-mentioned second subtask is the stream processing mode, at this time, the above-mentioned target data slice may contain data of multiple batches; for example, the batch identifiers of the data of multiple batches contained in the target data slice may specifically be an identifier that monotonically increases starting from a default value (such as 0 or 1) as the starting value; the value of the above-mentioned maximum batch identifier may be infinite, indicating that all batches of data contained in the target data slice are to be obtained.

[0097] If the computing mode adopted by the above-mentioned target downstream computing node to execute the second subtask is the batch processing mode, at this time, the above-mentioned target data slice may contain only data of a single batch; for example, the batch identifier corresponding to the data of the single batch may be a default value (such as 0 or 1) as the starting value; the value of the above-mentioned maximum batch identifier may be the batch identifier corresponding to the data of the single batch.

[0098] If the computing mode adopted by the above-mentioned target downstream computing node to execute the above-mentioned second subtask is the iterative graph processing mode, at this time, the second subtask itself is an iterative graph computing subtask. In this case, since the iterative graph processing mode usually includes an iterative graph computing mode executed in the way of stream processing and an iterative graph computing mode executed in the way of batch processing, it is necessary to discuss these two situations separately: On the one hand, if the computing mode adopted by the above-mentioned target downstream computing node to execute the second subtask is the iterative graph computing mode executed in the way of stream processing, at this time, the above-mentioned first subtask is the computing task of any round of iteration in the iterative graph computing, and the above-mentioned second subtask is the computing task of the next round of iteration of the any round of iteration; the above-mentioned target data slice may contain data of multiple batches, and the value of the above-mentioned maximum batch identifier may be infinite, indicating that all batches of data contained in the target data slice are to be obtained; On the other hand, if the computing mode adopted by the above-mentioned target downstream computing node to execute the above-mentioned second subtask is the iterative graph computing mode executed in the way of batch processing, at this time, the above-mentioned first subtask is still the computing task of any round of iteration in the iterative graph computing, and the above-mentioned second subtask is the computing task of the next round of iteration of the any round of iteration; the above-mentioned target data slice may contain only data of a single batch, and the value of the above-mentioned maximum batch identifier may be the batch identifier corresponding to the data of the single batch.

[0099] Step 206, send the target data to the target downstream computing node, so that the target downstream computing node continues to execute the second subtask with the target data as the input data.

[0100] After the above-mentioned target upstream node obtains the target data that meets the above-mentioned data acquisition condition from the data generated by executing the first subtask, it can send the target data to the above-mentioned target downstream computing node. Among them, the target data usually can include at least one batch of data obtained from the data shards allocated for the target downstream computing node.

[0101] After the above-mentioned target downstream computing node receives the above-mentioned target data sent by the above-mentioned target upstream node, it can use the target data as input data to continue executing the above-mentioned second subtask.

[0102] It should be noted that since multiple upstream computing nodes that execute the first subtask in parallel will each allocate a data shard for the target downstream computing node from the data generated by executing the first subtask, the above-mentioned target downstream computing node can specifically obtain the target data from the data shards allocated for itself by each upstream computing node among the multiple upstream computing nodes by sending a data acquisition request to the multiple upstream computing nodes, and then integrate the multiple target data obtained from the data shards allocated for itself by each upstream computing node as input data to continue executing the above-mentioned second subtask.

[0103] For example, please continue to refer to Figure 5 , for downstream computing node A, it can respectively obtain slice1-A generated by upstream computing node 1 for downstream computing node A, slice2-A generated by upstream computing node 2 for downstream computing node A, and slice3-A generated by upstream computing node 3 for downstream computing node A, and then integrate slice1-A, slice2-A and slice3-A as input data to continue executing the above-mentioned second subtask.

[0104] In some embodiments, since a unified Pull-based Shuffle data distribution method can be adopted for different computing modes to distribute data to downstream computing nodes, and a unified data structure can also be adopted to further abstract the data shards allocated for each downstream subtask from the output data generated by executing the upstream task into the form of a data stream; therefore, adopting this unified data distribution mode is equivalent to abstracting the batch data that needs to be processed in the batch processing mode into a data stream with a batch size of 1, and abstracting the data generated by N iterations that need to be processed in the iterative graph computing mode into a limited number of batches of data streams.

[0105] Based on this feature, when the above-mentioned target upstream computing node distributes the obtained above-mentioned target data to the target downstream computing node, specifically, it can distribute the data blocks in at least one batch of data included in the obtained target data to the above-mentioned target downstream computing node in a streaming transmission manner.

[0106] For example, please continue to refer to Figure 6 , each Message in at least one slice included in the obtained target data can be used as the smallest transmission unit, and can be distributed to the above-mentioned target downstream computing node one by one in a streaming transmission manner.

[0107] In some embodiments, since in the iterative graph computing mode, the above-mentioned target upstream computing node and the above-mentioned target downstream computing node usually correspond to the same computing node. At this time, the above-mentioned target upstream computing node and the above-mentioned target downstream computing node can usually be two different threads running on the same computing node. Therefore, if the computing mode adopted by the above-mentioned target downstream computing node to execute the second subtask is the iterative graph computing mode, at this time, the above-mentioned first subtask is usually the computing task of any round of iteration in the iterative graph computing, and the above-mentioned second subtask is the computing task of the next round of iteration of this arbitrary round of iteration.

[0108] In this case, when the thread corresponding to the above-mentioned target upstream computing node distributes the obtained above-mentioned target data to the target downstream computing node, specifically, it can distribute the above-mentioned target data to a pre-specified local storage space on the computing node where it is located; for example, the data blocks in at least one batch of data included in the obtained target data can be written into the above-mentioned local storage space one by one in a streaming transmission manner.

[0109] The thread corresponding to the above-mentioned target downstream computing node can specifically read the above-mentioned target data from this local storage space as the input data for executing the above-mentioned second subtask.

[0110] Among them, this local storage space can be a pre-specified memory space or disk storage space on this computing node.

[0111] In some embodiments, if the computing mode adopted by the above-mentioned target downstream computing node to execute the downstream task is the batch processing mode, and if the concurrency degrees of the upstream task and the downstream task are the same at this time, the second subtask executed by the above-mentioned target downstream computing node and the first subtask executed by the above-mentioned target upstream computing node can be deployed to the same computing node for running; that is to say, in the batch processing mode where the concurrency degrees of the upstream task and the downstream task are the same, the above-mentioned target upstream computing node and the above-mentioned target downstream computing node can correspond to the same computing node. At this time, the above-mentioned target upstream computing node and the above-mentioned target downstream computing node can be two different threads running on the same computing node.

[0112] In this case, since the upstream computing node usually starts before the downstream computing node in the batch processing mode, after the thread corresponding to the above-mentioned target upstream computing node starts, an additional independent thread can be started in advance on the computing node where it is located for the thread corresponding to the above-mentioned target downstream computing node; wherein, this independent thread will run independently of the thread corresponding to the above-mentioned target downstream computing node. After this independent thread starts, data acquisition requests can be sent to multiple upstream computing nodes used to execute the above-mentioned first subtask respectively in advance.

[0113] In this way, since the thread corresponding to the above-mentioned target upstream computing node can receive the data acquisition request sent by the thread corresponding to the above-mentioned target downstream computing node in advance, the thread corresponding to the above-mentioned target upstream computing node can obtain the target data from the data shard allocated to the thread corresponding to the above-mentioned target upstream computing node in advance, and can send the obtained target data to a pre-specified local storage space on the computing node where it is located in advance. Subsequently, after the thread corresponding to the above-mentioned target downstream computing node starts, it can directly read the above-mentioned target data from this local storage space as the input data for executing the above-mentioned second subtask, without the need to repeatedly send data acquisition requests.

[0114] Please refer to Figure 7 , Figure 7 which is a flowchart of another data distribution method for distributed graph computing shown in this specification; the computing nodes in the distributed system support executing distributed graph computing tasks according to multiple computing modes; the distributed graph computing task includes a first subtask as an upstream task and a second subtask as a downstream task; wherein, both the first subtask and the second subtask are concurrently executed by multiple computing nodes in the distributed system; the method is applied to any target downstream computing node among multiple downstream computing nodes that execute the second subtask; the method includes the following execution process: Step 702: Determine the data acquisition condition based on the computing mode adopted for executing the second subtask. The data acquisition condition is used to obtain the input data required for executing the second subtask from any target upstream computing node among multiple upstream computing nodes that execute the first subtask. Step 704: Generate a data acquisition request containing the data acquisition condition, and send the data acquisition request to the target upstream computing node, so that the target upstream computing node obtains target data that meets the data acquisition condition from the data generated by executing the first subtask. Step 704: Receive the target data sent by the target upstream computing node, and continue to execute the second subtask using the target data as the input data.

[0115] It should be noted that steps 202 - 204 included in the execution process shown above are a method flow with any target upstream computing node among multiple downstream computing nodes that execute the above-mentioned first subtask (i.e., the upstream task) as the execution entity, while Figure 2 different from this, steps 702 - 704 included in the execution process shown in Figure 2 are a method flow with any target downstream computing node among multiple downstream computing nodes that execute the above-mentioned second subtask as the execution entity. Since Figure 7 the implementation details of the method flow shown in Figure 7 are exactly the same as the implementation details of the method flow shown in Figure 2 , the implementation details related to each step included in Figure 7 will not be elaborated in this embodiment. Those skilled in the art can refer to the records of the previous embodiments.

[0116] In the above technical solution, a unified data distribution method can be provided for various computing modes supported by a distributed system that executes distributed graph computing. For different computing modes, there is no need to separately provide different data distribution methods, thereby reducing the complexity of system development and maintenance, and enabling the system to more flexibly adapt to different computing modes.

[0117] Corresponding to the embodiment of the foregoing method, this specification also provides embodiments of a device, an electronic device, and a storage medium.

[0118] Figure 8 is a schematic structural diagram of an electronic device provided by an exemplary embodiment. Please refer to Figure 8, at the hardware level, the device includes a processor 802, an internal bus 804, a network interface 806, a memory 808, and a non-volatile memory 810. Of course, it may also include other required hardware. One or more embodiments of this specification can be implemented in software. For example, the processor 802 reads the corresponding computer program from the non-volatile memory 810 into the memory 808 and then runs it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic unit and can also be hardware or a logic device.

[0119] As Figure 9 shown, Figure 9 is a block diagram of a data distribution device for distributed graph computing shown according to an exemplary embodiment of this specification. The device can run in an electronic device as Figure 8 shown to implement the technical solution of this specification. Among them, the computing nodes in the distributed system for executing the distributed graph computing task support executing the distributed graph computing task according to multiple computing modes; the distributed graph computing task includes a first subtask as an upstream task and a second subtask as a downstream task; both the first subtask and the second subtask are concurrently executed by multiple computing nodes in the distributed system; the device is applied to any target upstream computing node among the multiple upstream computing nodes executing the first subtask; the device 90 includes: A receiving module 901 that receives a data acquisition request sent by any target downstream computing node among the multiple downstream computing nodes for executing the second subtask; wherein, the data acquisition request includes data acquisition conditions determined by the target downstream computing node based on the computing mode adopted for executing the second subtask; the data acquisition conditions are used to acquire the input data required by the target downstream computing node for executing the second subtask from the data generated by the target upstream computing node for executing the first subtask; different computing modes adopted by the target downstream computing node for executing the second subtask respectively correspond to different data acquisition conditions; An acquisition module 902 that, in response to the data acquisition request, acquires target data that meets the data acquisition conditions from the data generated by executing the first subtask; A sending module 903 that sends the target data to the target downstream computing node so that the target downstream computing node continues to execute the second subtask with the target data as input data.

[0120] As Figure 10 shown, Figure 10FIG. 0 is a block diagram of another data distribution device for distributed graph computing shown in accordance with an exemplary embodiment of the present specification. This device may also run in an electronic device as shown in Figure 8 to implement the technical solutions of the present specification. Among them, the computing nodes in the distributed system for executing distributed graph computing tasks support executing the distributed graph computing tasks according to multiple computing modes; the distributed graph computing tasks include a first subtask as an upstream task and a second subtask as a downstream task; wherein, both the first subtask and the second subtask are concurrently executed by multiple computing nodes in the distributed system; the method is applied to any target downstream computing node among the multiple downstream computing nodes executing the second subtask; the device 100 includes: A determination module 1001 determines a data acquisition condition based on the computing mode used to execute the second subtask; wherein, the data acquisition condition is used to acquire the input data required to execute the second subtask from any target upstream computing node among the multiple upstream computing nodes executing the first subtask; A sending module 1002 generates a data acquisition request including the data acquisition condition and sends the data acquisition request to the target upstream computing node, so that the target upstream computing node acquires target data that meets the data acquisition condition from the data generated by executing the first subtask; An execution module 1003 receives the target data sent by the target upstream computing node and continues to execute the second subtask with the target data as the input data.

[0121] Correspondingly, the present specification also provides an electronic device, which includes a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the steps in all the method processes described above.

[0122] Correspondingly, the present specification also provides a computer-readable storage medium, on which executable computer program instructions are stored; wherein, when the instructions are executed by a processor, the steps in all the method processes described above are implemented.

[0123] Correspondingly, the present specification also provides a computer program product, on which executable computer program instructions are stored; wherein, when the computer program instructions are executed by a processor, the steps in all the method processes described above are implemented.

[0124] In the 1990s, it was obvious to distinguish whether an improvement in a technology was an improvement in hardware (e.g., an improvement in the circuit structure of diodes, transistors, switches, etc.) or an improvement in software (an improvement in the method flow). However, with the development of technology, many improvements in method flows today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. The designer can program by himself to "integrate" a digital system on a piece of PLD without asking the chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL). And there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that only by slightly logically programming the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit can the hardware circuit implementing the logical method flow be easily obtained.

[0125] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to implement the same function by logically programming the method steps so that the controller is in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0126] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technology, the computers implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.

[0127] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or terminal product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, the presence of additional identical or equivalent elements in the process, method, product or device comprising the said elements is not excluded. For example, if terms such as first and second are used to denote names, they do not denote any particular order.

[0128] For convenience of description, the above device is described by dividing it into various modules according to functions. Of course, when implementing one or more of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0129] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0130] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more of the blocks and / or processes. Figure 1 one or more of the processes and / or blocks Figure 1 specified in the one or more of the blocks and / or processes.

[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the one or more of the blocks and / or processes.

[0132] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0133] Memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0134] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0135] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0136] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0137] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples.

[0138] The above description is only for the embodiments of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.

Claims

1. A data distribution method for distributed graph computing, where computing nodes in a distributed system support executing distributed graph computing tasks according to multiple computing modes; the distributed graph computing tasks include a first subtask as an upstream task and a second subtask as a downstream task; among them, Both the first subtask and the second subtask are concurrently executed by multiple computing nodes in the distributed system; The method is applied to any target upstream computing node among multiple upstream computing nodes that execute the first subtask, and the method includes: Receiving a data acquisition request sent by any target downstream computing node among multiple downstream computing nodes that execute the second subtask; wherein, the data acquisition request includes data acquisition conditions determined by the target downstream computing node based on the computing mode adopted for executing the second subtask; the data acquisition conditions are used to acquire input data required for the target downstream computing node to execute the second subtask from the data generated by the target upstream computing node when executing the first subtask; different computing modes adopted by the target downstream computing node for executing the second subtask respectively correspond to different data acquisition conditions; In response to the data acquisition request, acquiring target data that meets the data acquisition conditions from the data generated by executing the first subtask; Sending the target data to the target downstream computing node, so that the target downstream computing node continues to execute the second subtask with the target data as input data.

2. The method according to claim 1, wherein the data generated by executing the first subtask includes data shards respectively allocated to multiple downstream computing nodes for executing the second subtask; the data shards include at least one batch of data; any one batch of data includes at least one data block; Among them, The number of batches of data included in the data shard allocated to any downstream computing node is the number corresponding to the computing mode adopted by the downstream computing node for executing the second subtask.

3. The method according to claim 2, wherein the data acquisition conditions include the shard identifier of the data shard allocated by the target upstream computing node to the target downstream computing node, the batch identifier of the starting data batch to be acquired, and the number of batches of data to be acquired; In response to the data acquisition request, acquiring target data that meets the data acquisition conditions from the data generated by executing the first subtask, includes: In response to the data acquisition request, determining a target data shard corresponding to the shard identifier from the data shards respectively allocated to the multiple downstream computing nodes; Determining a target data batch corresponding to the batch identifier from at least one batch of data included in the target data shard, and using this target data batch as the starting data batch, and sequentially acquiring at least one batch of data corresponding to the value of the number of batches as the target data.

4. The method according to claim 3, wherein the multiple computing modes include a stream processing mode, a batch processing mode, and an iterative graph computing mode; If the computing mode adopted by the target downstream computing node for executing the second subtask is the stream processing mode, the target data shard includes multiple batches of data, the batch identifier of the starting data batch is the batch identifier corresponding to the first batch of data among the multiple batches of data, and the value of the number of batches is infinite; If the computing mode adopted by the target downstream computing node to execute the second subtask is the batch processing mode, the target data slice contains data of only one batch, the batch identifier of the starting data batch is the batch identifier corresponding to the data of the only one batch, and the value of the number of batches is 1; Among them, The batch identifiers of the data of the multiple batches increase monotonically.

5. The method according to claim 4, wherein the iterative graph computing mode includes an iterative graph computing mode executed in a stream processing manner; and an iterative graph computing mode executed in a batch processing manner; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in a stream processing manner, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data slice contains data of multiple batches; the batch identifier of the starting data batch is the batch identifier corresponding to the data of the first batch among the data of the multiple batches; the value of the number of batches is infinite; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in a batch processing manner, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data slice contains data of only one batch, the batch identifier of the starting data batch is the batch identifier corresponding to the data of the only one batch, and the value of the number of batches is 1; Wherein, the batch identifiers of the data of the multiple batches increase monotonically.

6. The method according to claim 2, wherein the data acquisition condition includes the slice identifier of the data slice allocated by the target upstream computing node to the target downstream computing node, and the maximum batch identifier of all data batches to be acquired; Responding to the data acquisition request, acquiring target data that meets the data acquisition condition from the data generated by executing the first subtask includes: Responding to the data acquisition request, determining a target data slice corresponding to the slice identifier from the data slices respectively allocated to the multiple downstream computing nodes; Acquiring at least one batch of data whose batch identifier is not greater than the maximum batch identifier from at least one batch of data contained in the target data slice as the target data.

7. The method according to claim 6, wherein the multiple computing modes include a stream processing mode, a batch processing mode, and an iterative graph computing mode; If the computing mode adopted by the target downstream computing node to execute the second subtask is the stream processing mode, the target data slice contains data of multiple batches, and the value of the maximum batch identifier is infinite; If the computing mode adopted by the target downstream computing node to execute the second subtask is the batch processing mode, the target data slice contains only one batch of data, and the value of the maximum batch identifier is the batch identifier corresponding to the only batch of data; Among them, The batch identifiers of the multiple batches of data increase monotonically.

8. The method according to claim 7, wherein the iterative graph computing mode includes an iterative graph computing mode executed in a stream processing manner; and an iterative graph computing mode executed in a batch processing manner; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in a stream processing manner, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data slice contains multiple batches of data, and the value of the maximum batch identifier is infinity; If the computing mode adopted by the target downstream computing node to execute the second subtask is an iterative graph computing mode executed in a batch processing manner, the first subtask is the computing task of any round of iteration in the iterative graph computing, and the second subtask is the computing task of the next round of iteration of the any round of iteration; the target data slice contains only one batch of data, and the value of the maximum batch identifier is the batch identifier corresponding to the only batch of data; Wherein, the batch identifiers of the multiple batches of data increase monotonically.

9. The method according to claim 2, wherein the data acquisition request is an asynchronous request; In response to the data acquisition request, acquiring target data that meets the data acquisition condition from the data generated by executing the first subtask includes: Asynchronously responding to the data acquisition request, and acquiring target data that meets the data acquisition condition from the data generated by executing the first subtask.

10. The method according to claim 9, wherein the target data includes at least one batch of data acquired from the data slice allocated to the target downstream computing node; Sending the target data to the target downstream computing node includes: Sending the data blocks included in the acquired at least one batch of data to the target downstream computing node in a streaming transmission manner.

11. The method according to claim 1, wherein multiple downstream computing nodes for executing the second subtask execute the second subtask in different computing modes.

12. A data distribution method for distributed graph computing, where computing nodes in a distributed system support executing distributed graph computing tasks according to multiple computing modes; the distributed graph computing tasks include a first subtask as an upstream task and a second subtask as a downstream task; among them, Both the first subtask and the second subtask are concurrently executed by multiple computing nodes in the distributed system; The method is applied to any target downstream computing node among multiple downstream computing nodes for executing the second subtask, and the method includes: Determining a data acquisition condition based on the computing mode adopted to execute the second subtask; wherein, the data acquisition condition is used to acquire input data required to execute the second subtask from any target upstream computing node among multiple upstream computing nodes for executing the first subtask. Generate a data acquisition request including the data acquisition condition, and send the data acquisition request to the target upstream computing node, so that the target upstream computing node acquires target data that meets the data acquisition condition from the data generated by executing the first subtask; Receive the target data sent by the target upstream computing node, and continue to execute the second subtask using the target data as input data.

13. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.

14. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.

15. A computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Machine learning calculation optimization method and platform

    CN114418127A

  • Online data skew adjustment method and device for stream computing operation

    CN115437777A

  • Data processing method and device, electronic equipment and storage medium

    CN115857918A

  • Distributed computing scheduling system, task processing method, equipment and storage medium

    CN116360993A

  • Graph data processing method and graph calculation engine

    CN117972154A