BROADCASTING AND SCATTERING OPERATIONS
By encoding tree structures in message headers, the method allows dynamic adaptation in broadcast and scatter algorithms, improving performance in HPC and ML by eliminating synchronization requirements and optimizing data transmission in dynamic networks.
Patent Information
- Application Number
- DE112023004355
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-17
- Filing Date
- 2023-07-21
- Publication Date
- 2025-08-07
AI Technical Summary
Existing broadcast and scatter algorithms in high-performance computing (HPC) and machine learning (ML) require prior knowledge of the tree structure by all members, leading to synchronization issues and performance interruptions when adapting to dynamic network environments.
A method where only the root node knows the tree structure, dynamically encoding it in the message headers, allowing children to adapt without synchronization, optimizing memory references, and managing data transmission efficiently.
Enhances performance by eliminating synchronization needs and optimizing data organization in dynamic network environments, supporting broadcast, scatter, and combined operations with irregular message sizes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUNDThe present invention relates generally to data transmission networks, and more particularly to computer systems, computer implemented methods, and computer program products for performing distributed operations such as broadcast and broadcast operations.Broadcasting a message to a group of participants (e.g., processes, machines) over a network is a frequently used pattern in distributed systems. In this pattern, an original subscriber (sometimes referred to as a "root") sends the same message to a group of remotely located subscribers, often in a tree pattern. Spreading a message can be considered a variant of the broadcast operation. The original subscriber ("root") sends a different message to each remote subscriber, often using a tree pattern to improve performance over a centralized point-to-point data transfer pattern.Many applications for high-performance computing (HPC) and machine learning (ML) rely on the MPI (Message Passing Interface) standard and similar libraries for point-to-point data transfer and shared data transfer between distributed operations. The MPI standard defines the MPI_Bcast and MPI_Atter(v) operations for these two widely distributed common data transmissions. In addition, other common algorithms such as MPI_All and MPI_All are often based on broadcast and scatter operations to support the higher level algorithm.The shared broadcast and spread algorithms are also advantageous for any distributed system of permanent daemons for information distribution. HPC schedulers and jobstarts often use these patterns to update the distributed state and send jobstart messages to all remote systems, thereby launching the application. In particular, the latter example job launch benefits greatly from efficient broadcast and scatter algorithms that allow faster start times for user applications and higher machine load. In general, improvements to these two common communication patterns, particularly in dynamic networked environments such as the cloud, can provide significant performance advantages for client applications and middleware in data centers.SUMMARYEmbodiments of the present invention relate to a method for performing distributed data transfer operations. In one aspect, a method implemented by a computer includes receiving a request from a first data processing system to perform a distributed data transfer operation, and retrieving a tree structure from the first data processing system to perform the distributed data transfer operation, wherein the first data processing system is a root node of the tree structure. The method also includes generating, by the first data processing system, a message including header information and payload data for the distributed data transfer operation, and transmitting, by the first data processing system, a portion of the message to each child node of the first data processing system, wherein the portion transmitted to each child node is unique.Other embodiments of the present invention implement features of the method described above in computer systems and computer program products.Additional technical features and advantages are realized by the techniques of the present invention. Embodiments and aspects of the invention are described in detail herein and are considered part of the claimed subject matter. For a better understanding, reference is made to the description and the drawings.BRIEF DESCRIPTION OF THE DRAWINGSThe specifics of the exclusive rights described herein are particularly pointed out and expressly claimed in the claims at the end of the specification. The foregoing and other features and advantages of the embodiments of the invention will become apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which: FIG. 1 shows a block diagram of an example computer system for use in connection with one or more embodiments of the present invention; FIG. 2 is a block diagram of a tree structure for use in connection with one or more embodiments of the present invention; FIG. 3 is a block diagram illustrating a broadcast operation according to one or more embodiments of the present invention; FIG. 4 is a block diagram illustrating a spreading operation according to one or more embodiments of the present invention; FIG. 5 is a block diagram illustrating a combination of a broadcast operation and a spread operation according to one or more embodiments of the present invention; FIG. 6 is a flow diagram of a method for initiating a distributed data transfer operation, in accordance with one or more embodiments of the present invention; and FIG. 7 is a flow diagram of a method for performing distributed data transfer operations, in accordance with one or more embodiments of the present invention.DETAILED DESCRIPTIONAs discussed above, broadcast and scatter algorithms are increasingly used by HPC and ML processes. Existing broadcast and scatter algorithms assume that members prior to the start of the operation know the tree structure to understand their role in the data transfer protocol. If the tree structure needs to be adapted to a membership, conditions of the network and / or the size of a message, updates of the tree structure need to be distributed before the start of the joint operation. When switching between trees, this often requires hard synchronization, thereby interrupting network traffic. In spread transfer patterns, data organization in the buffer can impact operational performance, for example, when each stage in the algorithm must calculate different memory offsets and access different memory areas to compile a subset of the buffer for its children.In exemplary embodiments, improved broadcast and scatter algorithms are provided in which the members, except for the root, need not previously know the tree structure. The improved broadcast and scatter algorithms dynamically adapt to current network environments without requiring synchronization. Moreover, the enhanced broadcast and scatter algorithms are configured to manage data transmitted during execution of the broadcast and scatter algorithms to improve performance.In example embodiments, only the original party ("root") knows the tree structure when the broadcast and / or scatter operation begins. This allows the original user ("root") to dynamically adapt the tree structure to each message without the need to synchronize with the other users of the data transmission. The root encodes the tree structure for this data transfer operation in the header(s) of the message sent to its children. The non-original participants, the children, receive from their parent node in the tree a message containing instructions for the next hop in the tree with respect to themselves. The message includes a header describing the subordinate tree structure at and below this point in the tree, and instructions to remove irrelevant data for the next step in the data transfer operation. The non-original participants need only unpack their instructions for sending to their children in the tree, if any, regardless of the tree structure above them or below the children. Moreover, the data in each message is organized so that memory references are optimized in a scatter operation as it moves down the tree structure. In an exemplary embodiment, the method may be used for broadcast operations, spread operations with regular and irregular size messages per subscriber, and a combination of both in a combined operation.Various aspects of the present disclosure are described by illustrative text, flow diagrams, block diagrams of computer systems and / or block diagrams of machine logic in embodiments of computer program products (CPP). In all flowcharts, depending on the technology used, the operations may be performed in a different order than illustrated in a given flowchart. For example, again depending on the technology used, two operations depicted in successive blocks of the flow chart may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.An embodiment of a computer program product ("CPP embodiment" or "CPP") is a term used in the present disclosure to describe a set of one or more storage media (also called "media") that are collectively included in a set of one or more storage units that collectively include machine readable code corresponding to instructions and / or data for performing computer operations specified in a particular CPP claim. A "storage unit" is any physical unit that can store instructions for use by a computer processor. Without limitation, the computer readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing storage media. Some known types of memory units that include these media include: a floppy disk, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM), a static random access memory (SRAM), a portable compact disc read only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device (such as punch-cards or pits / bumps formed in a major surface of a storage disk), and any suitable combination thereof. A computer readable storage medium as used in the present disclosure is not intended to be storage in the form of transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses propagating through an optical waveguide, or electrical signals transmitted through a wire, and / or other transmission media. Those skilled in the art will appreciate that data is typically moved at certain times during normal operation of a storage unit, such as during access, defragmentation or garbage collection, but this does not make the storage unit a volatile unit because the data is non-volatile during storage.The data processing environment 100 includes an example of an environment for executing at least a portion of the computer code involved in performing the methods of the invention, such as the broadcast and scatter operations 150. In addition to block 150, the computing environment 100 includes, for example, the computer 101, the wide area network (WAN) 102, the end user unit (EUD) 103, the remote server 104, the public cloud 105, and the private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including processing circuitry 120 and cache 121), a communication structure 111, volatile memory 112, persistent storage 113 (including operating system 122 and above-referenced block 150), the peripheral unit set 114 (including user interface unit set 123, memory 124, and Internet of Things (IoT) sensor set 125), and the network module 115. Remote server 104 includes remote database 130. The public cloud 105 includes the gateway 140, the cloud orchestration module 141, the host physical machine set 142, the virtual machine set 143, and the container set 144.The COMPUTER 101 may be in the form of a desktop computer, laptop computer, tablet computer, smart phone, smart watch or other portable computer, mainframe computer, quantum computer, or other form of computer or mobile device now known or developed in the future that is / are capable of executing a program, accessing a network, or querying a database, such as the remote database 130. As is well known in computer technology, the performance of a method implemented by a computer may be distributed among multiple computers and / or sites depending on the technology. On the other hand, the focus of this representation of computing environment 100 is on the detailed description of a single computer, namely computer 101, to keep the representation as simple as possible. The computer 101 may be located in a cloud, although not shown in a cloud in FIG. 1. On the other hand, the computer 101 may not be in a cloud unless expressly stated.The PROCESSOR SET 110 includes one or more computer processors of any type now known or developed in the future. The processing circuit 120 may be distributed among multiple packages, for example, among multiple coordinated integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory residing in the / n processor chip package(s) and is typically used for data or code that should be available for rapid access by the threads or cores running on processor set 110. Caches are typically organized in multiple levels depending on the relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set may be off-chip. In some computing environments, the processor set 110 may be configured to operate with qubits and perform quantum computing.Computer readable program instructions are typically loaded into the computer 101 to execute a series of operations by the processor set 110 of the computer 101 and thereby execute a computer implemented method such that the instructions so executed instantiate the methods specified in flowcharts and / or illustrative descriptions of computer implemented methods in this document (collectively referred to as "the methods of the invention"). These computer readable program instructions are stored in various types of computer readable storage media, such as cache 121 and other storage media described below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the methods of the invention. In the computing environment 100, at least some of the instructions for performing the methods of the invention may be stored in a permanent memory 113 at block 150.The data transmission structure 111 is the signal line paths via which the various components of the computer 101 can exchange data. This structure typically consists of switches and electrically conductive paths, e.g., the switches and electrically conductive paths, which include buses, bridges, physical input / output ports, and the like. Other types of signal transmission paths may also be used, e.g., fiber optic data transmission paths and / or wireless data transmission paths.The VOLATILE MEMORY 112 is any type of volatile memory now known or developed in the future. Examples include dynamic types of random access memories (RAM) or static types of RAM. Typically, the volatile memory is random access, but this is not required unless expressly indicated. In the computer 101, the volatile memory 112 is in a single package and is integrated into the computer 101, but alternatively or additionally, the volatile memory may be distributed over multiple packages and / or located outside the computer 101.The PERSISTENT MEMORY 113 is any form of non-volatile memory for computers now known or developed in the future. Non-volatile in this memory means that the stored data is maintained regardless of whether the computer 101 and / or directly the permanent memory 113 are powered. The permanent memory 113 may be a read only memory (ROM), but typically at least a portion of the permanent memory allows data to be written, erased, and rewritten. Some known forms of permanent storage include magnetic disks and semiconductor memory units. The operating system 122 may take various forms, such as various known proprietary operating systems or portable operating system interface open source operating systems using a kernel. The code contained in block 150 typically comprises at least a portion of the computer code involved in performing the methods of the invention.The PERIPHERAL UNIT SET 114 includes the peripheral unit set of the computer 101. Communications links between the peripherals and the other components of the computer 101 may be implemented in various ways, such as Bluetooth links, near-field communication (NFC) links), cable links (e.g., Universal Serial Bus (USB) cables, plug-in links (e.g., secure digital card (SD) card), connections over local communications networks, and even connections over wide area networks such as the Internet. In various embodiments, the user interface unit set 123 may include components such as a display screen, a speaker, a microphone, wearable units (e.g., glasses and smart watches), a keyboard, a mouse, a printer, a touchpad, game controllers, and haptic units. The memory 124 is an external memory such as an external hard disk or a removable memory such as an SD card. The memory 124 may be permanent and / or volatile. In some embodiments, the memory 124 may take the form of a quantum processing memory unit for storing data in the form of qubits. In embodiments where the computer 101 needs a large amount of storage space (e.g., when the computer 101 locally stores and manages a large database), that storage space may be provided by peripheral storage units configured to store very large amounts of data, e.g., a storage area network (SAN) shared by multiple geographically distributed computers. The loT sensor set 125 consists of sensors that can be used in applications for the Internet of Things. One sensor can be, for example, a thermometer and another a motion detector.The NETWORK MODULE 115 is the combination of computer software, hardware, and firmware that allows the computer 101 to exchange data with other computers via a WAN 102. The network module 115 may include hardware, e.g., modems or WLAN transceivers, software for packetizing and / or depacketizing data for transmission over a communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, network control and forwarding functions of network module 115 are performed on the same physical hardware unit. In other embodiments (e.g., embodiments using software defined networking (SDN)), the control and forwarding functions of the network module 115 are performed on physically separate entities, such that the control functions manage multiple different hardware entities of the network. Computer readable program instructions for performing the methods of the present invention may typically be downloaded from an external computer or storage device to the computer 101 via a network adapter card or network interface included in the network module 115.The WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any technology for transmitting computer data that is already known or will be developed in the future. In some embodiments, the WAN may be replaced and / or supplemented with local area networks (LANs) configured to transmit data between devices in a local area, e.g., in a WLAN network. The WAN and / or LANs typically include computer hardware such as copper transmission cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.The end user device (EUD) 103 is a computer system used and controlled by an end user (e.g., a customer of a company operating the computer 101) and may take any of the forms described above in connection with the computer 101. The EUD 103 typically receives helpful and useful data from the operations of the computer 101. In a hypothetical case, for example, where the computer 101 is configured to provide a recommendation to an end user, this recommendation would typically be transmitted from the network module 115 of the computer 101 to the EUD 103 via the WAN 102. In this manner, the EUD 103 may display or otherwise present the recommendation to an end user. In some embodiments, the EUD 103 may be a client device, e.g., a thin client, heavy client, mainframe, desktop computer, etc.The remote server 104 is any computer system that provides at least some data and / or functions to the computer 101. Remote server 104 may be controlled and used by the same entity that also operates computer 101. Remote server 104 represents the machine(s) that collect / collect and store helpful and useful data for use by other computers, e.g., computer 101. For example, in a hypothetical case where the computer 101 is designed and programmed to make a recommendation based on historical data, that historical data may be provided to the computer 101 from the remote database 130 of the remote server 104.The public cloud 105 is a computer system that is available to multiple entities and enables demand-dependent availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically shares resources to achieve coherence and scale effects. The direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments executing on various computers that form the computers of the set of host physical machines 142, which represents the entirety of the physical computers in and / or for the public cloud 105. The virtual computing environments (VCEs) typically take the form of virtual machines from the set of virtual machines 143 and / or containers from the container set 144. It should be noted that these VCEs may be stored as mappings (images) and transmitted between the various physical machine hosts, either as mappings or after instantiating the VCE. The cloud orchestration module 141 manages the transfer and storage of mappings, implements new VCE instantiations, and manages active instantiations of VCE implementations. Gateway 140 is a combination of computer software, hardware, and firmware that enables public cloud 105 to transmit data over WAN 102.Some further explanation will now be given on virtualized computing environments (VCEs). VCEs may be stored as "maps.". From the mapping, a new active instance of the VCE may be created. Two known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system level virtualization. This relates to a function of the operating system in which the kernel allows the presence of several isolated instances in the user area, so-called containers. These isolated instances in the user area typically behave like real computers from the point of view of the programs executed therein. A computer program running on a common operating system may use all resources of that computer, e.g., connected devices, files and folders, shared network areas, CPU power, and quantifiable hardware capabilities. However, programs executing in a container may only use the contents of the container and the units assigned to the container. This function is referred to as containerization.The PRIVATE CLOUD 106 is similar to the PUBLIC CLOUD 105, except that the computing resources are available for use by only a single enterprise. Although the private cloud 106 is shown as being in communication with the WAN 102, in other embodiments, a private cloud may be completely disconnected from the Internet and only accessible via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public clouds) that are often implemented by different vendors. Each of the multiple clouds remains a separate and stand-alone entity, but the larger architecture of the hybrid cloud is interconnected by standardized or proprietary technology that enables orchestration, management, and / or data / application portability between the involved clouds. In this embodiment, the public cloud 105 and the private cloud 106 are both part of a larger hybrid cloud.Referring now to FIG. 2, a block diagram of a tree structure 200 for use in connection with one or more embodiments of the present invention is shown. In example embodiments, tree structure 200 includes a root node 202, also referred to as a source node or an original participant, and a plurality of children nodes 204, also referred to as participant nodes. As shown, one or more of the plurality of children 204 may include children 206 that depend on child 204, and children 206 may include children 208 that depend on child 206. In example embodiments, each of the nodes (root node 202 and children nodes 204, 206 and 208) may reside in a computer 101 as shown in FIG. 1.In example embodiments, prior to initiating a distributed data transfer operation, such as a broadcast or spread operation, the original subscriber 200 retrieves a tree structure 200 that is used for the distributed data transfer operation. In one embodiment, tree structure 200 is known only to original subscriber 202 and not to the other subscriber nodes. In example embodiments, the tree structure 200 may be different for the same subscriber group, depending on the data transfer operation itself. The tree structure need not be in a regular pattern, but may be irregular based on external inputs such as the conditions of the network. The method for creating the tree structure 200 is outside the scope of this invention, and any of a variety of known techniques may be used to create the tree structure 200.Referring now to FIG. 3, a block diagram illustrating a broadcast operation according to one or more embodiments of the present invention is illustrated. In a broadcast operation, the original subscriber 300 adds a header 302 to the payload of the message at a known location either before or after the payload 306 of the message. Header 302 describes the base address of payload 306, the length of payload 306, and children tree structure 308 for the receiving party ("children"). As seen in the illustration, the headers 312, 322, 342 transmitted to the children 310, 320, and 340, respectively, are different from each other because each contains a different child tree structure 308. In one embodiment, a cached child tree structure is used and a marker identifying this cached child tree is sent to the receiving party instead of the child tree structure.In example embodiments, once original subscriber 300 has assembled payload 306 and headers, it starts the broadcast operation. The original subscriber 300 transmits the header 312 and payload 306 to the child node 310, the header 322 and payload 306 to the child node 320, and the header 342 and payload 306 to the child node 340.In example embodiments, a participating child node receives the payload from its parent node that it does not know prior to the beginning of the message log. The participating child node checks the payload to determine the structure of the payload and the shape of the underlying child tree, if any. In example embodiments, a participating child node, when receiving the header and payload 306, corrects the header of information of the child tree that does not relate to the child tree to which it is sending, and transmits a new header and payload to its child nodes. For example, after node 310 receives header 312 and payload 306, node 310 corrects header 312 to generate headers 332 and 352, which are transmitted along with payload 306 to nodes 330 and 350, respectively.If the subordinate node is an end node, e.g. the node 370, the transmission of the payload 306 ends. When an acknowledgement is requested, each node sends an acknowledgement message to its immediate parent node (i.e., to the node from which it received the message). If the child node is not an end node and an acknowledgement is requested, the node waits until it receives acknowledgement messages from its child tree before forwarding this acknowledgement to its immediate parent node for this message. In example embodiments, the forwarding pattern continues until all subscribers have received the payload destined for them and sent all required acknowledgments.Referring now to FIG. 4, a block diagram illustrating a scatter operation 400 according to one or more embodiments of the present invention is illustrated. In example embodiments, when the original subscriber 402 initiates a scatter operation 400, it retrieves a tree structure to be used for the scatter operation 400. The tree structure comprises a plurality of nodes 403, 404, 405, 406, 407, 407, 408, and 409.After the tree structure is retrieved, the original subscriber 402 generates a message 410 that includes multiple headers 412, 416, 424 and multiple payloads 414, 418, 420, 422, 426, 428, 430. In example embodiments, the original subscriber 402 generates the payload 414, 418, 420, 422, 426, 428, 430 for each node 403, 404, 405, 406, 407, 408, and 409 in the tree structure. Similarly, the original subscriber 402 generates a header 412, 416, 424 for each node 403, 406, 404 of the tree structure that includes at least one child node. In example embodiments, for each of the nodes 403, 406, 404, the header 412, 416, 424 comprises a description of the children of the tree structure that depends on the node 403, 406, 404. The headers 412, 416, 424 may also describe the base address of the payload 414, 418, 420, 422, 426, 428, 430 and the length of the payload 414, 418, 420, 422, 426, 428, 430.In example embodiments, the message 410 is generated by the original participant 402 such that the portions of the message 410 transmitted to each child node are contiguous. For example, as seen in the illustration, headers 412, 416 and payload 414, 418, 420, 422 transmitted to child node 403 are contiguous. Similarly, header 424 and payload 426, 428 transmitted to child node 404 are contiguous. In one embodiment, a cached child tree structure is used and a marker identifying this cached child tree is sent to the receiving party instead of the child tree structure.Once a child node receives a portion of message 410, the child node is configured to extract the data required by the child node and split the remaining headers and payload using the information from the header for the child node. For example, once child node 403 receives the portion of the message from original subscriber 402, child node 403 extracts payload 414 needed by child node 403 and uses the information in header 412 to separate the remaining portion of message 410 into separate portions. The child node then forwards a part of the message, i.e. it sends only the subset of the headers and payloads destined for a particular child tree to that child tree. For example, child node 403 transmits header 416 and payload 418, 420 to child node 406 and payload 422 to child node 407.Referring now to FIG. 5, a block diagram illustrating a combination of a broadcast operation and a spread operation according to one or more embodiments of the present invention is illustrated. In example embodiments, when the original subscriber 502 initiates a combined broadcast / scatter operation 500, he retrieves a tree structure to be used for the combined broadcast / scatter operation 500. The tree structure includes a plurality of nodes 503, 504, 505, 506, 507, 508 and 509.After the tree structure is retrieved, the original subscriber 502 generates a message 510 that includes a broadcast header 512, broadcast payload 514, multiple scatter headers 516, 520, 528, and multiple scatter payloads 518, 522, 524, 526, 530, 532, 534. In example embodiments, the original subscriber 502 generates the scatter payload 518, 522, 524, 526, 530, 532, 534 for each node 503, 504, 505, 506, 507, 508 and 509 in the tree structure. Similarly, the original subscriber 502 generates a scatter header 516, 520, 528 for each node 503, 506, 504 of the tree structure that includes at least one child node. In example embodiments, for each of the nodes 503, 506, 504, the scatter header 516, 520, 528 includes a description of the sub-tree of the tree structure that depends on the node 503, 506, 504. The scatter headers 516, 520, 528 may also describe the base address of the scatter payload 518, 522, 524, 526, 530, 532, 534 and the length of the scatter payload 518, 522, 524, 526, 530, 532, 534.In example embodiments, message 510 is generated by original participant 502 such that the portions of message 510 transmitted to each child node are contiguous. For example, as seen in the illustration, headers 516, 520 and payload 518, 522, 524, 526 transmitted to child node 503 are contiguous. Similarly, header 528 and payload 530, 532 transmitted to child node 504 are contiguous.Once a child node receives a portion of message 510, the child node is configured to check the header corresponding to the child node to determine the structure of the payload and the shape of the child tree under the child node, if any. The child node is further configured to extract a copy of the broadcast payload 514 for its need and remove the scatter payload corresponding to the child node. For example, child node 503 checks broadcast header 512 and extracts a copy of broadcast payload 514, checks scatter header 516 and extracts scatter payload 518. Based on the information in broadcast header 512 and broadcast header 516, child node 503 generates messages and sends them to children nodes 506 and 507.In exemplary embodiments, the child node then forwards only a portion of the message to each child node that depends on it, i.e., the child node sends only the subset of the headers and payload destined for a particular child tree to that child tree. For example, child node 503 transmits broadcast header 512, broadcast payload 514, scatter header 520, and scatter payload 522, 524 to child node 506, and broadcast header 512, broadcast payload 514, and scatter payload 526 to child node 507.In example embodiments, each child node may be configured to add additional data to the header and / or payload that is forwarded to its child tree. Moreover, each child node may be configured to change the tree structure for its child tree. For example, a child node may know that a node in its child tree is offline or has an unexpected performance problem. In this case, the child node can replace that node in its child tree with another node.Referring now to FIG. 6, a flow diagram of a method for initiating a distributed data transfer operation according to one or more embodiments of the present invention is shown. In example embodiments, the distributed data transfer operation is either a broadcast operation or a spread operation, or a combination of broadcast and spread operations. As seen in the diagram, the method 600 includes, as shown in block 602, receiving a request to perform a distributed data transfer operation. As shown in block 604, the method 600 next includes retrieving a tree structure to perform the distributed data transfer operation. In example embodiments, a data processing system that initiates a distributed data transfer operation is a root node of the tree structure.As shown in block 606, the method 600 also includes generating a message including header information and payload data for the distributed data transfer operation. In example embodiments, the message is generated by organizing the header information and the payload based on the tree structure. In one embodiment, the header information and payload are organized such that a portion of the header information and a portion of the payload to be transmitted to a child node are contiguous. The method 600 further includes, as shown in block 608, transmitting a portion of the message to each child node of the first data processing system, wherein the portion transmitted to each child node is unique. In example embodiments, the portion of the message transmitted to each child node includes a child node header defining a child node tree structure.In one embodiment, the distributed data transfer operation is a broadcast operation, and the payload of the message transmitted to each child node comprises broadcast payload that is the same for each child node. In another embodiment, the distributed data transfer operation is a spreading operation, and the portion of the message transmitted to each child node includes spreading payload data retrieved based on the payload data. The scatter payload transmitted to each child node is different from the scatter payload transmitted to the other children nodes.Referring now to FIG. 7, a flow diagram of a method 700 for performing distributed data transfer operations according to one or more embodiments of the present invention is shown. As shown in block 702, the method 700 begins with a child node receiving a message from a parent node via a distributed data transfer operation. In example embodiments, the distributed data transfer operation is either a broadcast operation or a spread operation, or a combination of broadcast and spread operations. Once the child node receives the distributed data processing operation message, it retrieves information about its child tree from a header of the distributed data processing operation message. Next, as illustrated in decision block 704, the method 700 determines whether the child node is an end node based on the information about the child tree.Based on determining that the child node is an end node, the method 700 proceeds to block 710 and the child node extracts the payload needed by the child node. Based on determining that the child node is not an end node, the method 700 proceeds to block 706, and the child node queries data of the child tree from a header of the message and generates a message for each node that depends on the child node. In example embodiments, the message generated for each node includes only header information and payload data required for the child tree corresponding to the destination child node. The message transmitted to each child node may include a broadcast header, a scatter header, broadcast payload, and / or scatter payload. Next, as illustrated in block 708, the method 700 includes transmitting messages to each node that depends on the child node. As shown in block 712, the method 700 further includes transmitting an acknowledgement message to the parent node.The technical advantages and benefits include methods, systems, and computer program products for performing distributed data transfer operations using execution trees configured to dynamically adapt based on network and data processing conditions. The method for performing distributed data transfer operations only requires that the original party or root node know the tree structure used to perform the distributed data transfer operation. Thus, different tree structures may be used to perform different distributed data transfer operations, and the tree structures may be updated without synchronizing knowledge of the tree structure with each node of the tree.Various embodiments of the invention will be described herein with reference to the accompanying drawings. Alternative embodiments of the invention may be developed without departing from the scope of this invention. In the following description and drawings, various connections and positional relationships (e.g., above, below, next, etc.) between the elements are set forth. Unless otherwise indicated, these connections and / or positional relationships may be direct or indirect, and the present invention is not intended to be limiting in this respect. Accordingly, a connection of entities may refer to either a direct or indirect connection, and a positional relationship between entities may be a direct or indirect positional relationship. Moreover, the various tasks and process steps described herein may be incorporated into a more comprehensive method or process having additional steps or additional functionality, which are not described in detail herein.One or more of the methods described herein may be embodied with any one or combination of the following technologies known in the art: one or more discrete logic circuit(s) having logic gates for implementing logic functions on data signals, an application specific integrated circuit (ASIC) having suitable combinatorial logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.For the sake of brevity, conventional techniques associated with making and using aspects of the invention will not always be described in detail herein. In particular, various aspects of data processing systems and specific computer programs for implementing the various technical features described herein are known. Therefore, for brevity, numerous conventional implementation details are only briefly mentioned or omitted altogether without providing the known system and / or process detailsIn some embodiments, various functions or actions may occur at a particular location and / or in connection with the operation of one or more devices or systems. In some embodiments, a portion of a particular function or action may be performed at a first entity or location, and the remainder of the function or action may be performed at one or more additional entities or locations.The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a / an / an" and "the / s" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprise" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the following claims are intended to include all structures, materials, or acts for performing the function in combination with other claimed elements as specifically claimed. The present disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the form disclosed. It will be apparent to those skilled in the art that many changes and modifications are possible without departing from the scope of the disclosure. The embodiments were chosen and described in order to best explain the principles of the disclosure and the practical application, and to enable others skilled in the art to understand the disclosure for various embodiments with various modifications as are suited to the particular use contemplated.The diagrams illustrated herein are illustrative. There may be many modifications to the diagrams or steps (operations) described therein without departing from the scope of the disclosure. For example, the steps may be performed in a different order, or steps may be added, omitted, or changed. The term "connected" further describes a signal path between two elements and does not imply a direct connection between the elements without intervening elements / connections. All of these modifications are considered part of the present disclosure.For interpreting the claims and the description, the following definitions and abbreviations are to be used. As used herein, the terms "comprise," "comprise," "have," "contain," or variations thereof are intended to include non-exclusive inclusion. For example, a composition, a mixture, a process, a method, an article, or a device having / comprising a list of elements is not necessarily limited to these elements, but may also include other elements not expressly listed or associated with a / such composition, mixture, process, method, article, or device.Moreover, the term "exemplary" is used herein to mean "serving as an example, case, or illustration.". Embodiments or designs described herein as "exemplary" are not necessarily to be considered preferred or advantageous over other embodiments or designs. By the terms "at least one / one" and "one / more" is meant an integer greater than or equal to one, i.e., one, two, three, four, etc. By the term "a plurality" is meant an integer greater than or equal to two, i.e., two, three, four, five, etc. The term "connection" may include both an indirect "connection" and a direct "connection".The terms "about", "substantially", "approximately", and variations thereof are intended to include the degree of error associated with measuring the particular quantity based on the equipment present at the time of filing. For example, "about" may comprise a range of ± 8% or 5%, or 2% of a particular value.The present invention can be a system, a method and / or a computer program product at any possible technical level of detail of the integration. The computer program product may comprise computer readable storage medium(s) on / on which computer readable program instructions are / are stored for causing a processor to carry out aspects of the present invention.The computer readable storage medium may be a physical device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage unit, a magnetic storage unit, an optical storage unit, an electromagnetic storage unit, a semiconductor storage unit, or any suitable combination thereof. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions stored thereon, and any suitable combination thereof. A computer readable storage medium, as used herein, is not intended to be transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through a wire.Computer readable program instructions described herein may be downloaded from a computer readable storage medium to respective data processing / processing devices or via a network such as the Internet, a local area network, a wide area network, and / or a wireless network to an external computer or storage device. The network may include copper transmission cables, lightwave transmission conductors, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing unit receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the corresponding computing / processing unit.Computer readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, integrated circuit configuration data, or either source code or object code written in any combination of one or more programming languages, including object oriented programming languages such as Smalltalk, C++, or the like, and procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, via the Internet using an Internet Service Provider). In some embodiments, electronic circuits, including, for example, programmable logic circuits, field programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuits to perform aspects of the present invention.Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It is noted that each block of the flowcharts and / or the block diagrams, as well as combinations of blocks in the flowcharts and / or the block diagrams, may be executed by computer readable program instructions.These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus produce a means for implementing the functions / steps specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored on a computer readable storage medium that can control a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored thereon comprises an article of manufacture including instructions that implement aspects of / the function / step specified in the flowchart and / or block diagram block or blocks.The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of process steps to be performed on the computer or other programmable apparatus or other device to produce a computer executed process such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / steps specified in the flowchart and / or block diagram block or blocks.The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions that include one or more executable instructions for executing the particular logical function(s). In some alternative implementations, the functions specified in the block may occur in a different order than shown in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending on the appropriate functionality. It is further noted that each block of the block diagrams and / or flowcharts, as well as combinations of blocks in the block diagrams and / or flowcharts, may be implemented by special purpose hardware-based systems that perform the specified functions or steps, or perform combinations of special purpose hardware and computer instructions.The descriptions of the various embodiments of the present invention have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments. It will be apparent to those skilled in the art that many changes and modifications are possible without departing from the scope of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application, or technical improvement over technologies in the market, or to enable those skilled in the art to understand the embodiments described herein.
Claims
A computer implemented method comprising: receiving a request from a first data processing system to perform a distributed data transfer operation; retrieving a tree structure from the first data processing system to perform the distributed data transfer operation, wherein the first data processing system is a root node of the tree structure; generating a message by the first data processing system with header information and payload for the distributed data transfer operation; and transmitting a portion of the message by the first data processing system to each child node of the first data processing system, wherein the portion transmitted to each child node is unique.The method of claim 1, wherein the portion of the message transmitted to each child node comprises a child node header defining a child node's tree structure.The method of claim 1, wherein the distributed data transfer operation is a broadcast operation, and wherein the payload of the message transmitted to each child node comprises broadcast payload.The method of claim 1, wherein the distributed data transfer operation is a scatter operation, and wherein the portion of the message transmitted to each child node comprises scatter payloads retrieved based on the payload.The method of claim 1, wherein the distributed data transfer operation is a broadcast and spread operation.The method of claim 1, wherein the message is generated by organizing the header information and the payload based on the tree structure.The method of claim 6, wherein the header information and the payload are organized such that a portion of the header information and a portion of the payload to be transmitted to a child node are contiguous.A system comprising: a memory comprising computer readable instructions; and one or more processors to execute the computer readable instructions, wherein the computer readable instructions control the one or more processors to perform operations comprising: receiving a request from a first data processing system to perform a distributed data transfer operation; retrieving a tree structure from the first data processing system to perform the distributed data transfer operation, wherein the first data processing system is a root node of the tree structure; generating, by the first data processing system, a message including header information and payload data for the distributed data transfer operation; and transmitting, by the first data processing system, a portion of the message to each child node of the first data processing system, wherein the portion transmitted to each child node is unique.The system of claim 8, wherein the portion of the message transmitted to each child node comprises a child node header defining a child node's tree structure.The system of claim 8, wherein the distributed data transfer operation is a broadcast operation, and wherein the payload of the message transmitted to each child node comprises broadcast payload.The system of claim 8, wherein the distributed data transfer operation is a scatter operation, and wherein the portion of the message transmitted to each child node comprises scatter payloads retrieved based on the payload.The system of claim 8, wherein the distributed data transfer operation is a broadcast and spread operation.The system of claim 8, wherein the message is generated by organizing the header information and the payload based on the tree structure.The system of claim 13, wherein the header information and the payload are organized such that a portion of the header information and a portion of the payload to be transmitted to a child node are contiguous.A computer program product comprising a computer readable storage medium having program instructions embodied thereon, the program instructions executable by a processor to cause the processor to perform operations comprising: receiving a request from a first data processing system to perform a distributed data transfer operation; retrieving a tree structure from the first data processing system to perform the distributed data transfer operation, wherein the first data processing system is a root node of the tree structure; generating a message by the first data processing system with header information and payload for the distributed data transfer operation; and transmitting a portion of the message by the first data processing system to each child node of the first data processing system, wherein the portion transmitted to each child node is unique.The computer program product of claim 15, wherein the portion of the message transmitted to each child node comprises a child node header defining a child node tree structure.The computer program product of claim 15, wherein the distributed data transfer operation is a broadcast operation, and wherein the payload of the message transmitted to each child node comprises broadcast payload.The computer program product of claim 15, wherein the distributed data transfer operation is a scatter operation, and wherein the portion of the message transmitted to each child node comprises scatter payloads retrieved based on the payload.The computer program product of claim 15, wherein the distributed data transfer operation is a broadcast and spread operation.The computer program product of claim 15, wherein the message is generated by organizing the header information and the payload based on the tree structure.