Asymmetric data communication for host-device interface

By identifying the asymmetric bandwidth requirements of hardware devices in the system and dynamically configuring the bus channel, the problem of inefficient asymmetric data transmission in the prior art is solved, and data communication efficiency and system computing efficiency are improved.

CN113841132BActive Publication Date: 2025-05-02GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980096356.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-21
Filing Date
2019-11-15
Publication Date
2025-05-02
Estimated Expiration
2039-11-15

AI Technical Summary

Technical Problem

In the prior art, the asymmetry of data bandwidth requirements for equipment such as machine learning accelerators is not reflected in system interconnection sockets, component interfaces or hardware connections, resulting in inefficient allocation of symmetric bus channels at component interfaces.

Method used

By identifying the coupled hardware devices in the system, generating system topology, dynamically configuring asymmetric links, and dynamically allocating bus channels to meet asymmetric data transmission needs based on data transmission mode and bandwidth requirements.

Benefits of technology

It improves data communication efficiency, reduces the number of bus channels, more accurately reflects the actual asymmetric use, and improves the computing efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113841132B_ABST
    Figure CN113841132B_ABST
Patent Text Reader

Abstract

Disclosed are methods, systems, and apparatus for performing asymmetric data communications at a host-device interface of a system, including a computer program encoded on a computer storage medium. The method includes identifying a device coupled to a host of the system, and a system topology, the system topology generating a bus channel that identifies the connectivity of the device and enables data transmission at the system. The host determines that a connection between the host and a first device of a plurality of devices has an asymmetric bandwidth requirement. The host configures a set of bus channels of a data bus connecting the first device and the host to allocate a different number of bus channels for data outgoing from the host than for data entering to the host. The bus channels are configured to allocate different numbers of bus channels based on the asymmetric bandwidth requirement of the first connection.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority under 35 USC §119(e) to U.S. patent application serial number 62 / 851,052 filed on May 21, 2019. The entire contents of U.S. patent application serial number 62 / 851,052 are incorporated herein by reference in their entirety. Background Art

[0003] The present application generally relates to asymmetric data communications for various component interfaces of a system, such as a host device interface.

[0004] Devices such as machine learning accelerators, storage components, video transcoding accelerators, or neural network processors often have asymmetric bandwidth requirements. In some cases, when these devices are connected to a component of a system such as a host, the asymmetry of the bandwidth corresponds to an imbalance in the amount of data exchanged in a particular direction at the host-side interface of the system.

[0005] For example, the ingress data bandwidth of a machine learning accelerator may be ten times greater than the egress data bandwidth of the accelerator. The ingress data bandwidth of a machine learning accelerator may correspond to when a host sends a large amount of data to the accelerator to perform accelerated inference calculations at the accelerator, while the egress data bandwidth of the accelerator may correspond to when the accelerator sends a small amount of data to the host to indicate the result of the inference calculation. Summary of the invention

[0006] Asymmetry in data bandwidth requirements at a system is often not reflected in the configuration of the system interconnect sockets, component interfaces, or hardware connections. For example, current interconnect standards such as the Peripheral Component Interconnect Express (PCI-e) standard allocate the same number of data bus lanes for host-to-device communications as they do for device-to-host communications. Symmetric bus lane allocations at component interfaces lead to inefficiencies when there is an asymmetry between the amount of data transferred in either direction at the host-device interface.

[0007] Thus, this article describes a technique for implementing a software control loop for dynamically configuring asymmetric links at respective interconnected locations of a system. The technique includes identifying respective hardware devices of at least a host coupled to the system. The host is operable to generate a system topology identifying the connectivity of the respective devices. Information associated with the connectivity of the devices is used to determine the hardware configuration of the devices, including the asymmetric data transmission capabilities of the devices. The system topology also identifies the bus channels of the system and the asymmetric links of each device. The software loop references the connectivity of each device and the system topology to configure the asymmetry of bidirectional data transmission at the system.

[0008] One aspect of the subject matter described in this specification can be embodied in a method, the method comprising identifying a plurality of devices coupled to a host of a system, and generating a system topology identifying connectivity of the plurality of devices and identifying bus channels that enable data transfer at the system. The method also includes determining that a first connection between the host and a first device of the plurality of devices has an asymmetric bandwidth requirement. The method further includes configuring a first bus channel set of a first data bus connecting the first device and the host based on the asymmetric bandwidth requirement of the first connection, so as to allocate a different number of bus channels in the first bus channel set for data outgoing from the host and data entering to the host. Optionally, the bus channels are configured to allocate different numbers of bus channels based on the asymmetric bandwidth requirement of the first connection.

[0009] These and other implementations may optionally include one or more of the following features. For example, in some implementations, the method further includes: determining that a second connection between the host and a second device among the plurality of devices has an asymmetric bandwidth requirement; and configuring a second bus channel set of a second data bus connecting the second device and the host based on the asymmetric bandwidth requirement of the second connection, so as to allocate a different number of bus channels in the second bus channel set for data outgoing from the host and data incoming to the host.

[0010] The method may further include: determining a data transmission pattern at the system using the system topology; calculating an asymmetric bandwidth requirement for the first connection based on the data transmission pattern; and calculating an asymmetric bandwidth requirement for the second connection based on the data transmission pattern.

[0011] The method may further include: providing information describing data traffic at the system to a software agent; using the software agent to determine a data transmission pattern at the system based on a statistical analysis of the information or an inferential analysis of the information; using the software agent to generate a prediction indicating the distribution of data traffic for processing one or more workloads at the system; and calculating an asymmetric bandwidth requirement for the first connection based on the prediction indicating asymmetric data traffic at the first connection.

[0012] The method may further comprise calculating an asymmetric bandwidth requirement for the second connection based on a prediction indicative of asymmetric data traffic at the second connection.

[0013] In some implementations, each bus channel in the first set of bus channels is dynamically configured as a data-in channel or a data-out channel; and each bus channel in the second set of bus channels is dynamically configured as a data-in channel or a data-out channel.

[0014] The method may further include exchanging data between the host and the first device using bus channels in the first bus channel set allocated to data egress from the host and bus channels in the first bus channel set allocated to data ingress to the host.

[0015] In some implementations, the asymmetric bandwidth requirement of the first connection includes an M:N ratio of incoming bus lanes relative to outgoing bus lanes; and M has an integer value greater than an integer value of N.

[0016] In some implementations, the asymmetric bandwidth requirement of the second connection includes an N:M ratio of outgoing bus lanes relative to incoming bus lanes; and N has an integer value greater than the integer value of M.

[0017] In some implementations, the system includes a processor and an accelerator, and the method further includes: configuring the processor as the host; identifying the accelerator as the first device; and determining that the accelerator is configured to have connectivity including a bus channel configured for bidirectional data transfer with the host via the first connection.

[0018] In some implementations, the system includes a memory, and the method further includes: identifying the memory as the second device; and determining that the memory is configured to have connectivity including a bus channel configured for bidirectional data transfer with the host via the second connection.

[0019] Other implementations of this and other aspects include corresponding systems, apparatuses, and computer programs encoded on a non-transitory computer-readable storage device configured to perform the actions of the method. The system of one or more computers may be configured by software, firmware, hardware, or a combination thereof installed on the system operating to cause the system to perform the actions. One or more computer programs may be so configured by having instructions that, when executed by a data processing device, cause the device to perform the actions.

[0020] The subject matter described in this specification can be implemented in specific embodiments to achieve one or more of the following advantages. The techniques described herein can be used to implement an asymmetric configuration of bus channels in a data bus connection between devices of a system. The asymmetric configuration is based on an asymmetric bandwidth requirement generated by a system host using a derivation or prediction learned from analyzing a data traffic pattern of the system. The predictive analysis of the traffic pattern can produce an asymmetric bandwidth requirement that accurately reflects the traffic flow at the component interface of the system.

[0021] By using these techniques, the software control loop can more efficiently allocate the data transmission capacity of a given set of bus channels based on the traffic flow patterns observed at the system, thereby reflecting the actual asymmetric usage. Therefore, because the system can more accurately determine the relative magnitude of traffic, the system can be designed to include fewer bus channels for providing asymmetric data links. The system can also adjust the asymmetric configuration of certain bus channels throughout the system to achieve greater data communication efficiency.

[0022] Details of one or more implementations of the subject matter described in this specification are given in the accompanying drawings and the following description. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a block diagram of an example system for performing asymmetric data communications.

[0024] Figure 2 Shown for use Figure 1 An example process for a system to perform asymmetric data communication.

[0025] Figure 3 A list of devices of a system used to perform asymmetric data communication is shown.

[0026] Figure 4 Shown includes Figure 1 An example graph of information related to data traffic at a system.

[0027] Figure 5 An example hardware connection with asymmetric bidirectional bandwidth capability is shown.

[0028] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION

[0029] Figure 1 1 is a block diagram of an example hardware computing system 100 for performing asymmetric data communication. System 100 generally includes a first processor 102, a second processor 104, a first device 106, and a second device 108. Each of processor 102 or processor 104 can be configured as a system host ("host") that includes a host interface for coupling or interconnecting various hardware devices at system 100. For example, system 100 can include one or more dedicated hardware circuits, each of which includes a plurality of interconnect locations, such as interconnect sockets or card slots, for establishing connections with devices 106 and 108.

[0030] Processors 102 and 104 are central processing units (CPUs) or graphics processing units (GPUs) that form at least part of the hardware circuitry for executing software routines or control functions of the system. The host of system 100 may be represented by a single processor, multiple processors, or multiple processors of different types, such as a CPU, GPU, or a special purpose processor such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). Thus, although Figure 1 Two processors are shown in FIG. 1 , but system 100 may include multiple processors or specialized components.

[0031] In one implementation, processor 102 is configured as a main processor of a host, while processor 104 is configured as a secondary processor of the host. In another implementation, processors 102 and 104 are both configured as co-main processors of the host. The host may be a domain or software program that executes a software loop 130 (described below) for data processing operations, data services, asymmetric data flows, and different bandwidth requirements at the system 100. In some cases, the host is an example operating system (OS) running on one or more processors of the system 100. Further, in some examples described below, processor 102 may be referred to as host 102 to indicate an instance in which processor 102 is configured as a host or a main processor / co-main processor of a host.

[0032] Each of device 106 and second device 108 may be a respective hardware device (e.g., a peripheral device) that interacts with one or more components of system 100 at various interconnect locations of system 100. These interconnect locations may correspond to component interfaces of system 100. In some examples, interaction with a hardware or peripheral device via a component interface of system 100 may be described with reference to software, such as an application, script, or program, running on the device.

[0033] The system 100 includes a component interface 110 that defines a connection between a processor 102 (e.g., a host or host 102) and a device 106. Similarly, the system 100 includes another component interface 112 that defines a connection between a processor 104 and a device 108. In some implementations, the devices 106 and 108 are example peripheral devices that are each uniquely configured to include a bidirectional data link (described below), such as a bidirectional bus channel or communication channel, that enables various types of symmetric and asymmetric data transfers at the system 100.

[0034] The system 100 may include another component interface 114 that defines a connection between the processor 102 and the processor 104. In some implementations, the processor 102 is a host processor and the processor 104 is a co-processor used by the host 102. The system 100 may also include another component interface 116 that defines a connection between the processor 102 and a first type of memory 118 used by the host. The system 100 may also include another component interface 120 that defines a connection between the processor 102 and a second different type of memory 122 also used by the host. The system 100 may also include another component interface 124 that defines a connection between the processor 104 and a memory 126 used by the processor 104.

[0035] Each of the memories 118, 122, and 126 may be a different type of memory, the same type of memory, a combination of different types of memory, or part of the same memory structure. For example, each of the memories 118, 122, and 126 may be a dynamic / static random access memory (DRAM / SRAM), a non-volatile memory (NVM), an electrically erasable programmable read-only memory (EEPROM), a read-only memory (ROM), flash memory, or other known types of computer-readable storage media. In some implementations, the memory 118 and the memory 126 may be the same type of memory, or may be sub-segments of the same memory structure.

[0036] In general, system 100 may include a plurality of component interfaces, and each component interface is configured to allow data traffic to flow symmetrically or asymmetrically between components at the interface. In some examples, the component interface is a socket or card slot capable of receiving a card to interconnect or add internal components at system 100. For example, first component interface 110 may receive a PCIe 4.0 card to establish an asymmetric connection between a host such as processor 102 and a GPU hardware accelerator such as device 106.

[0037] The second component interface 112 may accommodate an example network card to establish an asymmetric connection between a secondary host processor such as the processor 104 and a networked hardware device such as the device 108. In an alternative implementation, the component interface 112 (or 110) may accommodate an NVLink bridge card to establish an asymmetric connection between a first GPU hardware accelerator such as the processor 104 and a second GPU hardware accelerator such as the device 108.

[0038] When processor 102 is configured as a host, interface 110 may represent a host-device interface corresponding to an interconnect location or connection point where data traffic flows asymmetrically between host 102 and device 106. Likewise, when system 100 configures processor 104 as a secondary host processor or co-primary host processor, interface 112 may represent an additional host-device interface corresponding to another interconnect location or connection point where data traffic also flows asymmetrically between host 102 and device 108.

[0039] System 100 includes one or more data buses. The data bus provides various interconnected data communication paths for routing data between various components of system 100. Each component interface of system 100 can be associated with a data bus, and each data bus includes a plurality of bus channels. Each data bus can have a bus channel set, which corresponds to a line set of a medium for transporting data within the system. A single bus channel in the bus channel set forming the data bus can be configured as a bidirectional link at system 100. Each component interface includes a bus channel set associated with a specific data bus of system 100. For example, the bus channels of component interfaces 110 and 112 can each be associated with data bus_1 of system 100, while the bus channels of component interface 114 can each be associated with data bus_2 of system 100.

[0040] To configure the directionality of a bus channel at a component interface, the system 100 may obtain a hardware definition file that defines the connectivity of one or more devices coupled at the interface. The hardware definition file describes or identifies various symmetric or asymmetric data link capabilities of an example peripheral device such as device 106 or 108. For example, the hardware definition file may indicate whether a bus link or interconnection point in a peripheral device is configured for bidirectional data transfer. In some implementations, the interconnection point of the peripheral device interacts with a bus channel included at the component interface to support asymmetric data transfer. For example, using software loop 130, the peripheral device may be configured to perform asymmetric data transfer using the corresponding bidirectional capabilities of each bus channel at the component interface.

[0041] In some cases, the system 100 transmits a ping packet to a peripheral device 106, such as a neural network processor, coupled to the host 102 at the component interface 110. In response to receiving the ping packet, the peripheral device can transmit a hardware definition file to the host for analysis at the host. In some implementations, the hardware definition file is only symmetric and asymmetric capabilities, such as the number of corresponding bus links at the device, each bus link configured for bidirectional data transfer, the maximum data bandwidth supported by each link or interconnection point at the peripheral device, the maximum data transfer rate of each link, or the maximum frequency supported by each link.

[0042] Figure 2 An example process 200 for performing asymmetric data communication is shown. The process 200 can be implemented or performed using the system 100 described above. Therefore, the description of the process 200 can refer to the computing resources of the system 100 mentioned above and other components described herein. In general, the computing steps or process flows included in the description of the process 200 can be grouped or arranged to occur in a different order and are not limited to the numerical order described herein.

[0043] Referring now to process 200, system 100 is configured to identify a plurality of devices coupled to at least a host of the system (202). In some implementations, the host management software control loop 130 is configured to monitor each interconnect location for establishing a connection with one or more of devices 106 and 108. The interconnect locations of the system may be represented by a socket or card slot integrated at a dedicated hardware circuit of the system. System 100 may reference a numbered list of identifiers for each interconnect location. Each interconnect location may correspond to a component interface for establishing a data connection between components of system 100, such as a processor 102 of the host and a hardware accelerator device 106.

[0044] The software loop can be operated to monitor each connection point at the component interface based on the position identifier of the interface. For example, the system 100 uses the software control loop 130 to monitor the connection activity at each component interface to identify or determine when a peripheral device is coupled to a host or another component at the system. The software control loop can analyze the numbered list of identifiers and the corresponding position of each identifier to determine the interconnection position of the peripheral device that establishes a data connection at the system. Using the software loop 130, the system 100 can be operated to determine a first interconnection position of an accelerator device 106 coupled to the host 102 and determine a second position of a peripheral device 108 coupled to the host.

[0045] The system 100 generates a system topology that identifies: i) the connectivity of multiple devices and ii) the bus channels that enable data transfer at the system (204). The host 102 can use an example BIOS or Linux command line to generate or execute a command, such as an "lspci" command, to identify the location of each peripheral device 106, 108 coupled to a connection point or component interface of the system 100. In some implementations, the host 102 generates a command that is processed by the operating system to display a detailed listing of information related to all data buses and devices in the system. For example, the listing can be based on a general portable interconnect library (e.g., libpci) that represents the interconnect configuration space of an operating system running on a processor of the host 102. The interconnect configuration space is referred to below in reference to Figure 3 Describe in more detail.

[0046] As discussed above, the system 100 can generate a ping packet that is provided to a peripheral device 106, such as a neural network processor. The device 106 can then transmit a hardware definition file to the host 102 in response to receiving the ping packet for analysis at the host. In some implementations, the hardware definition file is used to populate information in an interconnect library that describes the device configuration of an interconnect configuration space maintained by the operating system.

[0047] The device configuration indicates the symmetric and asymmetric capabilities of the buses and peripheral devices included at the system 100. Data and other information associated with the configuration space are used to generate a system topology. In some examples, the system topology identifies the connectivity of each device at the system 100 that has asymmetric data transmission capabilities, including the respective location identifiers of the devices. In other examples, the system topology also identifies the various data buses that enable asymmetric bidirectional data transmission at the system and the bus channels corresponding to each data bus.

[0048] The system 100 determines that a first connection between at least the host and a first device of the plurality of devices has an asymmetric bandwidth requirement (206). For example, connections between devices of the system 100 may have specific asymmetric bandwidth requirements based on predictions or inferences generated by analyzing data processing operations at the system 100. In some implementations, the software loop is operable to monitor data traffic at each component interface based at least on a location identifier of the interface.

[0049] The system 100 can use the software control loop 130 or the system topology to monitor the data traffic at each component interface to identify or determine the transmission mode at the system. For example, the system host interacts with the software control loop 130 and the software agent to determine the data transmission mode for a given set of data traffic using the system topology. The system host provides information describing the data traffic observed at the system 100 to the software agent managed by the software control loop 130. In some implementations, the software agent is represented by a data processing module including a trained machine learning model such as a machine learning engine or a statistical analysis engine. The data processing module is configured to analyze the information describing the data traffic at the system 100.

[0050] For example, a data processing module representing a software agent may include an artificial neural network, such as a deep neural network or a convolutional neural network, implemented on a GPU or a dedicated neural network processor. The neural network of the data processing module may be trained to generate a version of the software agent that is operable to perform inference calculations that produce data patterns for predicting asymmetric requirements at component interfaces of the system 100. For example, calculations for training the software agent may include processing inputs from training data through layers of a neural network to configure set weights to derive or predict asymmetries in observed data flows indicated by the data patterns. In some implementations, the software agent is based on an example support vector machine that uses a regression algorithm to perform statistical analysis to predict asymmetries in observed data flows indicated by one or more data flow patterns.

[0051] The information processed at the data processing module may include overall data bandwidth, data rate, data size, and the amount of data transmitted in an outgoing or incoming direction at each connection of the system, as well as the specific type of computing workload for which data is transmitted at the connection. The system 100 uses the software agent to determine the data transmission pattern at the system based on predictions or statistical analysis of the information. In some cases, the transmission pattern is obtained from a machine learning engine that processes the information using an example pattern mining algorithm. The machine learning engine derives or predicts the transmission pattern based on derivations calculated from the information. In some implementations, the software agent is used to generate a prediction of the distribution of data traffic that occurs when certain types of data analysis workloads are processed at the system 100.

[0052] In some implementations, the software control loop 130 causes a trained software agent (or machine learning model) to monitor and analyze data traffic at one or more component interfaces of the system 100. For example, the software agent analyzes data traffic observed at the component interface 110 for a given workload to determine various traffic flow patterns between the host 102 and the peripheral device 106. The workload may be an example classification and image recognition task, and the peripheral device 106 may represent a GPU hardware accelerator used to accelerate computations for the image recognition task.

[0053] The host 102 obtains a large set of image data from a networked hardware storage device represented by the peripheral device 108 via the component interface 112. The host 102 uses the set of incoming bus channels at the component interface 110 to provide the large set of image data to the GPU hardware accelerator (device 106). Similarly, the device 106 uses the set of outgoing bus channels at the component interface 110 to provide the output of its calculations, such as a text file describing objects recognized in multiple images. For this particular image recognition workload, the software agent analyzes the traffic flow patterns associated with the use of the host's incoming bus channels to provide image data to the hardware accelerator. Similarly, the software agent also analyzes the traffic flow patterns associated with the use of the accelerator's outgoing bus channels to provide recognition outputs to the host.

[0054] In some examples, the software agent analyzes inputs describing the data bandwidth of the connection, the data rate, the data size, and the relative amount of data transmitted in the outgoing or incoming direction at each bus channel or data link. For certain image recognition workloads, the software agent derives or predicts that the image data set provided via the incoming bus channel requires a signal bandwidth ranging from 250 gigabytes (GB) to 300GB, while the output text file provided via the outgoing bus channel requires a signal bandwidth ranging from 100 megabytes (MB) to 25GB. In some implementations, the system 100 has an example maximum signaling technology bandwidth of 100GB for each incoming bus channel (2x) at the component interface 100 and an example maximum signaling technology bandwidth of 100GB for each outgoing bus channel (2x) at the component interface 100. Therefore, the four bus channels can have a total signaling bandwidth of 400GB in total.

[0055] The system 100 uses a software control loop 130 to obtain predictions about data traffic patterns, including data transfer rates, incoming bandwidth requirements, outgoing bandwidth requirements, and the relative size of data routed via certain interfaces for a given workload. In particular, the control loop uses the software agent to calculate the asymmetric bandwidth requirements of the connection at the component interface 110. Based on the calculated bandwidth requirements, the software agent can operate to output a predicted ratio of incoming and outgoing bus channels that can most effectively cope with the predicted data traffic pattern. For example, the asymmetric bandwidth requirements of the component interface 110 can include a 3:1 ratio of incoming bus channels relative to outgoing bus channels. This ratio enables the incoming signaling bandwidth to be dynamically adjusted or increased to meet the example requirements of certain image recognition workloads that can range from 250GB to 300GB.

[0056] In other systems that do not implement the described techniques, the maximum signaling bandwidth in any one direction (e.g., ingress or egress) is limited to the symmetric signaling bandwidth or static data link configuration at the component interfaces of the system. For example, referring to the signaling bandwidth mentioned above, even if the actual data flow through the ingress path exceeds 200GB by a large margin, the ingress bandwidth in these other systems will be limited to 200GB based on their symmetric data links.

[0057] In contrast to these other systems, system 100 calculates an asymmetric bandwidth requirement for a connection based on a prediction indicating asymmetric data traffic at the connection. For example, the system calculates an asymmetric bandwidth requirement for a first connection corresponding to component interface 110 based on a data transmission pattern showing different types of asymmetric data traffic at interface 110. Similarly, the system may calculate an asymmetric bandwidth requirement for a second connection corresponding to component interface 112 based on a data transmission pattern showing different types of asymmetric data traffic at interface 112.

[0058] Reference again Figure 1 , the system 100 may have a dual socket (2S) configuration including at least component interfaces 110 and 112. In this configuration, the hardware accelerator 106 is coupled or connected to one socket (interface 110), and the networking device 108 is connected to the other socket (interface 112). Figure 1 As indicated in , the example software circuitry 130 is operable to manage and control operations involving symmetric and asymmetric data transfers at various data paths, devices, and component interfaces of the system 100 .

[0059] For example, at least one data path extends from the networking device 108 to CPU1 104 via interface 112 (e.g., a network card), to CPU0 via interface 114, and to the accelerator 106 via interface 110. In some examples, at least for component interface 110, the system host can use software control loop 130 and software agent to analyze data transfer patterns and generate predictions indicating that the bandwidth requirements of the host-accelerator interface 110 are completely or substantially asymmetric. For example, in each step, the accelerator 106 can be used to read a large amount of data from a distributed storage system to perform calculations using the data. The storage system can include memories 118, 122, and 126, or various combinations of these memory structures.

[0060] The accelerator 106 can read a large amount of data from the networking device 108 through the component interface 112, perform intensive calculations on the data, and output a small amount of data as a summary or result of the calculation. In these types of operations, the vast majority of data is transmitted across the system 100 in a particular direction, while a substantially smaller amount of data is transmitted across the system in the opposite direction. This difference in the amount of data transmitted indicates the asymmetry used to dynamically reconfigure the bus channel allocation.

[0061] Unlike a system bus configuration with equal bandwidth in both directions, the described techniques can be employed to exploit asymmetric bandwidth requirements so that more bus channels are configured for transmission in a specific direction where the majority of data is transmitted across the system. Exploiting the asymmetry and dynamic allocation of available channel bandwidth translates into improved bandwidth allocation and contributes to computing efficiency at system 100.

[0062] The system 100 can be operable to configure a first set of bus channels of a first data bus based on an asymmetric bandwidth requirement of a first connection (208). In some cases, the first connection defines a connection between the accelerator 106 and the host 102. The first set of bus channels of the first data bus is configured to allocate a different number of bus channels in the first set of bus channels for data egress from the host 102 than for data ingress to the host 102. For example, with respect to the host 102, each bus channel can be dynamically configured as a data ingress channel for receiving ingress data provided to the host 102 or a data egress channel for transmitting outgress data provided by the host 102.

[0063] Figure 3An example list 300 of components of system 100 used to perform asymmetric data communication is shown. For example, list 300 may correspond to a detailed list of information related to data buses and devices having connections at component interfaces of system 100. In some implementations, list 300 is displayed at an example command line in response to an operating system processing a command generated by host 102. The command may be processed to identify the location of each peripheral device (e.g., devices 106 and 108) coupled to the component interfaces of system 100. In some examples, list 300 represents an example interconnect configuration space stored in memory 122 and maintained by an operating system of host 102.

[0064] List 300 includes a first subset 302 of components associated with one or more peripheral devices in system 100, such as device 106 or device 108. List 300 also includes a second subset 304 of components that can each represent a corresponding peripheral device of system 100, such as a universal serial bus (USB) controller, a serial ATA (SATA) controller, an audio device, a memory controller, or a signal processing controller. In some implementations, each component in subset 302 can be a corresponding device, such as an accelerator device. The device can have a set of registers for storing performance data or information describing the results of data processing operations performed by the accelerator.

[0065] List 300 includes a subset 306 of the respective identifiers of each peripheral component or device coupled at system 100, such as processor 104, device 106, and each of device 108. Subset 306 may represent a portion of an enumerated list of identifiers of each interconnect location corresponding to a component interface where data flows asymmetrically between at least two devices of system 100. For example, "00:04.1" may represent a location identifier of a component interface that includes a connection between host 102 and network device 108. The connection may have a specific asymmetrical bandwidth requirement. Host 102 uses the asymmetrical bandwidth requirement to configure a data bus at the interface so that more bus lanes are allocated for data ingress from network device 108 to host 102 than for data egress from host 102 to network device 108.

[0066] Figure 4 Shown includes Figure 1 400 shows an example of information related to data traffic at a system. In some implementations, the MBx digital value indicates the traffic distribution of read operations and write operations for devices associated with a particular socket. Figure 4 The data depicted indicate the results of predictions related to bandwidth requirements for some future time frames, while in other examples, Figure 4The depicted data indicates a currently observed bandwidth pattern or a previously observed bandwidth pattern. In general, an example component interface of system 100 may correspond to socket 402 or socket 404. Each of sockets 402 and 404 may have a data connection supported by a data bus having a predefined number (or quantity) of bus channels, such as 20 bidirectional bus channels. For a given component interface, system 100 may configure or allocate a different number or quantity of bus channels for data egress relative to data ingress. The configured allocation may be based at least on the asymmetric bandwidth requirements of the component interface.

[0067] As described above, system 100 calculates an asymmetric bandwidth requirement at a connection based on a prediction indicating asymmetric data traffic at the connection. For example, each of sockets 402 and 404 may include a data connection having specific asymmetric bandwidth requirements 406 and 408, respectively. For socket 402, the software agent directly predicts that the asymmetric bandwidth requirement of the connection may include an M:N ratio of incoming bus channels relative to outgoing bus channels, where M has an integer value greater than an integer value of N. Similarly, for socket 404, the software agent directly predicts that the asymmetric bandwidth requirement of the connection may include an N:M ratio of outgoing bus channels relative to incoming bus channels, where N has an integer value greater than an integer value of M.

[0068] Figure 5 An example hardware connection 500 with asymmetric bidirectional bandwidth capability is shown. Connection 500 may be an example hardware connection based on a particular interconnect standard. The interconnect standard may specify configuration preferences for devices and data buses that support asymmetric data transfer at component interfaces of system 100.

[0069] Connection 500 may be an example data bus or a portion of a larger data bus that provides an interconnected data communication path for routing data between various components of system 100. Connection 500 includes individual bus channels that can each be configured as a bidirectional data transport medium at system 100. For example, connection 500 may have 20 individual bidirectional bus channels including a first subset of channels 502 (e.g., five channels) and a second subset of channels 504 (e.g., fifteen channels).

[0070] Connection 500 may also include a first interconnection location 506 (A) having a plurality of connection points or links for coupling to a first device (e.g., host 102) and a second interconnection location 508 (A) having a plurality of connection points or links for coupling to a second device (e.g., network device 108). Each of locations 506 and 508 may represent an example component interface. In some cases, each of locations 506 and 508 each represents a different but related component interface. As described above, each component interface of system 100 may be associated with a data bus and each data bus includes a plurality of bus channels. For example, each data bus may have a set of bus channels corresponding to a set of lines providing a medium for transporting data at the system.

[0071] Based on the described techniques, connection 500 has endpoints A and B including connection points for corresponding bus channels of example component interfaces. The bus channels of connection 500 can be configured to have a number of A→B bus channels, e.g., for data egress from A to B, that is not equal to a number of B→A bus channels, e.g., for data egress from A to B. In some implementations, the number of bus channels allocated for carrying data in a given direction corresponds to the magnitude, such as size, of the data traffic in the given direction. The unequal number of bus channels represents the asymmetry of the data traffic, which corresponds to the asymmetric bandwidth requirements of a particular connection at system 100.

[0072] In some implementations, the asymmetric configuration of the bus channels in connection 500 is achieved by reconfiguring less than half of the transceivers at one end of the data bus (connection point, A) as receivers and by reconfiguring the corresponding transceivers at the other end of the data bus (connection point, B) as transmitters. This reconfiguration can be done on a per-application basis. For example, in a GPU application, there may be more data going from host 102 to device 106, while in an FPGA application, there may be more data going from device 106 to host 102.

[0073] In some examples, the reconfiguration may occur at boot time, virtual machine migration time, or other related computing sessions of the system 100. The reconfiguration may also occur dynamically and in response to retraining one or more links of the connection. In this way, the described techniques may be used to improve the computational efficiency of bidirectional communications on a data bus, and thus improve the performance of certain applications, without requiring any increase in the original or fixed transmission rate of a given physical communication link.

[0074] Embodiments of the subject matter and functional operations described in this specification may be implemented in digital electronic circuits, computer software or firmware embodied in a tangible manner, computer hardware including the structures disclosed in this specification and their structural equivalents, or in a combination of one or more of them. Embodiments of the subject matter disclosed in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier to be executed by a data processing device or to control the operation of the data processing device.

[0075] Alternatively or in addition, the program instructions may be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to appropriate receiver apparatus for execution by a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0076] The processes and logic flows described in this specification can be implemented by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating (multiple) outputs. The processes and logic flows can also be performed by, and the apparatus can also be implemented as, a dedicated logic circuit such as an FPGA (field programmable gate array), an ASIC (application-specific integrated circuit), or a GPGPU (general purpose graphics processing unit).

[0077] A computer suitable for executing the included computer program may be based, for example, on a general or special purpose microprocessor or both, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential components of a computer are a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more large storage devices for storing data, or be operatively coupled thereto to receive data from it or to transmit data to it, or both, such as magnetic disks, magnetic optical disks, or optical disks. However, a computer is not required to have such a device.

[0078] Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM and flash memory devices; magnetic disks, such as internal hard disks or removable disks. The processor and memory can be supplemented by or integrated into a dedicated logic circuit.

[0079] Although the specification contains many specific implementation details, these should not be understood as limitations on any invention or the scope of protection that can be claimed, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described in the specification in the context of a single embodiment can also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment can also be implemented individually in multiple embodiments or in any appropriate sub-combination. In addition, although features may be described above as acting in a certain combination and even initially claimed as such, one or more features from the claimed combination can be separated from the combination in some cases and the claimed combination can be directed to a sub-combination or a variation of the sub-combination.

[0080] Similarly, although the operations are depicted in a particular order in the figures, this should not be understood as requiring such operations to be performed in the particular order shown or in a continuous order in order to achieve the desired result, or to perform all illustrated operations. In some cases, multitasking and parallel processing may be advantageous. In addition, the division of the various system modules and components described in the above-described embodiments should not be understood as requiring such division in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or be packaged as multiple software products.

[0081] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying drawings do not necessarily require a desired order or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.

Claims

1. A method for asymmetric data communication, comprising: identifying a plurality of devices coupled to a host of the system; generating a system topology that identifies connectivity of the plurality of devices and identifies bus channels that enable data transfer at the system; generating a prediction of a data pattern by analyzing data processing operations at the system according to the system topology, wherein generating the prediction of the data pattern comprises using a software control loop configured to monitor data traffic at each connection point between a device of the plurality of devices and the host according to the system topology, wherein a software agent analyzes observed data traffic at each connection point for a given workload to determine various data traffic patterns between the host and devices of the plurality of devices, and wherein the prediction is learned from analyzing the data traffic patterns of the system; determining that a first connection between the host and a first device of the plurality of devices has an asymmetric bandwidth requirement based on (i) the system topology and (ii) a prediction of a data traffic pattern generated from the system topology using the software control loop; and A first bus channel set of a first data bus connecting the first device and the host is configured based on an asymmetric bandwidth requirement of the first connection to allocate different numbers of bus channels in the first bus channel set for data egress from the host and data ingress to the host.

2. The method according to claim 1, further comprising: determining that a second connection between the host and a second device among the plurality of devices has an asymmetric bandwidth requirement; and A second bus channel set of a second data bus connecting the second device and the host is configured based on the asymmetric bandwidth requirement of the second connection to allocate different numbers of bus channels in the second bus channel set for data egress from the host and data ingress to the host.

3. The method according to claim 2, wherein: The prediction of the data mode includes prediction of the data transmission mode, further comprising: calculating an asymmetric bandwidth requirement for the first connection based on the prediction of the data transmission pattern; and An asymmetric bandwidth requirement for the second connection is calculated based on the prediction of the data transmission pattern.

4. The method according to claim 2, further comprising: providing information describing data traffic at the system to the software agent, the data traffic being monitored by the software control loop; determining, using the software agent, a data transmission pattern at the system based on a statistical analysis of the information or an inferential analysis of the information; generating, using the software agent, a prediction indicating a distribution of data traffic for processing one or more workloads at the system based on the data transfer pattern; and An asymmetric bandwidth requirement for the first connection is calculated based on a prediction indicative of a distribution of data traffic at the first connection.

5. The method according to claim 4, further comprising: An asymmetric bandwidth requirement for the second connection is calculated based on a prediction indicative of a distribution of data traffic at the second connection.

6. The method according to claim 2, wherein: Each bus channel in the first set of bus channels is dynamically configured as a data-in channel or a data-out channel; and Each bus channel in the second set of bus channels is dynamically configured as a data-in channel or a data-out channel.

7. The method according to claim 6, further comprising: Data is exchanged between the host and the first device using bus channels in the first set of bus channels allocated to data egress from the host and bus channels in the first set of bus channels allocated to data ingress to the host.

8. The method according to claim 1, wherein: The asymmetric bandwidth requirement of the first connection comprises an M:N ratio of incoming bus channels relative to outgoing bus channels; and M has an integer value greater than the integer value of N.

9. The method according to claim 2, wherein: The asymmetric bandwidth requirement of the second connection comprises an N:M ratio of outgoing bus channels relative to incoming bus channels; and N has an integer value greater than M's integer value.

10. The method of claim 2, wherein the system comprises a processor and an accelerator, and the method further comprises: configuring the processor as the host; identifying the accelerator as the first device; and It is determined that the accelerator is configured to have connectivity including a bus channel configured for bidirectional data transfer with the host via the first connection.

11. The method of claim 10, wherein the system comprises a memory, and the method further comprises: identifying the memory as the second device; and It is determined that the memory is configured to have connectivity including a bus channel configured for bidirectional data transfer with the host via the second connection.

12. A system for asymmetric data communication, comprising: one or more processors; as well as One or more non-transitory machine-readable storage media storing instructions executable by the one or more processors to cause operations to be performed, the operations comprising: identifying a plurality of devices coupled to a host of the system; generating a system topology that identifies connectivity of the plurality of devices and identifies bus channels that enable data transfer at the system; generating a prediction of a data pattern by analyzing data processing operations at the system according to the system topology, wherein generating the prediction of the data pattern comprises using a software control loop configured to monitor data traffic at each connection point between a device of the plurality of devices and the host according to the system topology, wherein a software agent analyzes observed data traffic at each connection point for a given workload to determine various data traffic patterns between the host and devices of the plurality of devices, and wherein the prediction is learned from analyzing the data traffic patterns of the system; determining that a first connection between the host and a first device of the plurality of devices has an asymmetric bandwidth requirement based on (i) the system topology and (ii) a prediction of a data traffic pattern generated from the system topology using the software control loop; and A first bus channel set of a first data bus connecting the first device and the host is configured based on an asymmetric bandwidth requirement of the first connection to allocate different numbers of bus channels in the first bus channel set for data egress from the host and data ingress to the host.

13. The system of claim 12, wherein the operations further comprise: determining that a second connection between the host and a second device among the plurality of devices has an asymmetric bandwidth requirement; and A second bus channel set of a second data bus connecting the second device and the host is configured based on the asymmetric bandwidth requirement of the second connection to allocate different numbers of bus channels in the second bus channel set for data egress from the host and data ingress to the host.

14. The system according to claim 13, wherein: The prediction of the data pattern includes prediction of a data transmission pattern, wherein the operations further include: calculating an asymmetric bandwidth requirement for the first connection based on the prediction of the data transmission pattern; and An asymmetric bandwidth requirement for the second connection is calculated based on the prediction of the data transmission pattern.

15. The system of claim 13, wherein: The data mode includes a data transmission mode, wherein the operations further include: providing information describing data traffic at the system to the software agent, the data traffic being monitored by the software control loop; determining, using the software agent, a data transmission pattern at the system based on a statistical analysis of the information or an inferential analysis of the information; generating, using the software agent, a prediction indicating a distribution of data traffic for processing one or more workloads at the system based on the data transfer pattern; and An asymmetric bandwidth requirement for the first connection is calculated based on a prediction indicative of a distribution of data traffic at the first connection.

16. The system of claim 15, wherein the operations further comprise: An asymmetric bandwidth requirement for the second connection is calculated based on a prediction indicative of a distribution of data traffic at the second connection.

17. The system of claim 13, wherein: Each bus channel in the first set of bus channels is dynamically configured as a data-in channel or a data-out channel; and Each bus channel in the second set of bus channels is dynamically configured as a data-in channel or a data-out channel.

18. The system of claim 17, wherein the operations further comprise: Data is exchanged between the host and the first device using bus channels in the first set of bus channels allocated to data egress from the host and bus channels in the first set of bus channels allocated to data ingress to the host.

19. The system of claim 12, wherein: The asymmetric bandwidth requirement of the first connection comprises an M:N ratio of incoming bus channels relative to outgoing bus channels; and M has an integer value greater than the integer value of N.

20. The system of claim 13, wherein: The asymmetric bandwidth requirement of the second connection comprises an N:M ratio of outgoing bus channels relative to incoming bus channels; and N has an integer value greater than M's integer value.

21. The system of claim 13, wherein the system comprises a processor and an accelerator, and the operations further comprise: configuring the processor as the host; identifying the accelerator as the first device; and It is determined that the accelerator is configured to have connectivity including a bus channel configured for bidirectional data transfer with the host via the first connection.

22. The system of claim 21, wherein the system comprises a memory, and the operations further comprise: identifying the memory as the second device; and It is determined that the memory is configured to have connectivity including a bus channel configured for bidirectional data transfer with the host via the second connection.

23. A non-transitory machine-readable storage medium storing instructions executable by one or more processors to cause operations to be performed, the operations comprising: identifying a plurality of devices coupled to a host of the system; generating a system topology that identifies connectivity of the plurality of devices and identifies bus channels that enable data transfer at the system; generating a prediction of a data pattern by analyzing data processing operations at the system according to the system topology, wherein generating the prediction of the data pattern comprises using a software control loop configured to monitor data traffic at each connection point between a device of the plurality of devices and the host according to the system topology, wherein a software agent analyzes observed data traffic at each connection point for a given workload to determine various data traffic patterns between the host and devices of the plurality of devices, and wherein the prediction is learned from analyzing the data traffic patterns of the system; determining that a first connection between the host and a first device of the plurality of devices has an asymmetric bandwidth requirement based on (i) the system topology and (ii) a prediction of a data traffic pattern generated from the system topology using the software control loop; and A first bus channel set of a first data bus connecting the first device and the host is configured based on an asymmetric bandwidth requirement of the first connection to allocate different numbers of bus channels in the first bus channel set for data egress from the host and data ingress to the host.

Citation Information

Patent Citations

  • Bandwidth configurable io connector

    CN103827841A