Distributed hardware tracing

The hardware tracing system addresses the complexity of analyzing distributed software performance by correlating and chronologically arranging hardware events across multiple processor components, enhancing debugging and performance analysis.

JP2025106560AActive Publication Date: 2025-07-15GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025067459
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-03-29
Filing Date
2025-04-16
Publication Date
2025-07-15
Estimated Expiration
2037-10-20

AI Technical Summary

Technical Problem

Effective performance analysis of distributed software running across multiple Central Processing Units (CPUs) or graphics processing units (GPUs) is complex due to the lack of efficient mechanisms for correlating and analyzing hardware events across these distributed components.

Method used

A hardware tracing system that monitors and stores hardware events across multiple processor components, generating a data structure to arrange these events chronologically for performance analysis, using dynamic triggers and unique trace identifiers to enhance correlation and debugging.

Benefits of technology

Enables efficient correlation and analysis of hardware events, improving computational efficiency and aiding in debugging and performance analysis of distributed software execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106560000001_ABST
    Figure 2025106560000001_ABST
Patent Text Reader

Abstract

To provide distributed hardware tracing.SOLUTION: A method performed by a hardware tracing system for capturing data describing hardware events, comprises the steps for: detecting a trigger that causes capture of data describing a hardware event, in a multicore processor including a plurality of cores; storing, in response to detection of the trigger, event data describing a trace activity at the multicore processor; and providing the stored event data to host software. The event data provided to the host software represents a timeline of operations, allowing the host software to be used for analyzing the performance of a program code.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is related to U.S. patent application Ser. No. 15 / 472, filed March 29, 2017, and entitled “Synchronous Hardware Event Collection.” No. 15 / 472,932, and Attorney Docket No. 16113-8129001. The entire disclosure of U.S. patent application Ser. No. 15 / 472,932 is expressly incorporated herein by reference in its entirety. [Background technology]

[0002] background This specification relates to analyzing the execution of program code.

[0003] Effective performance analysis of distributed software running within distributed hardware components can be a complex task. Distributed hardware components are comprised of two or more Central Processing Units (CPUs) (or graphics processors) that cooperate and interact to execute a larger software program or portion of program code. In some cases, the GPU may be a processor core of a graphics processing unit (GPU).

[0004] From a hardware perspective (e.g., in a CPU or GPU), there are broadly two types of information or functionality available for performance analysis: 1) hardware performance counters, and 2) hardware event traces. Summary of the Invention [Means for solving the problem]

[0005] overview Generally, one aspect of the subject matter described in this specification can be embodied in a computer-implemented method executed by one or more processors. The method includes monitoring the execution of program code by a first processor component, the first processor component being configured to execute at least a first portion of the program code, and the method further includes monitoring the execution of program code by a second processor component, the second processor component being configured to execute at least a second portion of the program code.

[0006] The method further includes storing, in at least one memory buffer, data that identifies one or more hardware events that occur across a processor unit that includes the first processor component and the second processor component by a computing system. Each hardware event represents at least one of a memory access operation of the program code, an issued instruction of the program code, or a data communication associated with an executed instruction of the program code. The data that identifies each of the one or more hardware events includes metadata that characterizes the hardware event and a hardware event timestamp. The method includes generating, by the computing system, a data structure that identifies the one or more hardware events, the data structure being configured to arrange the one or more hardware events in a chronological order of events associated with at least the first processor component and the second processor component.

[0007] The method further includes storing, in a memory bank of a host device, the generated data structure for use in analyzing the performance of program code executed by at least the first processor component or the second processor component.

[0008] These and other implementations are each optional and may include one or more of the following features. For example, in some implementations, the method further includes the computing system detecting a trigger function associated with a portion of program code executed by at least one of the first processor component or the second processor component, and in response to detecting the trigger function, the computing system initiating at least one trace event that causes data associated with one or more hardware events to be stored in at least one memory buffer.

[0009] In some implementations, the trigger function corresponds to at least one of a specific sequence step in the program code or a specific time parameter indicated by a global time clock used by the processor unit, and the step of initiating at least one trace event includes determining that a trace bit is set to a specific value, and at least one trace event is associated with a memory access operation that includes a plurality of intermediate operations occurring across processor units, and in response to determining that the trace bit is set to a specific value, data associated with the plurality of intermediate operations is stored in one or more memory buffers.

[0010] In some implementations, the step of storing data for identifying one or more hardware events includes storing a first subset of data for identifying the hardware events of the one or more hardware events in a first memory buffer of the first processor component. The storing step occurs in response to the first processor component executing a hardware trace instruction associated with at least a first portion of the program code.

[0011] In some implementation examples, the step of storing data for identifying one or more hardware events further includes storing, in a second memory buffer of a second processor component, a second subset of the data for identifying the hardware events of the one or more hardware events. The storing step occurs in response to the second processor component executing hardware trace instructions associated with at least a second portion of the program code.

[0012] In some implementation examples, the step of generating a data structure further includes the computing system comparing at least the hardware event timestamps of each event in a first subset of the data for identifying hardware events with at least the hardware event timestamps of each event in a second subset of the data for identifying hardware events, and the computing system providing, for presentation in the data structure, a correlated set of hardware events based at least in part on the comparison of each event in the first subset with each event in the second subset.

[0013] In some implementation examples, the generated data structure identifies at least one parameter indicative of a latency attribute of a particular hardware event, and the latency attribute indicates at least the duration of the particular hardware event. In some implementation examples, at least one processor of the computing system is a multi-core multi-node processor having one or more processor components, and the one or more hardware events correspond at least in part to data transfers occurring between a first processor component of at least a first node and a second processor component of a second node. having one or more processor components, and the one or more hardware events correspond at least in part to data transfers occurring between a first processor component of at least a first node and a second processor component of a second node.

[0014] In some implementation examples, the first processor component and the second processor component are one of a processor, a processor core, a memory access engine, or a hardware function of a computing system, one or more hardware events partially correspond to the movement of data packets between a source and a destination, and the metadata characterizing the hardware events corresponds to at least one of a source memory address, a destination memory address, a unique trace identification number, or a size parameter associated with a direct memory access (DMA) trace.

[0015] In some implementation examples, a specific trace ID number is associated with a plurality of hardware events occurring across processor units, the plurality of hardware events correspond to a specific memory access operation, the specific trace ID number is used to correlate one or more of the plurality of hardware events, and is used to determine a latency attribute of the memory access operation based on the correlation.

[0016] Other aspects of the subject matter described in this specification may be embodied in a distributed hardware tracing system that includes one or more processors including one or more processor cores and one or more machine-readable storage units for storing instructions. The instructions are executable by one or more processors to perform operations, the operations include monitoring the execution of program code by a first processor component, the first processor component is configured to execute at least a first portion of the program code, the operations further include monitoring the execution of program code by a second processor component, and the second processor component is configured to execute at least a second portion of the program code.

[0017] The method further includes storing, in at least one memory buffer, data that identifies one or more hardware events occurring across a processor unit that includes a first processor component and a second processor component. Each hardware event represents at least one of a memory access operation of program code, an issued instruction of program code, or a data communication associated with an executed instruction of program code. The data that identifies each of the one or more hardware events includes metadata that characterizes the hardware event and a hardware event timestamp. The method includes the computing system generating a data structure that arranges the one or more hardware events in a chronological order of events associated with at least the first processor component and the second processor component.

[0018] The method further includes storing the generated data structure in a memory bank of a host device for use in analyzing the performance of program code executed by at least the first processor component or the second processor component.

[0019] Other realizations of this and other aspects include corresponding systems, apparatuses, and computer programs configured to perform the actions of the method encoded on a computer storage device. One or more computer systems can be so configured by software, firmware, hardware, or combinations thereof installed in the system to cause the system to perform actions when operating. One or more computer programs can be so configured by having instructions that, when executed by a data processing device, cause the device to perform actions. One or more computer programs can be so configured by having instructions that, when executed by a data processing device, cause the device to perform actions.

[0020] The subject matter described in this specification can be realized in certain embodiments so as to achieve one or more of the following advantages. The hardware tracing system described enables efficient correlation of hardware events that occur during the execution of a distributed software program by a distributed processing unit that includes a multi-node multi-core processor. The hardware tracing system described further includes mechanisms that enable the collection and correlation of hardware events / trace data in a plurality of cross-node configurations.

[0021] The hardware tracing system increases computational efficiency by using dynamic triggers that are executed through hardware knobs / functions. Further, hardware events can be serially time-stamped using event descriptors such as unique trace identifiers, event timestamps, event source addresses, and event destination addresses. Such descriptors assist software programmers and processor design engineers in effectively debugging and analyzing software and hardware performance issues that can occur during source code execution.

[0022] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

[0024] Like reference numbers and names in the various drawings indicate like elements. **DETAILED DESCRIPTION**

[0025] Detailed Description The subject matter described in this specification generally relates to distributed hardware tracing. In particular, a computing system monitors the execution of program code executed by one or more processor cores. For example, a computing system can monitor the execution of program code executed by a first processor core and the execution of program code executed by at least a second processor core. The computing system stores data identifying one or more hardware events in a memory buffer. The stored data identifying the events corresponds to events that occur across a distributed processor unit including at least the first and second processor cores.

[0026] For each hardware event, the stored data includes metadata and an event timestamp that characterize the hardware event. The system generates a data structure identifying the hardware events. The data structure arranges the events in chronological order and associates the events with at least the first or second processor core. The system stores the data structure in a memory bank of a host device and uses the data structure to analyze the performance of program code executed by the first or second processor core.

[0027] FIG. 1 shows a block diagram of an exemplary computing system 100 for distributed hardware tracing. As used in this specification, distributed hardware system tracing corresponds to the storage of data that identifies events occurring within components and sub-components of an exemplary processor microchip. Further, as used herein, a distributed hardware system (or tracing system) corresponds to a set of processor microchips or processing units that cooperate to execute respective portions of software / program code configured for distributed execution among a set of processor microchips or distributed processing units.

[0028] System 100 may be a distributed processing system having one or more processors or processing units that execute a software program distributively, i.e., by executing different portions of the program code on different processing units of system 100. The processing units may include two or more processors, processor microchips, or processing units, such as at least a first processing unit and a second processing unit.

[0029] In some embodiments, two or more processing units may be distributed processing units when a first processing unit receives and executes a first portion of the program code of a distributed software program and a second processing unit receives and executes a second portion of the program code of the same distributed software program.

[0030] In some embodiments, different processor chips of system 100 may form respective nodes of the distributed hardware system. In an alternative embodiment, a single processor chip may include one or more processor cores and hardware functions that can each form respective nodes of that processor chip.

[0031] For example, in the context of a central processing unit (CPU), the processor chip may include at least two nodes, and each node may be a respective core of the CPU. Alternatively, in the context of a graphics processing unit (GPU), the processor chip may include at least two nodes, and each node may be a respective streaming multiprocessor of the GPU. Computing system 100 may include a plurality of processor components. In some implementations, the processor components may be at least one of the processor chips, processor cores, memory access engines, or at least one hardware component of the entire computing system 100.

[0032] In some examples, a processor component such as a processor core may be a fixed-function component configured to execute at least one specific operation based on at least one issued instruction of the program code being executed. In other examples, a processor component such as a memory access engine (MAE) may be configured to execute program code at a lower level of detail or granularity than the program code executed by other processor components of system 100.

[0033] For example, the program code executed by the processor core can cause an MAE descriptor to be generated and sent to the MAE. After receiving the descriptor, the MAE can execute a data transfer operation based on the MAE descriptor. In some implementations, the data transfer executed by the MAE may include, for example, moving data between components of system 100 via a certain data path or interface component of the system, or issuing a data request to an exemplary configuration bus of system 100.

[0034] In some implementation examples, each tensor node of the exemplary processor chip of system 100 may have at least two "front ends" that can be hardware blocks / functions for processing program instructions. As described in more detail below, the first front end can correspond to the first processor core 104, while the second front end can correspond to the second processor core 106. Thus, the first and second processor cores may also be described herein as the first front end 104 and the second front end 106, respectively.

[0035] As used in this specification, a trace chain may be a specific physical data communication bus on which trace entries can be placed for transmission to an exemplary chip manager within system 100. The received trace entry may be a data word / structure that includes a plurality of bytes and a plurality of binary values or binary numbers. For this reason, the descriptor "word" refers to a fixed-size piece of binary data that can be treated as one unit by the hardware devices of an exemplary processor core.

[0036] In some implementation examples, the processor chip of the distributed hardware tracing system is a multi-core (i.e., having multiple cores) processor that executes a part of the program code on each core of the chip. In some implementation examples, a part of the program code can correspond to vectorized calculations for the inference workload of an exemplary multi-layer neural network. On the other hand, in alternative implementation examples, a part of the program code can generally correspond to software modules associated with conventional programming languages.

[0037] Computing system 100 generally includes a node manager 102, a first processor core (FPC) 104, and a second processor core (second processor core: SPC) 106, a node fabric (NF) 110, a data router 112, and a host interface block (HIB) 114. In some implementations, system 100 may include a memory mux 108 configured to perform signal switching, multiplexing, and demultiplexing functions. System 100 further includes a tensor core 116 within which the FPC 104 is disposed. The tensor core 116 may be an exemplary computing device configured to perform vectorized computations on a multi-dimensional data array. The tensor core 116 may include a vector processing unit (VPU) 118, which may interact with a matrix unit (MXU) 120, a transpose unit (XU) 122, and a reduction and permutation unit (RPU) 124. In some implementations, computing system 100 may include one or more execution units of a conventional CPU or GPU, such as a load / store unit, an arithmetic logic unit (ALU), and a vector unit. The components of system 100 collectively include a large set of hardware performance counters and support hardware that facilitates completion of trace activity within the components. As described in more detail below, by each processor core of system 100 it may include one or more execution units of a conventional CPU or GPU.

[0038] The components of system 100 collectively include a large set of hardware performance counters and support hardware that facilitates completion of trace activity within the components. As described in more detail below, by each processor core of system 100 The program code to be executed may include an embedded trigger used to enable multiple performance counters simultaneously during code execution. Generally, the detected trigger causes trace data to be generated for one or more trace events. The trace data can correspond to incremental parameter counts stored in counters and analyzed to identify performance characteristics of the program code. Data for each trace event can be stored in an exemplary storage medium (e.g., a hardware buffer) and may include a timestamp generated in response to the detection of the trigger.

[0039] Furthermore, trace data can be generated for various events occurring within the hardware components of system 100. Exemplary events can include inter-node and cross-node communication operations such as direct memory access (DMA) operations and synchronization flag updates (each described in more detail below). In some implementations, system 100 can include a globally synchronized timestamp counter, generally referred to as a Global Time Counter (GTC). In other implementations, system 100 can include other types of global clocks such as a Lamport clock.

[0040] The GTC can be used for an accurate correlation between program code execution and the performance of software / program code executed in a distributed processing environment. Additionally, and somewhat related to the GTC, in some implementations, system 100 can include one or more trigger mechanisms used by a distributed software program to start and stop data tracing in a very coordinated manner in a distributed system.

[0041] In some implementation examples, host system 126 compiles program code that may include embedded operands. When an operand is detected, it triggers the capture and storage of trace data associated with a hardware event. In some implementation examples, host system 126 provides the compiled program code to one or more processor chips of system 100. In an alternative implementation example, the program code may be compiled by an exemplary external compiler (using the embedded trigger) and loaded onto one or more processor chips of system 100. In some examples, the compiler can set one or more trace bits (described below) associated with a certain trigger embedded in a portion of the software instructions. The compiled program code may be a distributed software program executed by one or more components of system 100.

[0042] Host system 126 may include a monitoring engine 128 configured to monitor the execution of program code by one or more components of system 100. In some implementation examples, monitoring engine 128 enables host system 126 to monitor the execution of program code executed by at least FPC 104 and SPC 106. For example, during code execution, host system 126 can monitor the performance of the execution code via monitoring engine 128 by receiving, at least, a periodic timeline of hardware events based on the generated trace data. Although a single block is shown for host system 126, in some implementation examples, system 126 may include multiple hosts (or host subsystems) associated with multiple processor chips or chip cores of system 100.

[0043] In another implementation example, when data traffic traverses the communication path between the FPC 104 and an exemplary third processor core / node, cross-node communication involving at least three processor cores may cause the host system 126 to monitor the data traffic with one or more intermediate "hops". For example, the FPC 104 and the third processor core may be the only cores executing program code during a given period. Thus, when data is transferred from the FPC 104 to the third processor core, the SPC 106 can generate trace data for the intermediate hop as the data is transferred from the FPC 104 to the third processor core. Stated another way, during data routing in the system 100, data traveling from the first processor chip to the third processor chip may need to cross the second processor chip, and thus, the execution of the data routing operation may cause trace entries to be generated for routing activity at the second chip.

[0044] When the compiled program code is executed, the components of the system 100 can interact to generate a timeline of hardware events that occur in a distributed computer system. Hardware events can include in-node and cross-node communication events. Exemplary nodes of the distributed hardware system and their associated communications are described in more detail below with reference to FIG. 2. In some implementation examples, a data structure is generated that identifies a set of hardware events for at least one hardware event timeline. The timeline enables the reconstruction of events that occur in the distributed system. In some implementation examples, event reconstruction can include correct event ordering based on the analysis of timestamps generated during the occurrence of specific events.

[0045] In general, an exemplary distributed hardware tracing system may include the above-described components of system 100 and at least one host controller associated with host system 126. The performance or debugging of data obtained from the distributed tracing system can be useful when event data is correlated, for example, in a time series or in an ordered manner. In some implementations, a plurality of stored hardware events corresponding to connected software modules are stored, and then data correlation can occur when they are ordered for structured analysis by host system 126. For implementations that include multiple host systems, the correlation of data obtained via different hosts may be performed, for example, by a host controller.

[0046] In some implementations, FPC 104 and SP 106 are each separate cores of one multi-core processor chip. On the other hand, in other implementations, FPC 104 and SP 106 are the respective cores of separate multi-core processor chips. As described above, system 100 may include a distributed processor unit having at least FPC 104 and SPC 106. In some implementations, the distributed processor unit of system 100 may include one or more hardware or software components configured to execute at least a portion of a larger distributed software program or program code.

[0047] Data router 112 is an inter-chip interconnect (ICI) that provides a data communication path between the components of system 100. In particular, router 112 can provide a communication link or connection between FPC 104 and SPC 106 and between the respective components associated with cores 104, 106. Node fabric 110 interacts with data router 112 to move data packets within the distributed hardware components and sub-components of system 100.

[0048] The node manager 102 is a high-level device that manages the low-level node functions in a multi-node processor chip. As described in more detail below, one or more nodes of the processor chip include a chip manager controlled by the node manager 102 to manage hardware event data and store it in a local entry log. It may include. The memory mux 108 is a multiplexing device that can perform switching, multiplexing, and demultiplexing operations on data signals provided to or received from an exemplary external high bandwidth memory (HBM). In some implementations, when the mux 108 switches between the FPC 104 and the SPC 106, exemplary trace entries (described below) can be generated by the mux 108. The memory mux 108 can affect the performance of certain processor cores 104, 106 that do not have access to the mux 108. For this reason, the trace entry data generated by the mux 108 can help understand the spikes that result in the latency of certain system activities associated with each core 104, 106. In some implementations, the hardware event data (e.g., trace points described below) that occurs within the mux 108 can be grouped with the event data for the node fabric 110 in an exemplary hardware event timeline. Event grouping can occur when a certain trace activity causes event data for multiple hardware components to be stored in an exemplary hardware buffer (e.g., the trace entry log 218 described below).

[0049]

[0050] In system 100, the performance analysis hardware includes FPC 104, SPC 106, mux 108, node fabric 110, data router 112, and HIB 114. Each of these hardware components or units includes a hardware performance counter and a hardware event trace mechanism and functionality. In some implementations, VPU 118, MXU 120, XU 122, and RPU 124 do not include their own dedicated performance hardware. Rather, in such implementations, FPC 104 can be configured to provide the necessary counters for VPU 118, MXU 120, XU 122, and RPU 124.

[0051] VPU 118 may include an internal design architecture that supports localized high-bandwidth data processing and arithmetic operations associated with the vector elements of an exemplary matrix-vector processor. MXU 120 is a matrix multiplication unit configured to perform, for example, matrix multiplications of up to 128×128 on a vector data set of the multiplicand.

[0052] XU 122 is a transpose unit configured to perform, for example, matrix transpose operations of up to 128×128 on vector data associated with matrix multiplication operations. RPU 124 may include a sigma unit and a replacement unit. The sigma unit performs sequential reduction on vector data associated with matrix multiplication operations. The reduction may include sums and various types of comparison operations. The replacement unit can completely replace or duplicate all elements of the vector data associated with matrix multiplication operations.

[0053] In some implementation examples, the program code executed by the components of system 100 may represent machine learning, neural network inference calculations, and / or one or more direct memory access functions. The components of system 100 may be configured to execute one or more software programs that include instructions to cause one or more functions to be executed on a processing unit or device of the system. The term "component" is intended to include any data processing device, or storage device such as a control status register, or any other device capable of processing and storing data.

[0054] System 100 generally includes a plurality of processing units or devices that may include one or more processors (e.g., microprocessors or central processing units (CPUs)), graphics processing units (GPUs), application specific integrated circuits (ASICs), or combinations of different processors In alternative embodiments, system 100 may each include other computing resources / devices (e.g., cloud-based servers) that provide additional processing options for performing calculations related to the hardware trace functions described herein.

[0055] The processing unit or device may further include one or more memory units or memory banks (e.g., registers / counters). In some implementations, the processing unit executes programmed instructions stored in memory for the system 100's devices to perform one or more of the functions described in this specification. The memory unit / bank may include one or more non-transitory machine-readable storage media. Non-transitory machine-readable storage media may include solid-state memory, magnetic disks, and optical disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (e.g., EPROM, EEPROM, or flash memory), or any other tangible medium capable of storing information.

[0056] FIG. 2 shows a block diagram of a trace chain and respective exemplary nodes 200, 201 used for distributed hardware tracing executed by system 100. In some implementations, nodes 200, 201 of system 100 may be different nodes within a single multi-core processor. In other implementations, node 200 may be the first node in a first multi-core processor chip, and node 201 may be the second node in a second multi-core processor chip.

[0057] Although two nodes are illustrated in the implementation example of FIG. 2, in alternative implementation examples, system 100 may include multiple nodes. For implementation examples with multiple nodes, cross-node data transfer can generate trace data at intermediate hops along an exemplary data path traversing the multiple nodes. For example, an intermediate hop can correspond to a data transfer passing through a separate node in a particular data transfer path. In some examples, trace data associated with ICI traces / hardware events can be generated for one or more intermediate hops that occur during cross-node data transfer passing through one or more nodes.

[0058] In some implementation examples, node 0 and node 1 are tensor nodes used for vectorized calculations associated with a part of the program code for the inference workload. As used in this specification, a tensor is a multi-dimensional geometric object, and exemplary multi-dimensional geometric objects include matrices and data arrays.

[0059] As shown in the implementation example of FIG. 2, node 200 includes a trace chain 203 that interacts with at least a subset of the components of system 100. Similarly, node 201 includes a trace chain 205 that interacts with at least a subset of the components of system 100. In some implementation examples, nodes 200, 201 are exemplary nodes of the same subset of components, while in other implementation examples, nodes 200, 201 are the respective nodes of separate subsets of components. Data router / ICI 112 includes a trace chain 207 that generally converges with trace chains 203 and 205 to provide trace data to chip manager 216.

[0060] In the implementation example of FIG. 2, nodes 200, 201 may each include a respective subset of components having at least FPC 104, SPC 106, node fabric 110, and HIB 114. Each component of nodes 200, 201 includes one or more trace muxes configured to group trace points (described below) generated by a particular component of the node. FPC 104 is a trace mux It includes 204, the node fabric 110 includes trace muxes 210a / b, the SPC 106 includes trace muxes 206a / b / c / d, the HIB 214 includes trace mux 214, and the ICI 212 includes trace mux 212. In some implementations, the trace control register for each trace mux enables individual trace points to be enabled and disabled. In some examples, for one or more trace muxes, their corresponding trace control registers may include individual enable bits and broader trace mux control.

[0061] Generally, the trace control register may be a conventional control status register (CSR) that receives and stores trace instruction data. Broader Regarding broader trace mux control, in some implementations, tracing can be enabled and disabled based on CSR writes executed by the system 100. In some implementations, tracing can be dynamically started and stopped by the system 100 based on the value of a global time counter (GTC), the value of an exemplary trace mark register in the FPC 104 (or core 116), or the value of a step mark in the SPC 106.

[0062] Details and explanations related to a computing system and a method implemented by a computer for dynamically starting and stopping trace activity and for synchronized hardware event collection are described in related U.S. Patent Application No. 15 / 472,932, titled "Synchronized Hardware Event Collection," filed on March 29, 2017, and Attorney Docket No. 16113 - 8129001. The entire disclosure of U.S. Patent Application No. 15 / 472,932 is hereby expressly incorporated by reference in its entirety.

[0063] In some implementation examples, for core 116, FPC 104 can use trace control parameters to define a trace window associated with event activities occurring within core 116. The trace control parameters enable the trace window to be defined by a lower limit and an upper limit for the GTC, as well as a lower limit and an upper limit for the trace mark register.

[0064] In some implementation examples, system 100 may include a function that enables a reduction in the number of generated trace entries, such as a trace event filtering function. For example, each of FPC 104 and SPC 106 may include a filtering function that restricts the rate at which each core sets trace bits in an exemplary generated trace descriptor (described below). HIB 114 may include a similar filtering function, such as an exemplary DMA rate limiter that restricts the trace bits associated with the capture of a certain DMA trace event. In addition, HIB 114 may include control (e.g., via an enable bit) for restricting the queuing of source DMA trace entries.

[0065] In some implementation examples, descriptors for DMA operations may have trace bits set by an exemplary compiler of host system 126. When the trace bits are set, hardware functions / knobs that determine and generate trace data are used to complete an exemplary trace event. In some examples, the last trace bit in DMA may be a logical OR operation of the trace bits statically inserted by the compiler and the trace bits dynamically determined by a specific hardware component. Thus, in some examples, the trace bits generated by the compiler can provide a mechanism for reducing the overall amount of generated trace data, apart from filtering.

[0066] For example, the compiler of host system 126 may perform one or more remote DMA operations (for For example, it may be determined to set only the trace bits for DMA spanning between at least two nodes and clear the trace bits for one or more local DMA operations (e.g., DMA within a specific tensor node such as node 200). In this way, the amount of trace data generated can be reduced based on trace activities limited to cross-node (i.e., remote) DMA operations rather than trace activities including both cross-node and local DMA operations.

[0067] In some implementations, at least one trace event initiated by system 100 may be associated with a memory access operation that includes a plurality of intermediate operations occurring across system 100. A descriptor for the memory access operation (e.g., an MAE descriptor) may include trace bits that cause data associated with the plurality of intermediate operations to be stored in one or more memory buffers. Thus, the trace bits can be used to "tag" intermediate memory operations at intermediate hops of the DMA operation as data packets traverse system 100 and generate a plurality of trace events.

[0068] In some implementations, ICI 112 may include a set of enable bits and a set of packet filters that provide control functionality for each of the ingress and egress ports of specific components of nodes 200, 201. These enable bits and packet filters enable ICI 112 to enable and disable trace points associated with specific components of nodes 200, 201. In addition to enabling and disabling trace points, ICI 112 may be configured to filter trace data based on event source, event destination, and trace event packet type.

[0069] In some implementations, in addition to using step markers, GTCs, or trace markers, each trace control register for processor cores 104, 106, and HIB 114 may also include an "everyone" trace mode. This "everyone" trace mode may enable tracing across the entire processor chip to be controlled by either trace mux 204 or trace mux 206a. In the everyone trace mode, trace muxes 204 and 206a can send a "window-in" trace control signal that identifies whether that particular trace mux, i.e., either mux 204 or mux 206a, is within the trace window.

[0070] The window-in trace control signal can be sent in parallel or broadcast to all other trace muxes within, for example, one processor chip or across multiple processor chips. If either mux 204 or mux 206a is performing trace activity, all tracing can be enabled by broadcasting to the other trace muxes. In some implementations, each trace mux associated with processor cores 104, 106, and HIB 114 includes a trace window control register that identifies when and / or how the "everyone trace" control signal is generated.

[0071] In some implementations, trace activity in trace muxes 210a / b and trace mux 212 is generally enabled based on whether a trace bit is set in a data word for a DMA operation or a control message crossing the ICI / data router 112. The DMA operation or control message may be a fixed-size binary data structure that may have a trace bit within a binary data packet set based on a certain situation or software state.

[0072] For example, if a DMA operation is a trace type DMA command to FPC 104 (or SP If it is started at C106) and the initiator (processor core 104 or 106) is within the trace window, the trace bit will be set in that particular DMA. In another example, for FPC104, if FPC104 is within the trace window and a trace point for storing trace data is enabled, a control message for writing data to another component within system 100 will set the trace bit.

[0073] In some implementations, zero-length DMA operations provide an example of a broader DMA implementation within system 100. For example, some DMA operations can generate non-DMA activity within system 100. The execution of non-DMA activity can also be traced (e.g., generate trace data) as if the non-DMA activity were a DMA operation (e.g., a DMA activity including non-zero-length operations). For example, a DMA operation that is started at a source location but has no data to be transmitted or transferred (e.g., is zero-length) may instead send a control message to the destination location. The control message will indicate that there is no data to be received or worked on at the destination. And the control message itself will be traced by system 100 such that non-zero-length DMA operations are traced.

[0074] In some examples, for SPC106, zero-length DMA operations can generate a control message, and the trace bit associated with that message will be set only if the DMA sets the trace bit, i.e., only if the control message does not have a zero length. Generally, if HIB114 is within the trace window, a DMA operation started from host system 126 will set the trace bit.

[0075] In the implementation example of FIG. 2, trace chain 203 receives trace entry data for a subset of components aligned with node 0, while trace chain 205 receives trace entry data for a subset of components aligned with node 1. Each of the trace chains 203, 205, 207 is a separate data communication path used by respective nodes 200, 201, and ICI112 to provide trace entry data to the exemplary trace entry data log 218 of chip manager 216. For this reason, the endpoints of trace chains 203, 205, 207 are chip manager 216 where trace events can be stored in an exemplary memory unit.

[0076] In some implementation examples, at least one memory unit of chip manager 216 can be 128 bits wide and have a memory depth of at least 20,000 trace entries. In alternative implementation examples, at least one memory unit may have a larger or smaller bit width and may have a memory depth that can store more or fewer entries.

[0077] In some implementation examples, chip manager 216 may include at least one processing device that executes instructions for managing the received trace entry data. For example, chip manager 216 can execute instructions to scan / analyze the timestamp data for each hardware event of the trace data received via trace chains 203, 205, 207. Based on the analysis, chip manager 216 can populate trace entry log 218 to include data that can be used to identify (or generate) the chronological order of hardware trace events. The hardware trace events can correspond to the movement of data packets that occur at the component and sub-component levels when the processing unit of system 100 executes an exemplary distributed software program. It can correspond to the movement of data packets that occur at the component and sub-component levels when the processing unit of system 100 executes an exemplary distributed software program.

[0078] In some implementation examples, the hardware units of system 100 may generate trace entries (and corresponding timestamps) that populate an exemplary hardware trace buffer out of order (i.e., in any order). For example, chip manager 216 can cause a plurality of trace entries with generated timestamps to be inserted into entry log 218. Among the plurality of inserted trace entries, each trace entry may not be serialized with respect to each other. In this implementation example, the out-of-order trace entries can be received by an exemplary host buffer of host system 126. When received by the host buffer, host system 126 can execute instructions related to performance analysis / monitoring software to scan / analyze the timestamp data for each trace entry. The executed instructions can be used to sort the trace entries and to construct / generate a timeline of hardware trace events.

[0079] In some implementation examples, during a tracing session, trace entries can be removed from entry log 218 via host DMA operations. In some examples, host system 126 may not remove DMA entries from trace entry log 218 as quickly as they are added to the log. In other implementation examples, entry log 218 may include a predefined memory depth. When the memory depth limit of entry log 218 is reached, additional trace entries may be lost. To control which trace entries are lost, entry log 218 can operate in a first-in-first-out (FIFO) mode or, alternatively, in an overwrite recording mode.

[0080] In some implementation examples, the overwrite recording mode may be used by the system 100 to support performance analysis associated with post-debugging. For example, the program code may be executed for a period of time with trace activity enabled and the overwrite recording mode enabled. In response to a post-software event (such as program corruption) within the system 100, the monitoring software executed by the host system 126 can analyze the data content of an exemplary hardware trace buffer to understand the hardware events that occurred prior to the program corruption. As used in this specification, post-debugging relates to the analysis or debugging of program code after the code has been corrupted or after it has generally ceased to execute / operate as intended.

[0081] In FIFO mode, if the entry log 218 is full and the host system 126 removes the stored log entries within a certain time frame, new trace entries may not be stored in the memory unit of the chip manager 216 to conserve memory resources. On the other hand, in overwrite recording mode, if the entry log 218 is full and the host system 126 removes the stored log entries within a certain time frame, new trace entries can be overwritten onto the oldest trace entry stored in the entry log 218 to conserve memory resources. In some implementation examples, trace entries are moved to the memory of the host system 126 in response to the DMA operation using the processing function of HIB114.

[0082] As used in this specification, a trace point is the origin of a trace entry and the data associated with that trace entry received by the chip manager 216 and stored in the trace entry log 218. In some implementation examples, a multi-core multi-node processor microchip includes three trace chains within the chip It may also be that the first trace chain receives a trace entry from chip node 0, the second trace chain receives a trace entry from chip node 1, and the third trace chain receives a trace entry from the ICI router of the chip.

[0083] Each trace point has a unique trace identification number within its trace chain that it inserts into the header of the trace entry. In some implementations, each trace entry identifies the trace chain in which it occurred in a header indicated by one or more byte / bit data words. For example, each trace entry may include a data structure having a defined field format (such as a header, payload, etc.) that conveys information about a particular trace event. Each field in the trace entry corresponds to useful data applicable to the trace point that generated the trace entry.

[0084] As described above, each trace entry may be written or stored in the memory unit of the chip manager 216 associated with the trace entry log 218. In some implementations, the trace points may be individually enabled or disabled, and multiple trace points may generate trace entries having different trace point identifiers but of the same type.

[0085] In some implementations, each trace entry type may include a trace name, a trace description, and a header that identifies an encoding for specific fields and / or a set of fields within the trace entry. These names, descriptions, and headers together provide a description of what the trace entry represents. From the perspective of the chip manager 216, this description can also identify that a particular trace entry belongs to a particular trace chain 203, 205, 207 within a particular processor chip. Thus, the fields within the trace entry represent pieces of data (e.g., in byte / bit units) regarding the description, and may be trace entry identifiers used to determine which trace point generated a particular trace entry.

[0086] In some implementations, the trace entry data associated with one or more of the stored hardware events can partially correspond to data communications that occur a) between at least node 0 and node 1, b) between components within at least node 0, and c) between components within at least node 1. For example, the stored hardware event can partially correspond to data communications that occur at least at one of 1) between the FPC 104 of node 0 and the FPC 104 of node 1, between the FPC 104 of node 0 and the SPC 106 of node 0, 2) between the SPC 106 of node 0 and the SPC 106 of node 1.

[0087] FIG. 3 shows a block diagram of an exemplary trace multiplexing design architecture 300 and an exemplary data structure 320. The trace multiplexing design 300 generally includes a trace bus input 302, a bus arbiter 304, and a local trace point arbiter 306, a bus FIFO 308, at least one local trace event queue 310, a shared trace event FIFO 312, and a trace bus output 314.

[0088] The multiplexed design 300 corresponds to an exemplary trace mux disposed within the components of the system 100. The multiplexed design 300 may include the following functionality. The bus in 302 may be associated with local trace point data that is temporarily stored in the bus FIFO 308 until the trace data is placed on an exemplary trace chain by time arbitration logic (e.g., arbiter 304). One or more trace points for the component can insert trace event data into at least one local trace event queue 310. The arbiter 306 provides a first level of arbitration and enables selection of events from the local trace events stored in the queue 310. The selected events are placed in a shared trace event FIFO 312 that also functions as a storage queue.

[0089] The arbiter 304 receives local trace events from the FIFO queue 312 and provides a second level of arbitration that merges the local trace events onto specific trace chains 203, 205, 207 via the trace bus out 314. In some implementations, trace entries may be pushed into the local queue 310 faster than they can be merged into the shared FIFO 312. Or, alternatively, trace entries may be pushed into the shared FIFO 312 faster than they can be merged onto the trace bus 314. If these scenarios occur, each queue 310 and 312 will become full with trace data.

[0090] In some implementations, when any of the queues 310 or 312 becomes full of trace data, the system 100 may be configured such that the latest trace entry is dropped and not stored or merged into a particular queue. In other implementations, rather than dropping trace entries when a queue (e.g., queue 310, 312) becomes full, the system 100 may be configured to stall an exemplary processing pipeline until the queue that has become full again has available queue space to receive an entry.

[0091] For example, a processing pipeline that uses queues 310, 312 may be stalled until a sufficient or threshold number of trace entries are merged on the trace bus 314. The sufficient or threshold number can correspond to a particular number of merged trace entries that results in available queue space for one or more trace entries to be received by queues 310, 312. Implementations where the processing pipeline is stalled until downstream queue space becomes available can provide higher-fidelity trace data based on the fact that trace entries are not dropped but preserved.

[0092] In some implementations, the local trace queue is as wide as required by the trace entry so that each trace entry occupies only one location in the local queue 310. However, the shared trace FIFO queue 312 can use a unique trace entry line encoding such that some trace entries can occupy two locations in the shared queue 312. In some implementations, if any data of a trace packet is dropped, the entire packet is dropped so that a partial packet does not appear in the trace entry log 218.

[0093] Generally, a trace is a timeline of activities or hardware events associated with a particular component of system 100. Unlike performance counters which are aggregate data (described below), a trace includes detailed event data that provides insight into the hardware activities that occur within a specified trace window. The hardware systems described enable extensive support for distributed hardware tracing, including generation of trace entries, temporary storage of trace entries in a hardware management buffer, static and dynamic enabling of one or more trace types, and streaming of trace entry data to host system 126.

[0094] In some implementations, a trace can be generated for hardware events that are executed by components of system 100, such as generation of DMA operations, execution of DMA operations, issue / execution of an instruction, or update of a synchronization flag. In some examples, trace activity can be used to track DMA through the system or to track instructions executed on a particular processor core.

[0095] System 100 can be configured to generate at least one data structure 320 that identifies one or more hardware events 322, 324 from a timeline of hardware events. In some implementations, data structure 320 arranges one or more hardware events 322, 324 in chronological order of events associated with at least FPC 104 and SPC 106. In some examples, system 100 can store data structure 320 in a memory bank of a host control device of host system 126. Data structure 320 can be used to evaluate the performance of program code executed by at least processor cores 104 and 106.

[0096] As shown by hardware event 324, in some implementations, a particular trace identification (ID) number (e.g., trace ID ‘003) can be associated with multiple hardware events that occur across distributed processor units. The multiple hardware events can correspond to a particular memory access operation (e.g., DMA), and the particular trace ID number is used to correlate one or more hardware events.

[0097] For example, as shown by event 324, a single trace ID for a DMA operation can include multiple timestamps corresponding to multiple different points in the DMA. In some examples, trace ID ‘003 can have “issued,” “executed,” and “completed” events that are identified as being somewhat temporally separated from each other. Thus, in this regard, the trace ID can further be used to determine the latency attributes of the memory access operation based on correlation and with reference to the timestamps.

[0098] In some implementations, generating data structure 320 can include, for example, the system 100 comparing the event timestamp of each event in a first subset of hardware events with the event timestamp of each event in a second subset of hardware events. Generating data structure 320 can further include the system 100 providing a correlated set of hardware events for presentation in the data structure, based in part on the comparison of the first subset of events and the second subset of events.

[0099] As shown in FIG. 3, data structure 320 can identify at least one parameter indicating the latency attributes of specific hardware events 322, 324. The latency attribute can at least indicate the duration of a specific hardware event. In some implementations, data structure 320 is generated by software instructions executed by a control device of host system 126. In some examples, structure 320 can be generated in response to the control device storing trace entry data in a memory disk / unit of host system 126.

[0100] FIG. 4 is a block diagram 400 showing exemplary trace activities for direct memory access (DMA) trace events executed by system 100. For DMA tracing, data about an exemplary DMA operation occurring from a first processor node to a second processor node can proceed through ICI 112 and can generate intermediate ICI / router hops along the data path. When a DMA operation crosses ICI 112, the DMA operation will generate trace entries at each node within the processor chip and at each hop. To reconstruct the time evolution of the DMA operation along the nodes and hops, information is captured by each of these generated trace entries. The exemplary DMA operation can be associated with the process steps shown in the implementation example of FIG. 4. For this operation, local DMA transfers data from virtual memory 402 (vmem402) associated with at least one of processor cores 104, 106 to HBM 108. The numbering shown in block diagram 400 corresponds to the steps in table 404 and generally represents activities in node fabric 110 or activities initiated by node fabric 110.

[0101]

[0102] ​The steps of Table 404 generally describe the associated trace points. An exemplary operation will generate six trace entries for this DMA. Step 1 includes the first DMA request from the processor core to the node fabric 110, which generates a trace point in the node fabric. Step 2 includes a read command for the node fabric 110 to request to transfer data to the processor core, which generates another trace point in the node fabric 110. When vmem402 completes the read of the node fabric 110, the exemplary operation does not have a trace entry for Step 3.

[0103] Step 4 includes the node fabric 110 performing a read resource update to cause a synchronization flag update in the processor core, which generates a trace point in the processor core. Step 5 includes a write command for the node fabric 110 to notify the memory mux 108 that the next data is being written to the HBM. The notification via the write command generates a trace point in the node fabric 110, while in Step 6, the completion of the write to the HBM also generates a trace point in the node fabric 110. In Step 7, the node fabric 110 performs a write resource update to cause a synchronization flag update in the processor core, which generates a trace point in the processor core (e.g., in FPC104). In addition to the write resource update, the node fabric 110 can perform a receive confirmation update (ack update) where the data completion for the DMA operation is signaled back to the processor core. The ack update can generate a trace entry similar to the trace entry generated by the write resource update.

[0104] In another exemplary DMA operation, when a DMA command is issued at the node fabric 110 of the source node, a first trace entry is generated. Additional trace entries may be generated at the node fabric 110 to capture the time used to read data for the DMA and write that data to the transmit queue. In some implementations, the node fabric 110 can packetize the DMA data into smaller chunks of data. For the data packetized into smaller chunks, read and write trace entries can be generated for the first and last data chunks. Optionally, in addition to the first and last data chunks, all data chunks can be set to generate trace entries.

[0105] For remote / non-local DMA operations that may require ICI hops, the first and last data chunks can generate additional trace entries at the ingress and egress points at each intermediate hop along the ICI / router 112. When the DMA data arrives at the destination node, trace entries similar to the previous node fabric 110 entries are generated at the destination node (e.g., read / write of the first and last data chunks). In some implementations, the last step of the DMA operation may include causing the executed command associated with the DMA to trigger an update of the synchronization flag at the destination node. When the synchronization flag is updated, a trace entry indicating the completion of the DMA operation can be generated.

[0106] ​In some implementation examples, to enable trace points to be executed, DMA tracing is initiated by FPC104, SPC106, or HIB114 when each component is in trace mode. Components of system 100 can enter trace mode based on global control in FPC104 or SPC106 via a trigger mechanism. In response to the occurrence of a specific action or state associated with the execution of program code by a component of system 100, a trace point is triggered. For example, a part of the program code may include an embedded trigger function that is detectable by at least one hardware component of system 100.

[0107] Components of system 100 can be configured to detect a trigger function associated with a part of the program code executed by at least one of FPC104 or SPC106. In some examples, the trigger function can correspond to at least one of 1) a specific sequence step in a part or module of the executed program code, or 2) a specific time parameter indicated by the GTC used by the distributed processor units of system 100.

[0108] In response to the detection of the trigger function, a specific component of system 100 can initiate, trigger, or execute at least one trace point (e.g., a trace event) that causes trace entry data associated with one or more hardware events to be stored in at least one memory buffer of a hardware component. As described above, the stored trace data can then be provided to chip manager 216 via at least one of trace chains 203, 205, 207.

[0109] FIG. 5 is a process flow diagram of an exemplary process 500 for distributed hardware tracing using the component functions of system 100 and one or more nodes 200, 201 of system 100. Thus, process 500 may be implemented using one or more of the above-described computing resources of system 100, including nodes 200, 201.

[0110] Process 500 begins at block 502 and includes the step of the computing system 100 monitoring the execution of program code executed by one or more processor components (including at least FPC 104 and SPC 106). In some implementations, the execution of program code that generates trace activity may be monitored at least in part by multiple host systems, or subsystems of a single host system. Thus, in these implementations, system 100 can perform multiple processes 500 related to the analysis of trace activity for hardware events occurring across distributed processing units.

[0111] In some implementations, a first processor component is configured to execute at least a first portion of the program code being monitored. At block 504, process 500 includes the step of the computing system 100 monitoring the execution of program code executed by a second processor component. In some implementations, the second processor component is configured to execute at least a second portion of the program code being monitored.

[0112] Each component of computing system 100 includes at least one memory It may include a buffer. Block 506 of process 500 includes the step of the system 100 storing data identifying one or more hardware events in at least one memory buffer of a particular component. In some implementations, the hardware events occur across a distributed processor unit including at least a first processor component and a second processor component. Each of the stored data identifying the hardware events may include metadata characterizing the hardware event and a hardware event timestamp. In some implementations, the set of hardware events corresponds to timeline events.

[0113] For example, the system 100 can store data identifying one or more hardware events that partially correspond to the movement of data packets between a source hardware component and a destination hardware component within the system 100. In some implementations, the stored metadata characterizing the hardware event can correspond to at least one of: 1) the source memory address, 2) the destination memory address, 3) a unique trace identification number associated with the trace entry that stores the hardware event, or 4) a size parameter associated with a direct memory access (DMA) trace entry.

[0114] In some implementations, the step of storing data identifying the set of hardware events includes, for example, storing event data in a memory buffer of the FPC 104 and / or the SPC 106 corresponding to at least one local trace event queue 310. The stored event data may indicate a subset of the hardware event data that can be used to generate a larger timeline of the hardware events. In some implementations, the storing of the event data occurs in response to at least one of the FPC 104 or the SPC 106 executing a hardware trace instruction associated with a portion of the program code executed by a component of the system 100.

[0115] In block 508 of process 500, system 100 generates a data structure, such as structure 320, that identifies one or more hardware events from a set of hardware events. The data structure arranges the one or more hardware events in a chronological order of events associated with at least a first processor component and a second processor component. In some implementations, the data structure identifies a hardware event timestamp for a particular trace event, a source address associated with that trace event, or a memory address associated with that trace event.

[0116] In block 510 of process 500, system 100 stores the generated data structure in a memory bank of a host device associated with host system 126. In some implementations, the stored data structure can be used by host system 126 to analyze the performance of program code executed by at least the first processor component or the second processor component. Similarly, the stored data structure can be used by host system 126 to analyze the performance of at least one component of system 100.

[0117] For example, a user or host system 126 can analyze the data structure to detect or determine whether there are performance issues associated with the execution of a particular software module within program code. Exemplary issues can include the software module not completing execution within an allotted execution time window.

[0118] Furthermore, the user or host device 126 can detect or determine whether a particular component of the system 100 is operating above or below a threshold performance level. Exemplary problems related to component performance can include a particular hardware component performing an event but generating result data outside an acceptable parameter range for the result data. In some implementations, this result data may not match the result data generated by other related components of the system 100 performing substantially the same operation.

[0119] For example, during the execution of program code, a first component of the system 100 may be required to complete an operation and to generate a result. Similarly, a second component of the system 100 may be required to complete a substantially similar operation and to generate a substantially similar result. Analysis of the generated data structure may reveal that the second component generated a result that is significantly different from the result generated by the first component. Similarly, the data structure may indicate result parameter values for the second component that are significantly outside the range of acceptable result parameters. These results may potentially indicate a performance problem with the second component of the system 100.

[0120] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more combinations of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a receiver apparatus suitable for execution by a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations of them.

[0121] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, and the apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array), an ASIC (application specific integrated circuit ), or a GPGPU (general purpose graphics processing unit).

[0122] Computers suitable for the execution of a computer program include, by way of example, general purpose or special purpose microprocessors or both, or any other kind of central processing unit, and may be based thereon. Generally, the central processing unit will receive instructions and data from read-only memory or random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions, and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from or transfer data to such mass storage devices, or both. However, a computer need not have such devices. Combined.

[0123] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks. Processors and memories may be supplemented or incorporated by special purpose logic circuitry.

[0124] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of the invention or the claims, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable partial combination. Furthermore, features may be described as acting in a certain combination and may thus be initially claimed as such, but in some cases, one or more features from the claimed combination may be deleted from that combination, and the claimed combination may be directed to a partial combination or a variation of a partial combination.

[0125] Similarly, operations are shown in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in the particular order or in a sequential order shown in order to achieve the desired result, or that all of the illustrated operations be performed. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0126] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired result. As one example, the processes illustrated in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

Claims

1. A computer-implemented method, executed by a computing system having one or more processors, the method comprising: monitoring the execution of program code by a first processor component, the first processor component being configured to execute at least a first portion of the program code, the method further comprising: monitoring the execution of the program code by a second processor component, the second processor component being configured to execute at least a second portion of the program code, the method further comprising: the computing system storing data identifying one or more hardware events occurring across processor units including the first processor component and the second processor component, each hardware event representing at least one of a memory access operation of the program code, an issued instruction of the program code, or a data communication associated with an executed instruction of the program code, the data identifying each of the one or more hardware events including metadata characterizing the hardware event and a hardware event timestamp, the method further comprising: the computing system generating a data structure identifying the one or more hardware events, the data structure being configured to arrange the one or more hardware events in a chronological order of events associated with at least the first processor component and the second processor component, the method further comprising: the computing system storing the generated data structure in a memory bank of a host device.

2. the computing system detecting a trigger function associated with a portion of program code executed by at least one of the first processor component or the second processor component; In response to the step of detecting the trigger function, the method according to claim 1, further comprising the step of starting at least one trace event that causes the computing system to store data associated with the one or more hardware events in at least one memory buffer.

3. The trigger function corresponds to at least one of a specific sequence step in the program code or a specific time parameter indicated by a global time clock used by the processor unit. The step of starting the at least one trace event includes the step of determining that a trace bit is set to a specific value. The at least one trace event is associated with a memory access operation including a plurality of intermediate operations occurring across the processor units. In response to the step of determining that the trace bit is set to the specific value, data associated with the plurality of intermediate operations is stored in one or more memory buffers. The method according to claim 2.

4. The step of storing data for identifying the one or more hardware events further includes the step of storing a first subset of data for identifying the hardware events of the one or more hardware events in a first memory buffer of the first processor component. The storing step occurs in response to the first processor component executing a hardware trace instruction associated with at least the first portion of the program code. The method according to claim 1.

5. The step of storing data for identifying the one or more hardware events further includes the step of storing a second subset of data for identifying the hardware events of the one or more hardware events in a second memory buffer of the second processor component. The storing step occurs in response to the second processor component executing a hardware trace instruction associated with at least the second portion of the program code. The method according to claim 4.

6. The step of generating the data structure further comprises The computing system compares at least the hardware event timestamps of each event in the first subset of data identifying hardware events with at least the hardware event timestamps of each event in the second subset of data identifying hardware events; The method according to claim 5, further comprising the computing system providing a correlated set of hardware events for presentation in the data structure, based at least in part on a comparison of each event in the first subset with each event in the second subset. **Claim 7** The method according to claim 1, wherein the generated data structure identifies at least one parameter indicative of a latency attribute of a particular hardware event, the latency attribute indicating at least a duration of the particular hardware event. **Claim 8** The method according to claim 1, wherein at least one processor of the computing system is a multi-core multi-node processor having one or more processor components, and the one or more hardware events correspond at least in part to data transfers occurring between the first processor component of the first node and the second processor component of the second node. **Claim 9** The first processor component and the second processor component are one of a processor, a processor core, a memory access engine, or a hardware function of the computing system, and the one or more hardware events correspond at least in part to the movement of data packets between a source and a destination. The method according to claim 1, wherein the metadata characterizing the hardware event corresponds to at least one of a source memory address, a destination memory address, a unique trace identification number, or a size parameter associated with a direct memory access (DMA) trace. **Claim 10** A distributed hardware tracing system, One or more processors including one or more processor cores; One or more machine-readable storage units for storing instructions, the instructions being executable by the one or more processors to perform operations, the operations being Including monitoring the execution of program code by a first processor component, the first processor component being configured to execute at least a first portion of the program code, the operations further being Including monitoring the execution of the program code by a second processor component, the second processor component being configured to execute at least a second portion of the program code, the operations further being Including the computing system storing data identifying one or more hardware events occurring across processor units including the first processor component and the second processor component, each hardware event Represents at least one of a memory access operation of the program code, an issued instruction of the program code, or data communication associated with an executed instruction of the program code, the data identifying each of the one or more hardware events including metadata characterizing the hardware event and a hardware event timestamp, the operations further being Including the computing system generating a data structure identifying the one or more hardware events, the data structure being configured to arrange the one or more hardware events in a chronological order of events associated with at least the first processor component and the second processor component, the operations further being A distributed hardware tracing system including the computing system storing the generated data structure in a memory bank of a host device.

11. The operations further are The computing system detecting a trigger function associated with a portion of program code executed by at least one of the first processor component or the second processor component, The distributed hardware tracing system of claim 10, comprising: in response to detecting the trigger function, starting, by the computing system, at least one trace event that causes data associated with the one or more hardware events to be stored in at least one memory buffer.

12. The trigger function corresponds to at least one of a specific sequence step in the program code or a specific time parameter indicated by a global time clock used by the processor unit. Starting the at least one trace event includes determining that a trace bit is set to a specific value. The at least one trace event is associated with a memory access operation that includes a plurality of intermediate operations occurring across the processor units. In response to determining that the trace bit is set to the specific value, data associated with the plurality of intermediate operations is stored in one or more memory buffers. The distributed hardware tracing system of claim 11.

13. Storing data identifying the one or more hardware events further includes storing, in a first memory buffer of the first processor component, a first subset of data identifying the hardware events of the one or more hardware events. Storing occurs in response to the first processor component executing a hardware trace instruction associated with at least the first portion of the program code. The distributed hardware tracing system of claim 10.

14. Storing data identifying the one or more hardware events further includes storing, in a second memory buffer of the second processor component, a second subset of data identifying the hardware events of the one or more hardware events. Storing occurs in response to the second processor component executing a hardware trace instruction associated with at least the second portion of the program code. The distributed hardware tracing system of claim 13.

15. Generating the data structure further comprises: the computing system comparing at least the hardware event timestamps of respective events in the first subset of data identifying hardware events with at least the hardware event timestamps of respective events in the second subset of data identifying hardware events; and the computing system providing a correlated set of hardware events for presentation in the data structure, based at least in part on a comparison of the respective events in the first subset with the respective events in the second subset. The distributed hardware tracing system according to claim 14 **Claim 16** The generated data structure identifies at least one parameter indicative of a latency attribute of a particular hardware event, the latency attribute indicating at least a duration of the particular hardware event. The distributed hardware tracing system according to claim 10 **Claim 17** The at least one processor is a multi-core multi-node processor having one or more processing components, and the one or more hardware events correspond at least in part to data communication occurring between the first processor component of the first node and the second processor component of the second node. The distributed hardware tracing system according to claim 10 **Claim 18** The first processor component and the second processor component are one of a processor, a processor core, a memory access engine, or a hardware function of the computing system, the one or more hardware events corresponding at least in part to the movement of data packets between a source and a destination, and metadata characterizing the hardware event corresponding to at least one of a source memory address, a destination memory address, a unique trace event identification (ID) number, or a size parameter associated with a direct memory access trace request. The distributed hardware tracing system according to claim 10 **Claim 19** ​ A specific trace ID number is associated with a plurality of hardware events occurring across said processor units, said plurality of hardware events corresponding to a specific memory access operation, said specific trace ID number being used to correlate one or more of said plurality of hardware events and being used to determine a latency attribute of said memory access operation based on the correlation, the distributed hardware tracing system according to claim 18.

20. A non-transitory computer storage unit disposed in a data processing device and encoded with a computer program, said program including instructions that, when executed by one or more processors, cause said one or more processors to perform operations, said operations including monitoring the execution of program code by a first processor component, said first processor component being configured to execute at least a first portion of said program code, said operations further including monitoring the execution of said program code by a second processor component, said second processor component being configured to execute at least a second portion of said program code, said operations further including the computing system storing data identifying one or more hardware events occurring across a processor unit including said first processor component and said second processor component, each hardware event representing at least one of a memory access operation of said program code, an issued instruction of said program code, or a data communication associated with an executed instruction of said program code, said data identifying each of said one or more hardware events including metadata characterizing said hardware event and a hardware event time stamp, said operations further including The computing system includes generating a data structure that identifies the one or more hardware events, the data structure being configured to arrange the one or more hardware events in a chronological order of events associated with at least the first processor component and the second processor component, and the operation further comprises The non-transitory computer storage unit includes the computing system storing the generated data structure in a memory bank of a host device.

Citation Information

Patent Citations

  • Computer system

    JP2008176477A

  • Process control device and system, and soundness determination method therefor

    JP2016106298A

  • Method and System for Monitoring and Debugging Access to a Bus Slave Using One or More Throughput Counters

    US20120226839A1