Method and apparatus for telemetry of system on chip

By designing computing systems for die arrays, microcontrollers and controllers in high-performance computing systems, the challenge of real-time monitoring and processing of telemetry data to determine die performance and implement correction measures is solved, real-time monitoring and management of die performance in computing systems is achieved, and system efficiency is improved.

CN120019363APending Publication Date: 2025-05-16TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069621.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-28
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In high-performance computing systems, especially in computing systems with large quantities of processing dies, there are challenges in real-time monitoring and processing of telemetry data to determine die performance and implement corrections.

Method used

A computing system is designed that includes a die array, a microcontroller and a controller. The microcontroller receives telemetry data and transmits it to the controller, which processes it to determine die performance and implements corrections such as deactivation or throttling a specific die based on the performance metrics satisfying thresholds.

Benefits of technology

Real-time monitoring and management of die performance in computing systems is realized, system efficiency is improved, and inefficient utilization of computing resources caused by performance degradation is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019363A_ABST
    Figure CN120019363A_ABST
Patent Text Reader

Abstract

The disclosure relates to systems and methods for monitoring computing resources of a computing system. The computing resource may include a plurality of systems on wafer (SoWs), each system on wafer including an array of die. The monitoring is based on receiving telemetry data from a die included in the SoW of the computing resource. An example computing system includes a SoW; a microcontroller communicatively coupled with the SoW and receiving telemetry data associated with at least one of the dies; and a controller configured to obtain the data from the microcontroller, determine a performance of a die of the SoW, and apply a corrective measure in response to determining a performance degradation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 378,013, filed on September 30, 2022, and entitled “Method and Apparatus for Telemetry Display of an On-Chip System,” the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to an apparatus for collecting telemetry data from a system on wafer (SOW) and processing the collected telemetry data. Background Art

[0004] Certain computing systems may be used and / or specifically configured for high performance computing and / or computing intensive applications, such as neural network training, neural network reasoning, machine learning, artificial intelligence, complex simulations, etc. In some applications, a computing system may be used to perform neural network training. For example, such neural network training may generate data for an autonomous driving system of a vehicle (e.g., a car), other autonomous vehicle functions, or advanced driver assistance system (ADAS) functions.

[0005] In high performance computing systems, there may be a high density of processing die. It may be desirable to obtain telemetry data associated with the processing die. In computing systems with a large number of processing die, there are technical challenges associated with processing telemetry data. Summary of the invention

[0006] The innovations described in the claims each have several aspects, no single one of which is solely responsible for its desirable attributes. Without limiting the scope of the claims, some of the prominent features of the disclosure will now be briefly described.

[0007] One aspect of the present disclosure is a computing system comprising: an array of dies, the array of dies being included on a system on a wafer (SoW); a microcontroller configured to receive telemetry data associated with at least one die of the array of dies; and a controller configured to obtain data including the telemetry data from the microcontroller, determine a performance metric of a particular die of the array of dies by processing the obtained data, and apply a corrective measure in response to determining that the performance metric satisfies a threshold. The die of the array is configured to output the telemetry data.

[0008] In computing systems, corrective action may include disabling specific dies.

[0009] In a computing system, corrective action may include throttling specific die.

[0010] In a computing system, a controller may be configured to identify a particular die that generated telemetry data.

[0011] In the computing system, the microcontroller may be configured to provide telemetry data to the controller at a first resolution in a first mode and at a second resolution in a second mode.

[0012] In a computing system, a microcontroller may be configured to communicate with two dies in an array of dies.

[0013] In a computing system, a controller may be configured to receive data from a plurality of SoWs.

[0014] In a computing system, telemetry data may include data associated with operating temperature, voltage, and current of at least one die.

[0015] In a computing system, the controller may also be configured to generate a graphical representation of the processed data.

[0016] In the computing system, the controller may also be configured to aggregate the telemetry data to perform post-processing on the aggregated data.

[0017] In a computing system, a controller may be configured to partition the dies of a SoW to perform parallel tasks.

[0018] Another aspect of the present disclosure is a method of monitoring a computing system. The method includes obtaining telemetry data from the computing system and determining a performance metric of an individual die of at least one SoW in a plurality of SoWs by processing the obtained telemetry data. The computing system includes a plurality of systems on wafers (SoWs). In addition, each SoW in the plurality of SoWs includes an array of dies.

[0019] In the method, the method may also include applying a corrective measure in response to determining that the performance metric of the particular die meets the threshold. The corrective measure may include deactivating the particular die. Furthermore, the corrective measure may include throttling the particular die.

[0020] In the method, the method may further include switching a mode of the microcontroller from a first mode to a second mode such that the microcontroller provides telemetry data associated with a particular die of a SoW in the plurality of SoWs at a different resolution in the second mode than in the first mode.

[0021] In the method, the telemetry data may include at least one of operating temperature, voltage, current, and power consumption of the individual dies.

[0022] In this method, the method may further include generating a graphical representation of the performance metrics of the individual dies.

[0023] Another aspect of the present disclosure is a non-transitory computer-readable storage medium. The storage medium includes instructions that, when executed by one or more processors, cause a method of monitoring a computing system to be performed. The method includes obtaining telemetry data from the computing system, and determining a performance metric of an individual die of at least one SoW in a plurality of SoWs by processing the obtained telemetry data. The computing system includes a plurality of systems on wafers (SoWs). In addition, each SoW in the plurality of SoWs includes an array of dies.

[0024] Another aspect of the present disclosure is a method of providing a visualization of performance metrics of dies of a system on wafer (SoW). The method includes obtaining telemetry data from dies of the SoW, wherein the SoW includes an array of dies, determining a performance metric of each of the dies of the SoW based on processing the telemetry data, and providing a graphical representation of the performance metric of each of the dies of the SoW based on the determination.

[0025] Another aspect of the present disclosure is a non-transitory computer-readable storage medium. The storage medium includes instructions that, when executed by one or more processors, cause the following method to be performed: obtaining telemetry data from a die of a SoW, determining a performance metric for each of the dies of the SoW based on processing the telemetry data, and providing a graphical representation of the performance metric for each of the dies of the SoW based on the determination. The SoW includes an array of dies.

[0026] Another aspect of the present disclosure is a system comprising an array of dies on a system on a wafer (SoW), each die of the array being configured to output telemetry data, and a microcontroller configured to receive telemetry data from at least two dies of the array of dies. The microcontroller is operable in at least a first mode and a second mode, such that the microcontroller outputs telemetry data at different resolutions in the first mode and the second mode.

[0027] In the system, the microcontroller may be configured to output the telemetry data along with information identifying the respective die of the array of die associated with the respective portions of the telemetry data.

[0028] To summarize the present disclosure, certain aspects, advantages, and novel features of the innovations are described herein. It should be understood that not all of these advantages may be achieved according to any particular embodiment. Thus, the innovations may be embodied or implemented in a manner that achieves or optimizes one advantage or group of advantages taught herein without necessarily achieving other advantages taught or suggested herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Embodiments of the present disclosure will be described by way of non-limiting examples with reference to the accompanying drawings.

[0030] Figure 1 An example of a cabinet including one or more processing systems according to embodiments disclosed herein is illustrated.

[0031] Figure 2 An example of a computing system according to embodiments disclosed herein is illustrated.

[0032] Figure 3A An example SoW of an array of included dies is illustrated.

[0033] Figure 3B An array of nodes included in each die is illustrated.

[0034] Figure 4 An example of the interaction between the controller and the SoW and the die is illustrated.

[0035] Figure 5 An example of partitioning the die of a SoW is illustrated.

[0036] Fig. 6A Illustrated is an example of a graphical representation of processed telemetry data at the SoW level according to embodiments disclosed herein.

[0037] Figure 6B Illustrated is an example of a graphical representation of processed telemetry data at the die level according to embodiments disclosed herein.

[0038] Figure 7 An example of a controller is illustrated.

[0039] Fig. 8A An example of interaction between a controller, a group of microcontrollers, a SoW, and a die according to embodiments disclosed herein is illustrated.

[0040] Figure 8B An example of interaction between a die and a microcontroller according to embodiments disclosed herein is illustrated.

[0041] Figure 8C An example of interaction between a microcontroller and two dies according to embodiments disclosed herein is illustrated.

[0042] Fig.8D An example of interaction between a microcontroller and a die according to embodiments disclosed herein is illustrated. DETAILED DESCRIPTION

[0043] The following detailed description of certain embodiments presents various descriptions of specific embodiments. However, the innovations described herein may be embodied in a variety of different ways, for example, as defined and covered by the claims. In this specification, reference is made to the accompanying drawings in which the same reference numerals and / or terms may indicate the same or functionally similar elements. It should be understood that the elements shown in the figures are not necessarily drawn to scale. In addition, it should be understood that certain embodiments may include more elements than shown in the drawings and / or a subset of the elements shown in the drawings. In addition, some embodiments may combine any suitable combination of features from two or more drawings.

[0044] The present disclosure relates to an apparatus for collecting and displaying data from a system on a wafer (SoW) in real time and for playback. Telemetry data sent by the SoW can be configured for debugging and / or operating status. A method may include collecting high-volume telemetry data and distributing it to an apparatus endpoint.

[0045] A SoW may include an array of integrated circuit dies packaged together. A SoW may achieve high computing density. A SoW may include an integrated cooling system. A system tray may include an array of SOWs supported by a common structure and interconnected. A system tray may be arranged in a computer cabinet. In a computing system, multiple adjacent computer cabinets may be interconnected.

[0046] As the demand for computing resources of computing systems increases, high-density computing systems are needed. For example, a computing system may include one or more SoWs, and each SoW may include an array of integrated circuit dies. However, monitoring the operation (or performance) and operating environment of each die can be challenging. For example, the die can generate heat during operation, and a portion of the SoW (e.g., a group of dies in the SoW) can have a relatively higher temperature than other portions. In another example, one or more dies in the SoW can have a relatively lower power supply than other dies in the SoW. Therefore, the performance of these dies with lower power supplies can be degraded.

[0047] Conventionally, it is not possible to detect in real time integrated circuit dies that have degraded performance due to the operation or operating environment of each die integrated into a SoW. If one or more dies of a SoW become inoperable or have degraded performance, the lack of real-time monitoring can cause the performance of the SoW to be degraded. Furthermore, if a portion of a SoW (e.g., a group of dies in a SoW) has degraded performance, this can cause the entire SoW to have degraded performance. This situation can cause inefficient utilization of computing resources within a computing system.

[0048] To address at least a portion of the above technical challenges, one or more aspects of the present disclosure correspond to a computing system including one or more SoWs and one or more controllers. Each SoW may be communicatively coupled to a controller and send data to the controller. The data may include telemetry data. The telemetry data may include information collected from each die or portion of the SoW (e.g., a group of dies within the SoW). The telemetry data may include environmental information, such as, but not limited to, one or more of the following: ambient and / or operating temperature of the die(s), operating parameters such as power supply for each die, current and / or voltage measurements of the die(s), and performance information such as usage, bandwidth, or latency of the die(s). Although the present disclosure describes telemetry data using specific types of data, these descriptions are provided only as examples, and the present disclosure does not limit the types of telemetry data.

[0049] Telemetry data may be sent from individual dies of the SoW to one or more microcontrollers on the control plane. As an example, a SoW may include 25 dies, and 13 microcontrollers may receive telemetry data from individual dies. In this example, 12 microcontrollers may receive telemetry data from two dies of the SoW, and 1 microcontroller may receive telemetry data from 1 die of the SoW. The microcontrollers may provide the telemetry data to the controller along with information identifying the individual dies associated with portions of the telemetry data. The controller may process and aggregate telemetry data from microcontrollers associated with one or more dies of the SoW. The controller may generate one or more graphical representations based on the processed telemetry data. Alternatively or additionally, the controller may direct one or more corrective measures based on the processed telemetry data.

[0050] Although embodiments disclosed herein may relate to computing systems with SoWs, any suitable principles and advantages disclosed herein may be applied to computing systems including multiple dies partitioned to perform computing tasks.

[0051] As described herein, the computing system may also be configured to collect telemetry data and provide an interface for visualizing the collected telemetry data and / or one or more performance metrics derived from the collected telemetry data. More specifically, the computing system may include a controller, and the controller may collect telemetry data. In some embodiments, the controller may collect telemetry data from one or more SoWs, multiple SoWs, individual dies of a SoW, and / or one or more portions of a die within a SoW. The controller may process the collected data and provide a graphical representation of the processed collected telemetry data. The graphical representation may provide various resolutions to describe the SoW as a whole and its individual dies or portions of a die. For example, the system's graphical interface may provide functionality that allows a user to examine the processed telemetry data at the SoW level and / or at the die level within each SoW.

[0052] As disclosed herein, a computing system may control operating parameters of each die, a portion of a SoW (e.g., a group of dies in a SoW), and / or each SoW based on the processing results of the collected telemetry data. For example, the computing system may apply one or more correction instructions based on the processing results of the collected telemetry data. For example, the system may control powering of the die(s) within the SoW based on the processing results of the collected telemetry data. As another example, the system may throttle one or more dies of the SoW based on the processing results of the collected telemetry data.

[0053] The principles and advantages disclosed herein can be applied to any suitable computing system. In certain applications, the disclosed system can be applied to a SoW, each of which includes an array of smaller die. The system can monitor the operation and operating environment of the die by receiving telemetry data from each die. The system can also provide a visual representation of the processed telemetry data at different resolutions, such as at each die level or at the SoW level. Therefore, a user or operator can identify the operation of each die and its environment in real time. In addition, once a problem in the operation or operating environment of a die is identified, the system can automatically implement one or more corrective measures. As a result, the disclosed system can improve efficiency within a computing system.

[0054] In some aspects of the present disclosure, multiple SoWs are used as computing resources of a computing system. For example, Figure 1 An example of two cabinets 100 that may include multiple computing tiles 110 is shown. Figure 1 As shown, the cabinet 100 may include one or more slots 102 or rails, and a system tray including an array of compute tiles 110 may be inserted into the cabinet 100 via the slots 102 . Figure 1Each cabinet 100 includes a cabinet structure arranged to receive two system trays, wherein each system tray includes an array of computing tiles 110 (e.g., a 2×3 array of computing tiles 110). Therefore, multiple computing tiles 110 can be integrated into the cabinet 100, and the present disclosure does not limit the type and number of slots, and these types and numbers can be determined based on specific applications.

[0055] In some embodiments, the computing tile 110 may include one or more of a control board structure, a cooling system, a voltage regulator module, a frame structure, a SoW, and a heat dissipation structure. An example of a computing tile 110 is disclosed in International PCT Application No. PCT / US2022 / 040420, entitled "Connector System and Related Method for Connecting Processor Systems", the disclosure of which is incorporated herein by reference in its entirety. In some applications, the SoW of the computing tile 110 includes a microcontroller that obtains telemetry data from the die of the SoW. In various applications, the computing tile 110 may include a circuit board including a plurality of microcontrollers thereon that obtain telemetry data from the die of the SoW of the computing tile 110. According to some other applications, a microcontroller external to the computing tile 110 on a circuit board located on a system tray may obtain telemetry data from the die of the SoW.

[0056] Multiple SoWs may be implemented in one cabinet, and the present disclosure provides systems and methods for monitoring the operation (and / or operating environment) of each SoW and its die, generating a graphical representation of the monitoring results, and controlling the operation of each SoW and its die based at least on the monitoring results.

[0057] Figure 2 An example computing system 200 is illustrated. Figure 2 As shown, the computing system 200 may include a SoW array 210. Each SoW 212 of the SoW array 210 may correspond to a computing tile 110 (e.g., Figure 1 ). The SoW array 210 may include SoWs 212 on one or more system trays and / or in one or more cabinets. The SoW array 210 may be used as a computing resource for the computing system 200. The number of SoWs 212 of the SoW array 210 may be determined based on a specific application.

[0058] like Figure 2 As shown, each SoW 212 may include an array of dies 232, and additionally Figure 2 Any suitable number of dies 232 may be included in SoW 212 .

[0059] In addition, Figure 2As shown, computing system 200 may include electronic module array 220. In some embodiments, the electronic module is a voltage regulator module (VRM) 222. Die 232 may be interconnected with electronic module array 220. In some embodiments, array of dies 232 and electronic module array 220 are stacked vertically. Figure 1 The computing tile 110 may include a SoW 212 stacked with an electronic module array 220 of VRMs 222. Each VRM 222 may be connected to and powered by its corresponding die 232. For example, die 1, 2, 3, and 4 may be connected to each corresponding VRM 1, 2, 3, and 4, respectively. Each VRM 222 may be configured to supply input power (e.g., input voltage) to the corresponding die 232. For example, each VRM 222 may be powered by an external power supply (e.g., an external power supply). Figure 2 ) receives a direct current (DC) power supply voltage and generates an output voltage to supply to the corresponding die 232.

[0060] Figure 3A An example of a SoW 212 is illustrated that includes an array of die 232. The die 232 may be an integrated circuit die. The die 232 may be implemented on a SoW 212 packaged with a wafer level packaging structure.

[0061] like Figure 3B As shown, in some embodiments, the tube core 232 may include an array of nodes. The array of nodes may include a computing node 306 and a global node 308. In some embodiments, the computing node 306 may include a circuit for performing a processing task. The global node 308 may generate telemetry data for the tube core 232. The global node 308 may not include a circuit for performing a processing task. For example, the global node 308 may include a pressure, voltage and temperature (PVT) sensor to monitor the operating conditions of the tube core 232. In some embodiments, the computing node 306 and the global node 308 may include a communication interface to enable communication with adjacent nodes. For example, each global node 308 can monitor the operating voltage of the surrounding nodes by receiving the current power supply voltage from the adjacent node via the communication interface. In some embodiments, the communication interface of the computing node 306 may be the same as the communication interface of the global node 308.

[0062] In some embodiments, each die 232 may generate telemetry data. Telemetry data may refer to information generated from each die. The information may include, but is not limited to, one or more environmental information, such as ambient and / or operating temperature of the die(s), operating parameters, such as the power supply of each die, current and / or voltage measurements of the die(s), and performance information, such as usage, bandwidth, and / or latency of the die(s). For example, each die 232 may be configured to communicate data with the SoW 212 via an input / output interface 312. In this example, the global node 308 may be configured to provide its data to the SoW 212 via the interface 312. In some scenarios, the global node 308 may continuously monitor the operating parameters of the computing node 306 by enabling a communication interface with an adjacent computing node 306. In addition, the global node 308 may continuously measure the operating temperature of the die 232.

[0063] Figure 4 4 illustrates an example of a computing system 400 according to one or more embodiments disclosed herein. Figure 4 As shown, computing system 400 may include SoW array 210 and controller 410. For example, Figure 3A and Figure 3B As described, each SoW 212 of the SoW array 210 may include an array 232 of dies, and each die may include an array 306 of computing nodes. The controller 410 may include a microcontroller or a computing device capable of performing a computing process, and the present disclosure does not limit the type of the controller 410. In addition, one or more controllers 410 may be implemented to perform one or more embodiments disclosed herein.

[0064] In some embodiments, the controller 410 may be configured to partition the SoW array 210 into one or more partitions. For example, the controller 410 may partition the SoW 212 of the SoW array 210 into two partitions, and each partition may be used as a computing resource for a specific computing task. The controller 410 may also partition the SoW array 210 based on the computing resources specified for each task. For example, if there are two tasks, Task A and B, where Task A involves more computing resources, the controller 410 may partition the SoW array 210 into two partitions, each containing a different number of SoWs 212. In this example, the partition with the higher number of SoWs 212 may be used to execute Task A.

[0065] Figure 5 An example of partitioning of the SOW array 210 is shown. Figure 5 As shown, the SoW 212 in the SoW array 210 is divided into two partitions, namely, partition 510 and partition 520. Figure 5, partition 520 may include more SoWs 212 than partition 510. Therefore, partition 520 may provide more computing resources than partition 510. In some embodiments, controller 410 may estimate the number of tasks and the computing resources required to perform each task. Then, controller 410 may partition SoW array 210 and assign partitions based on the computing resources required for each task. Figure 5 A specific number of SoWs 212 and two partitions are described, but these are provided as examples only, and the present disclosure does not limit the number of dies and partitions. In addition, the number of dies and partitions can be determined based on the specific application.

[0066] like Figure 4 As shown, the controller 410 can be configured to exchange data with the die 232 of the SoW 212. In some embodiments, each die 232 can send telemetry data to the controller 410. For example, the global node 308 (in Figure 3B 232) may collect telemetry data and transmit the telemetry data to controller 410. Controller 410 may be capable of receiving telemetry data from multiple die 232. In some cases, controller 410 may receive telemetry data from multiple die 232 simultaneously.

[0067] In a computing system having an SoW 210 array, a large amount of telemetry data may be generated. Processing and organizing such telemetry data presents technical challenges.

[0068] In some embodiments, the controller 410 can identify each die and store the telemetry data by linking the received data to the specific die 232 that transmitted the telemetry data. For example, the controller 410 can store a unique address associated with each die 232 and its SoW 212. In addition, in some applications, each die 232 can transmit its own identifier when sending telemetry data. This can enable the controller 410 to match the telemetry data with the corresponding die that generated the telemetry data. In some embodiments, each die 232 can periodically transmit its telemetry data. In these embodiments, the time period can be provided by the system operator of the computing system 400. In some other embodiments, the time period can be determined based on the clock frequency of the die.

[0069] Controller 410 can monitor and process the received telemetry data in real time. In some applications, controller 410 can filter the received data based on one or more metrics. For example, controller 410 can use a metric like temperature. In this case, controller 410 can process telemetry data related to the operating temperature of each die and filter out dies based on the temperature, such as excluding those dies operating above a specified threshold temperature.

[0070] In some embodiments, the controller 410 may process telemetry data received from each die at different resolutions. For example, the controller 410 may process telemetry data received from a first die at a higher resolution than telemetry data received from a second die. The controller 410 may process telemetry data at different resolutions in different modes.

[0071] In various embodiments, the controller 410 may apply one or more corrective measures based on the results of processing the telemetry data received from the die 232. The corrective measures may include, but are not limited to, performing throttling, disabling the die(s), restarting the die(s), controlling power, etc. In some embodiments, the controller 410 may utilize one or more mathematical operations when applying the corrective measures. For example, when the telemetry data is transmitted from a particular die, the controller 410 may add the current data and temperature data from the received telemetry data. The controller 410 may then determine one or more corrective measures, such as throttling and controlling the power supplied to the die.

[0072] In some scenarios, controller 410 may reconfigure the partitioning of the die based on the results of monitoring telemetry data received from die 232. For example, if controller 410 initially partitioned the die into two sections, where the first section contained more die than the second section, it may repartition the die 232 if it determines that one or more dies should be disabled due to overheating.

[0073] The controller 410 may be configured to communicate with the SoW 212. In some embodiments, the controller 410 may store telemetry data by classifying the data based on the identification of the SoW 212 and the die 232. For example, each die 232 may have an identifier indicating the SoW 212 to which the die belongs and the specific identification of the die within the SoW. The identifier may be composed of a specific protocol, such as the Internet Protocol, etc. The present disclosure does not limit the type of identification or its protocol. In some embodiments, the controller 410 may aggregate telemetry data from multiple dies 232 and / or multiple SoWs 212.

[0074] In various embodiments, the controller 410 may post-process the aggregated telemetry data for each die. In some embodiments, the controller 410 may determine the correlation between the telemetry data and the performance of each die. For example, by analyzing the aggregated data, the controller 410 may identify the threshold operating temperature of the die that will degrade computing performance. In these embodiments, the controller 410 may be configured to automatically apply one or more corrective measures in response to detecting that the operating temperature of the die meets a threshold. In another example, the controller 410 may establish a correlation between the operating time of the die and its operating temperature. In this scenario, the controller 410 may automatically apply one or more corrective measures at the specific operating time of the die corresponding to a predetermined threshold temperature.

[0075] In some embodiments, controller 410 may perform post-processing to determine the external operating environment of computing system 400. This may involve processing aggregated telemetry data related to one or more external factors, such as the power supply to each die, the cooling structure of the die, etc. This type of analysis can help optimize the environment in which computing system 400 is operated. For example, if computing system 400 is housed in a data center, Figure 1 In the cabinet 100, the post-processing results can provide information about the ambient temperature of the data center, the level of power supplied to the computing system, and other relevant factors.

[0076] The controller 410 may also include a graphics interface 412. The graphics interface 412 may generate a graphical representation of the performance metrics. In some embodiments, the graphics interface 412 may generate the graphical representation at two different resolutions, such as at the SoW level and at the die level.

[0077] Fig. 6A An example graphical representation of performance metrics of a SoW with die-level resolution is shown. Fig. 6A As shown, the temperature properties of each die of the SoW can be indicated by the level of the varying square bars. Fig. 6A The graphical representation can be used to debug and / or enhance the performance of the computing system.

[0078] Figure 6B Another graphical representation of performance metrics of a computing system's SoW with die-level resolution is shown. Figure 6B , the temperature properties of each die can be represented by a different pattern (e.g., pattern 610) such that the die depicted by the pattern has a higher operating temperature than the die without the pattern. Thus, a graphical representation of the performance metrics of the die of multiple SoWs can summarize a large amount of performance data of the die. Figure 6BEach graphical representation in may include data associated with a performance metric for each die of the six SoWs. In this example, each SoW includes 25 dies. The six SoWs may be included on one system tray. Thus, Figure 6B The graphical interface may indicate the performance metrics of each die of each SoW on the system tray.

[0079] like Figure 6B As shown, selection element 620 can be used to select from a plurality of performance metrics to be displayed on the graphical interface. As shown, temperature is the selected performance metric. Selection element 620 can be used to select another performance metric, such as voltage, and then the graphical interface can display the voltage data of each die.

[0080] Figure 7 An example of a controller 410 is illustrated. In some embodiments, the controller 410 may include at least a storage medium 710 and a main processor 720. The storage medium 710 is a non-transitory storage medium. In some applications, the storage medium 710 is a non-volatile storage medium. As disclosed herein, the storage medium 710 may contain various instructions to perform one or more embodiments. It may also store and aggregate telemetry data received from the die 232. The main processor 720 may be used to execute instructions stored in the storage medium 710. In addition, the main processor 720 may utilize its computing resources to process the received telemetry data and generate a graphical representation of the attributes of the processed telemetry data.

[0081] Fig. 8A An example of a computing system 800 according to one or more embodiments disclosed herein is illustrated. The computing system 800 may include a SoW array 210 and a microcontroller group 820. In some embodiments, each microcontroller 810 may be configured to receive and process telemetry data from one or more die 232. For example, Figure 8B An example of a microcontroller group 820 configuration is shown. Figure 8B As shown, SoW 212 includes 25 die 232, and microcontroller group 820 includes 13 microcontrollers 810. In this example, each of the 12 microcontrollers 820 can receive telemetry data from 2 die 232 (e.g., Figure 8C ), and the remaining microcontroller 820 can receive telemetry data from the remaining die 232 (as shown Fig.8D 100). This configuration is provided as an example only, and the present disclosure is not limited to this configuration. In some embodiments, the SoW may include a microcontroller and a die that provides telemetry data to the microcontroller. The microcontroller group 820 may be integrated into the computing tile 110 ( Figure 1 as shown) or integrated as an external device.

[0082] like Fig. 8A As shown, each SoW 212 of the SoW array 210 may include an array 232 of dies, and each die may include an array 306 of compute nodes, such as Figure 3A and Figure 3B shown.

[0083] In addition, Fig. 8A As shown, each microcontroller 810 can be configured to communicate data with the die 232. In some embodiments, each die 232 can send telemetry data to the microcontroller 810. For example, the global node 308 (in Figure 3B ) can generate telemetry data and transmit the telemetry data to the microcontroller 810. In some embodiments, the microcontroller 810 can receive telemetry data from one or more die 232, such as Figure 8C and Fig.8D For example, the microcontroller group 820 may include 13 microcontrollers 810. In this example, the 12 microcontrollers may receive telemetry data from the two dies 232, such that each microcontroller receives telemetry data from the two dies 232, as shown in FIG. Figure 8C Then, the remaining microcontroller 810 receives telemetry data from the remaining die 232, as shown. Fig.8D shown.

[0084] In some embodiments, each microcontroller 810 can request a specific type of telemetry data from one or more die 232. In certain applications, the microcontroller 810 can dynamically select telemetry data associated with one or more specific metrics to obtain from the die 232. The microcontroller 810 can also control the resolution of the telemetry data and / or the frequency at which the telemetry data is obtained. The microcontroller 810 can process the telemetry data received from the die 232 by one or more of transforming, filtering, discarding, applying mathematical operations, etc.

[0085] In some embodiments, the microcontroller 810 can identify the die associated with the telemetry data and store the telemetry data by associating the received data with the specific die 232 that transmitted the telemetry data. For example, the microcontroller 810 can store a unique address or identifier associated with each die 232. In addition, when sending telemetry data, each die 232 can transmit its own identifier. This can enable the microcontroller 810 to match the telemetry data with the corresponding die that generated the telemetry data. In some embodiments, each die 232 can periodically transmit its telemetry data. In some of these embodiments, the time period can be provided by the system operator of the computing system 800. Alternatively or additionally, the time period can be determined based on the clock frequency of the die.

[0086] In some embodiments, the microcontroller 810 can operate in multiple modes associated with different resolutions of telemetry data. This can customize the resolution of telemetry data. When more accurate telemetry data associated with a specific performance metric is desired, the microcontroller can obtain and / or process higher resolution telemetry data associated with the specific performance metric. When telemetry data associated with a large number of performance metrics is required, lower resolution telemetry data associated with these specific performance metrics can be obtained and / or processed by the microcontroller. The microcontroller 810 can operate in a first mode, where lower resolution telemetry data is associated with more performance metrics. The microcontroller 820 can operate in a second mode, where higher resolution telemetry data is associated with fewer performance metrics. Therefore, setting the mode of the microcontroller 810 can control the resolution of telemetry data. In some applications, one or more microcontrollers 810 can process telemetry data from different die at different resolutions. For example, the microcontroller 810 can process telemetry data received from the first die at a higher resolution than the telemetry data received from the second die.

[0087] In addition, Fig. 8A As shown, computing system 800 may also include controller 830. Controller 830 may be any suitable controller that communicates with microcontroller 810. In some cases, controller 830 may be in a cabinet (e.g., Figure 1 In some applications, the controller 820 may be outside the cabinet of the computing system. The controller 830 may exchange data with the microcontroller 810. In some embodiments, the controller 830 may be configured to partition the SoW array 210, such as Figure 5 For example, the controller 830 may divide the SoW array 210 into a plurality of partitions, and each partition may be used as a computing resource for a specific task.

[0088] In some embodiments, the controller 830 can receive telemetry data for each die from each microcontroller 810 and also process the received telemetry data. For example, the controller 830 can process the telemetry data received from the microcontroller 810 by one or more of transforming, filtering, discarding, applying mathematical operations, etc.

[0089] The controller 830 may apply one or more corrective measures when determining the performance degradation of the die based on the processing results of the received telemetry data. For example, the controller 830 may include a threshold corresponding to the performance metric of a specific die in the array of dies. In this example, if the telemetry data received from the die indicates that the performance metric meets the threshold, the performance degradation of the die can be determined. In response to determining that the performance metric associated with the specific die meets the threshold, a corrective measure may be applied. The corrective measures may include, but are not limited to, performing throttling, disabling the die (one or more), restarting the die (one or more), controlling the power supply (e.g., adjusting the parameters of the VRM), etc. In some embodiments, when applying the corrective measures, the controller 820 may utilize mathematical and / or logical operations. For example, when telemetry data is transmitted from a specific die, the controller 830 may add current data and temperature data from the received telemetry data. Then, the controller 830 may determine one or more corrective measures, such as throttling and / or reducing the power supplied to the die.

[0090] The controller 830 may also enhance the telemetry data with more information, such as identifying the partition of the SoW 312 that includes the die 232 associated with the received telemetry data. The telemetry data may be aggregated for the partitions, and then measures may be applied at the partition level. For example, one partition may be throttled to provide more power to one or more other partitions. The controller 830 may receive one or more signals from other hardware in the data center, such as substations and cooling infrastructure. This may enable advanced data center analytics and be associated with the performance of one or more SoWs.

[0091] Data associated with telemetry and / or performance may be stored by controller 830. Such data may be accessed later for various purposes, including but not limited to debugging and fault analysis.

[0092] Controller 830 may use telemetry data and other system information to implement power- and / or thermal-aware scheduling algorithms for computing resources that may enhance hardware utilization and / or establish safety monitoring and alteration control loops.

[0093] In some scenarios, controller 820 may reconfigure the partitioning of the die based on the results of monitoring telemetry data received from die 232. For example, if controller 820 initially partitioned the die into two partitions, where the first partition contained more die than the second partition, it may repartition the die 232 if it determines that one or more die are being deactivated due to overheating.

[0094] The controller 820 may also be configured to communicate with the SoW 212. In some embodiments, the controller 820 may store telemetry data by sorting the data based on identification of specific die 232 associated with portions of the telemetry data. For example, each die 232 may have an index, and the microcontroller 810 may have an address. In this example, the controller 820 may identify a specific die 232 associated with specific telemetry data based on the index of the die 232 and the address of the associated microcontroller 810. The identification may be performed in a specific protocol such as an Internet protocol. The present disclosure does not limit the type of identification and its protocol. In some embodiments, the controller 820 may aggregate the telemetry data.

[0095] In various embodiments, the controller 820 may post-process the aggregated telemetry data for each die. In some embodiments, the controller 820 may determine the correlation between the telemetry data and the performance of each die. For example, by analyzing the aggregated data, the controller 820 may identify the threshold operating temperature of the die that may degrade computing performance. In these embodiments, the controller 820 may be configured to automatically apply one or more corrective measures when it is detected that the operating temperature of the die is at or near a threshold temperature. In another example, the controller 820 may establish a correlation between the operating time of the die and its operating temperature. In this scenario, the controller 820 may automatically apply corrective measures at the specific operating time of the die corresponding to a predetermined threshold temperature.

[0096] In some embodiments, controller 820 may perform post-processing to determine the external operating environment of computing system 800. This may involve processing aggregated telemetry data related to one or more external factors, such as, but not limited to, the power supply to each die, the cooling structure of the die, etc. This type of analysis may help optimize the environment in which computing system 800 is operated. For example, if computing system 800 is housed in a rack 100 (e.g., a rack 100 in a data center) Figure 1 As shown in the figure), the post-processing results can provide information about the ambient temperature of the data center, the power level supplied to the computing system, and other relevant factors.

[0097] like Fig. 8A As shown, the controller 820 may also include a graphics interface 412. The interface 412 may be configured to generate a graphical representation of the telemetry data attributes. In some embodiments, the graphics interface 412 may generate graphical representations at two different resolutions, such as at the SoW level and at the die level.

[0098] The computing systems disclosed herein can be implemented in various processing systems. Such processing systems can be used and / or specifically configured for high-performance computing and / or computationally intensive applications, such as neural network training, neural network reasoning, machine learning, artificial intelligence, complex simulations, etc. In some applications, the processing system can be used to perform neural network training. For example, such neural network training can generate data for an autonomous driving system of a vehicle (e.g., a car), other autonomous vehicle functions, or advanced driver assistance system (ADAS) functions.

[0099] Unless the context clearly requires otherwise, throughout the specification and claims, the words "comprise", "comprising", "include", "including", etc. should be understood as inclusive, as opposed to exclusive or exhaustive; that is, in the sense of "including, but not limited to". The word "coupled" as generally used herein refers to two or more elements that can be directly connected or connected through one or more intermediate elements. Similarly, the word "connected" as generally used herein means that two or more elements can be directly connected or connected through one or more intermediate elements. In addition, when used in this application, the words "herein", "above", "below" and words of similar meaning shall refer to the entirety of this application and not to any particular part of this application. Where the context permits, words used in the above detailed description in the singular or plural may also include the plural or singular, respectively. The word "or" when referring to a list of two or more items covers all of the following interpretations of the word: any item in the list, all items in the list, and any combination of items in the list.

[0100] In addition, conditional language used herein, such as "can", "could", "might", "may", "eg", "for example", "such as", etc., unless otherwise specifically stated or otherwise understood in the context of use, is generally intended to convey that some embodiments include, while other embodiments do not include, certain features, elements, and / or states. Therefore, such conditional language is generally not intended to imply that one or more embodiments require features, elements, and / or states in any way.

[0101] The foregoing description has been described with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the disclosure to the precise forms described. In view of the above teachings, many modifications and variations are possible. Therefore, others skilled in the art will be able to best utilize these techniques and various embodiments with various modifications to suit a variety of uses.

[0102] Although the present disclosure and examples have been described with reference to the accompanying drawings, various changes and modifications will become apparent to those skilled in the art. These changes and modifications should be understood to be included within the scope of the present disclosure.

Claims

1. A computing system, comprising: an array of dies, the array of dies being included on a system on a wafer (SoW), wherein the dies of the array are configured to output telemetry data; a microcontroller configured to receive telemetry data associated with at least one die of the array of dies; as well as A controller is configured to obtain data including the telemetry data from the microcontroller, determine a performance metric for a particular die of the array of dies by processing the obtained data, and apply a corrective action in response to determining that the performance metric satisfies a threshold.

2. The system of claim 1, wherein the corrective action comprises disabling the particular die.

3. The system of claim 1, wherein the corrective action comprises throttling the particular die. 4 . The system of claim 1 , wherein the controller is configured to identify the specific die that generated the telemetry data.

5. The system of claim 1, wherein the microcontroller is configured to provide the telemetry data to the controller at a first resolution in a first mode and a second resolution in a second mode.

6. The system of claim 1, wherein the microcontroller is configured to communicate with two dies in the array of dies.

7. The system of claim 1, wherein the controller is configured to receive data from a plurality of SoWs.

8. The system of claim 1, wherein the telemetry data includes data associated with operating temperature, voltage, and current of the at least one die.

9. The system of claim 1, wherein the controller is further configured to generate a graphical representation of the processed data.

10. The system of claim 1, wherein the controller is further configured to aggregate the telemetry data for post-processing the aggregated data.

11. The system of claim 1, wherein the controller is configured to partition the die of the SoW to perform parallel tasks.

12. A method of monitoring a computing system, the method comprising: obtaining telemetry data from the computing system, wherein the computing system comprises a plurality of systems on wafers (SoWs), and wherein each of the plurality of SoWs comprises an array of dies; as well as A performance metric of an individual die of at least one SoW of the plurality of SoWs is determined by processing the obtained telemetry data.

13. The method of claim 12, further comprising applying corrective action in response to determining that the performance metric of a particular die meets a threshold. The method of claim 13 , wherein the corrective action comprises disabling the particular die.

15. The method of claim 13, wherein the corrective action comprises throttling the particular die.

16. The method of claim 12, further comprising switching a mode of a microcontroller from a first mode to a second mode such that the microcontroller provides telemetry data associated with a particular die of a SoW in the plurality of SoWs at a different resolution in the second mode than in the first mode.

17. The method of claim 12, wherein the telemetry data comprises at least one of operating temperature, voltage, current, and power consumption of the individual die.

18. The method of claim 12, further comprising generating a graphical representation of the performance metric of the individual die.

19. A non-transitory computer-readable storage medium comprising instructions that, when executed by one or more processors, cause the method of claim 12 to be performed.

20. A method of providing visualization of performance metrics of a die of a system on wafer (SoW), the method comprising: obtaining telemetry data from a die of the SoW, wherein the SoW comprises an array of die; determining the performance metric for each of the dies of the SoW based on processing the telemetry data; as well as Based on the determined performance metric, a graphical representation of the performance metric for each of the dies of the SoW is provided.

21. A non-transitory computer-readable storage medium comprising instructions which, when executed by one or more processors, cause the method of claim 20 to be performed.

22. A system, comprising: an array of die on a system-on-wafer (SoW), each die of the array being configured to output telemetry data; as well as A microcontroller is configured to receive telemetry data from at least two dies of the array of dies, the microcontroller being operable in at least a first mode and a second mode such that the microcontroller outputs the telemetry data at different resolutions in the first mode and the second mode.

23. The system of claim 22, wherein the microcontroller is configured to output the telemetry data together with information identifying a corresponding die of the array of dies associated with a portion of the telemetry data.