Apparatus and method
By controlling the processor core to execute workloads and using machine learning to infer physical layouts, optimizing the spatial layout of the processor core, solving the problem of processor manufacturers not being publicly laid out and improving the performance and life of the processor.
Patent Information
- Application Number
- CN202510124650.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-08
- Filing Date
- 2025-01-26
- Publication Date
- 2025-08-08
AI Technical Summary
Processor manufacturers often do not disclose physical layout details on the processor die, making it difficult to optimize the spatial layout of the processor core to improve performance and reduce thermal choke.
By controlling the processor core to execute workloads and record temperatures, the physical layout of the processor core is inferred using machine learning and cluster analysis, thereby optimizing load distribution to uniform temperature distribution, reducing cooling system energy consumption and extending processor life.
Achieve higher processor performance and energy efficiency, reduce heat-blocking risks, and extend the life of the processor.
Smart Images

Figure CN120449802A_ABST
Abstract
Description
Background Art
[0001] Processor manufacturers typically keep secret the details of the spatial layout on a processor die, particularly how multiple processor cores are arranged. For example, the physical configuration of the cores and other circuit elements can significantly affect performance, power efficiency, and thermal distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Some examples of apparatus and / or methods will be described below by way of example only and with reference to the accompanying drawings, in which:
[0003] Figure 1A A block diagram illustrating an example of an apparatus;
[0004] Figure 1B A block diagram illustrating an example of an apparatus;
[0005] Figure 2A A flowchart illustrating an example of a method;
[0006] Figure 2B A flowchart illustrating an example of a method;
[0007] Figure 3A graphically illustrating temperature distribution over a physical layout of a first processing circuitry when the first core is stressed;
[0008] Figure 3B Figure 1 illustrates temperature distribution over the physical layout of the first processing circuitry with the second core stressed;
[0009] Figure 4 shows a temperature profile of a stressed processor over time and a temperature profile of an idle processor core over time;
[0010] Figure 5 The graph shows the linear regression between the collected temperature data of the stressed core and the temperature data of the idle core;
[0011] Figure 6 The graph shows the cluster analysis performed on the obtained regression coefficients for one core;
[0012] Figure 7 illustrates embodiments of standard configurations of cores belonging to different clusters or dies in a 1D, 2D or 3D layout; and
[0013] Figure 8 The diagram illustrates three possible placements for any given core within the physical layout of the first processing circuitry. DETAILED DESCRIPTION
[0014] Some examples are now described in more detail with reference to the accompanying drawings. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of features and equivalents and substitutes of features. In addition, the terms used in this article to describe certain examples should not limit other possible examples.
[0015] Throughout the description of the drawings, the same or similar reference numerals refer to the same or similar elements and / or features, which may be the same or implemented in a modified form while providing the same or similar functions. For clarity, the thickness of lines, layers and / or regions in the drawings may also be exaggerated.
[0016] When two elements A and B are combined using "or", it is understood that this discloses all possible combinations, i.e., only A, only B, and A and B, unless otherwise explicitly defined in individual cases. As alternative wording for the same combination, "at least one of A and B" or "A and / or B" can be used. This applies equally to combinations of more than two elements.
[0017] If singular forms such as "a / an" and "the" are used and the use of only a single element is neither explicitly nor implicitly defined as mandatory, further examples may also use several elements to achieve the same function. If a function is described below as being achieved using multiple elements, further examples may use a single element or a single processing entity to achieve the same function. It is further understood that when the terms "include," "including," "comprise," and / or "comprising" are used, they describe the presence of specified features, integers, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0018] In the following description, specific details are set forth, but examples of the technology described herein can be implemented without these specific details. Well-known circuits, structures, and techniques are not shown in detail to avoid obscuring the understanding of this description. "An example / example," "various examples / examples," "some examples / examples," etc. may include features, structures, or characteristics, but not every example necessarily includes these specific features, structures, or characteristics.
[0019] Some examples may have some, all, or none of the features described for other examples. "First," "second," "third," and the like describe common elements and indicate different instances of the same element being referenced. Such adjectives do not imply that the elements so described must be in a given order, either temporally or spatially, in ranking, or in any other manner. "Connected" may indicate that elements are in direct physical or electrical contact with each other, and "coupled" may indicate that elements cooperate or interact with each other, but these elements may or may not be in direct physical or electrical contact.
[0020] As used herein, the terms "operate," "execute," or "run" are used interchangeably when referring to software or firmware in connection with a system, device, platform, or resource, and may refer to software or firmware stored in one or more computer-readable storage media accessible to the system, device, platform, or resource, even if the instructions contained in the software or firmware are not being actively executed by the system, device, platform, or resource.
[0021] The specification may use the phrases "in an example / example," "in various examples," "in some examples / examples," and / or "in various examples / examples," each of which may refer to one or more of the same or different examples. Furthermore, the terms "comprising," "including," "having," and the like as used with respect to examples of the present disclosure are synonymous.
[0022] As described above, processor manufacturers may not publish the physical layout of on-die processing circuitry that includes multiple processor cores. In some examples, core identifiers (IDs) (e.g., central processing unit identifiers (CPUIDs)) may not provide information about the physical layout of the processor cores to the operating system (OS) (e.g., a core with ID "0" and a core with ID "1" may not be adjacent to each other). However, knowledge of the physical core layout from the operating system can be used for various purposes.
[0023] The proposed concept can execute a power-intensive workload on a single (processor) core while all other cores are idle. Then, according to the proposed concept, the temperatures of all (or some) cores can be recorded. It can then be inferred which cores are close to the core that generates the (primary) heat. By repeating this process for each core in the processing circuitry, the core layout of the die can be inferred. For example, machine learning techniques can be used in this regard.
[0024] Based on the inferred physical layout of the processing circuitry ("core layout"), this can be used for various purposes. For example, the OS can choose to distribute workloads to cores that are physically farther apart. By spreading the temperature more evenly across the die, the processor cooling system may require less energy (improved energy efficiency) and the cores may be less likely to encounter thermal throttling (higher performance). In addition, thermal degradation of parts will be reduced (longer life). In yet another example, surrounding cores can be used to execute a software-only thermal shmoo map (a thermal shmoo map can be a graphical representation used to characterize and analyze the performance and stability of a processor under various temperature conditions).
[0025] Figure 1 illustrates a block diagram of an example of an apparatus 100 or device 100. The apparatus 100 includes circuitry configured to provide the functionality of the apparatus 100. For example, Figure 1A Apparatus 100 includes interface circuitry 120, processing circuitry 130, and (optionally) storage circuitry 140. For example, processing circuitry 130 may be coupled to interface circuitry 120 and, optionally, to storage circuitry 140.
[0026] For example, processing circuitry 130 may be configured to provide the functionality of apparatus 100 in conjunction with interface circuitry 120. Interface circuitry 120 is configured to exchange information, for example, with other components internal or external to apparatus 100 and storage circuitry 140. Likewise, apparatus 100 may include apparatus configured to provide the functionality of apparatus 100.
[0027] The components of the apparatus 100 are defined as component means that may correspond to or be implemented by corresponding structural components of the apparatus 100. For example, Figure 1ADevice 100 includes: means for processing 130, which may correspond to or be implemented by processing circuitry 130; means for communicating 120, which may correspond to or be implemented by interface circuitry 120; and (optionally) means for storing information 140, which may correspond to or be implemented by storage circuitry 140. Hereinafter, the functionality of device 100 will be illustrated with respect to device 100. Therefore, features described in conjunction with device 100 may also apply to the corresponding device 100.
[0028] Generally speaking, the functionality of processing circuitry 130 or means for processing 130 may be implemented by executing machine-readable instructions by processing circuitry 130 or means for processing 130. Thus, any features attributed to processing circuitry 130 or means for processing 130 may be defined by one or more of the plurality of machine-readable instructions. Apparatus 100 or device 100 may include machine-readable instructions (e.g., within storage circuitry 140 or means for storing information 140).
[0029] The interface circuitry 120 or means for communicating 120 may correspond to one or more inputs and / or outputs for receiving and / or transmitting information within a module, between modules, or between modules of different entities, which may be in the form of digital (bit) values according to a specified code. For example, the interface circuitry 120 or means for communicating 120 may include circuitry configured to receive and / or transmit information.
[0030] For example, the processing circuit system 130 or the means for processing 130 may be implemented using one or more processing units, one or more processing devices, or any means for processing, such as a processor, a computer, or a programmable hardware component operable with correspondingly adapted software. In other words, the functionality of the processing circuit system 130 or the means for processing 130 described may also be implemented using software, which is then executed on one or more programmable hardware components. Such hardware components may include general-purpose processors, digital signal processors (DSPs), microcontrollers, and the like.
[0031] For example, the storage circuit system 140 or the device for storing information 140 may include at least one element of the group of computer-readable storage media, such as magnetic or optical storage media, such as a hard drive, a flash memory, a floppy disk, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a network storage device.
[0032] Processing circuitry 130 is configured to control each of the plurality of processor cores of the first processing circuitry to execute a corresponding workload. Furthermore, processing circuitry 130 is configured to obtain temperature measurement data from each of the plurality of processor cores. The temperature measurement data is obtained while the corresponding processor core of the plurality of processor cores executes the corresponding workload. For example, each of the processor cores of the first processing circuitry is controlled to execute the corresponding workload sequentially.
[0033] The workload executed by a processor core may consist of specific computing tasks and processes that the processor core is responsible for handling, such as running applications, performing calculations, or processing data. The execution of the workload may generate heat by the processor core due to the electrical energy used by the electrical components (such as transistors) of the processor core during operation. Therefore, the more intensive the workload executed by the stressed processor core, the greater the energy consumption and the more heat that may be generated. For example, each of the corresponding workloads executed by each of the multiple processor cores may be the same. In another example, the corresponding workload may be different for each of the multiple processor cores.
[0034] In some examples, the temperature measurement data includes one or more individual temperature measurements of one or more processor cores in the plurality of processor cores of the first processing circuitry, the one or more individual temperature measurements being measured by one or more dedicated sensors assigned to the one or more processor cores. For example, the individual temperature measurements of the processor cores may be expressed in degrees Celsius, degrees Fahrenheit, or Kelvin, etc. In some examples, the temperature measurement data may be acquired at predetermined time intervals. For example, the individual temperature measurements of the processor cores in the plurality of processor cores may be measured by the dedicated sensors every 5 ms, 20 ms, 100 ms, 200 ms, etc.
[0035] In some examples, for each processor core in the plurality of processor cores, the obtained temperature measurement data includes a corresponding temperature measurement data set corresponding to the corresponding processor core. The temperature measurement data set corresponding to a particular processor core may include temperature measurement data of the plurality of processor cores obtained when the particular processor core executes a corresponding workload. In other words, the temperature measurement data may include as many temperature measurement data sets as the number of processors, since there is one temperature measurement data set for each processor in the plurality of processors (see also Figure 3A and Figure 3B ). Each of the temperature measurement data sets may include temperature measurements for each of the plurality of processor cores obtained when the particular processor core (to which the temperature measurement data set corresponds) executes a corresponding workload.
[0036] In some examples, the processing circuit system 130 can be configured to control each processor core of the plurality of processor cores to execute a corresponding workload for a predetermined time. For example, the predetermined time can be 15 seconds or 20 seconds or 30 seconds or 40 seconds or 1 minute, etc. For example, after the predetermined time period ends, the processing circuit system 130 causes the corresponding processor core to stop executing its corresponding workload and continues to control the next processor core to execute its corresponding workload until each processor core has executed its corresponding workload. For example, the processing circuit system 130 can be configured to start controlling the first processor core of the plurality of processor cores to execute its corresponding workload for a predetermined time, and then continue with the second processor core to execute its corresponding workload for a predetermined time until the last processor core has executed its corresponding workload for a predetermined time. For example, during each execution phase of the corresponding processor core, the individual temperatures of all processor cores are measured and recorded (see also below). Figure 4 ).
[0037] In some examples, the processing circuit system 130 is configured to control each of the multiple processor cores to execute a corresponding workload until the change in temperature of the corresponding temperature measurement data obtained for the processor core executing the workload is below a predetermined threshold. For example, the change in temperature is determined as a change in the temperature of the processor core executing the workload over time. For example, if the change in the temperature of the processor core executing the workload is less than 0.5° or 1° or 1.5°, etc., the execution of the workload by the corresponding processor core is stopped. For example, after the change in temperature drops below a predetermined threshold, the processing circuit system 130 causes the corresponding processor core to stop executing its corresponding workload and continues to control the next processor core to execute its corresponding workload until each processor core has executed its corresponding workload. For example, during each execution phase of the corresponding processor core, the individual temperatures of all processor cores are measured and recorded (see also below). Figure 4 ).
[0038] Furthermore, processing circuitry 130 is configured to infer a physical layout of the first processing circuitry based on temperature measurement data obtained from each of the plurality of processor cores of the first processing circuitry. In some examples, the physical layout includes spatial positioning of the processor cores within the first processing circuitry. For example, the physical layout of the first processing circuitry includes a primary geometric configuration (a line, a square, a cube, etc.) and spatial coordinates defining the positions of the processor cores of the first processing circuitry. For example, the physical layout can be defined in one-dimensional (1D), two-dimensional (2D), or three-dimensional (3D) space.
[0039] In a 1D physical layout, the processor cores may be arranged on a single line (e.g., as a straight line). In a 2D physical layout, the processor cores may be arranged on a single plane, such as in a square or rectangular arrangement, such as in rows and columns forming a grid pattern. In a 2D physical layout of processor cores, the processor cores may share edges, corners, or no shared borders at all, depending on their placement within the layout. Shared edges provide longer continuous borders for the cores, thereby facilitating a larger area for heat transfer between adjacent processor cores. This means that the more edges a core shares with an adjacent core, the more heat the core can potentially receive from the adjacent core. When cores share a corner, they are connected at a single point, which results in less direct interaction and minimal heat transfer compared to edge sharing. If the cores do not share any borders, the heat transfer between them may be lower.
[0040] In a 3D physical layout, the processor cores can be arranged into cubes or rectangles. For example, the processor cores can be arranged into stacked layers, where each layer can be arranged into rows and columns. In a 3D physical layout, the processor cores can share edges, corners, and surfaces, or have no shared boundaries. Surface sharing between stacked cores allows for even more significant heat transfer than edge sharing due to the larger contact area. Similarly, sharing edges and corners in 3D space also helps with heat transfer, but to a lesser extent than surface sharing. Cores that do not share any direct contact in a 3D arrangement are minimally affected by the heat generated from their neighbors, similar to the dynamics in a 2D configuration.
[0041] In some examples, the first processing circuitry is the same as processing circuitry 130. In another example, the first processing circuitry and processing circuitry 130 are different processing circuitry. For example, processing circuitry 130 may obtain data such as temperature measurement data, etc., from the first processing circuitry via interface circuitry 120.
[0042] For example, the inferred physical layout of the first processing circuitry can be used for various purposes. In some examples, processing circuitry 130 (or the first processing circuitry if the first processing circuitry is different from processing circuitry 130) can be configured to distribute the processing load among the multiple processor cores of the first processing circuitry based on the inferred physical layout. For example, processing circuitry 130 (or the first processing circuitry if the first processing circuitry is different from processing circuitry 130) can distribute the workload across multiple cores that are physically farther apart to achieve a more even temperature distribution within the first processing circuitry. By distributing the temperature more evenly across the first processing circuitry (i.e., the die), the cooling system can require less energy, resulting in improved energy efficiency. Further, by distributing the temperature more evenly across the first processing circuitry, the multiple processor cores may be less likely to experience thermal throttling, resulting in higher performance. Further, by distributing the temperature more evenly across the first processing circuitry, thermal degradation of the first processing circuitry can be reduced, resulting in a longer service life.
[0043] In some examples, the processing circuit system 130 is configured to determine a numerical value for each processor core in the plurality of processor cores. For a particular processor core, the numerical value for the particular processor core may describe a corresponding relationship between a temperature pattern of temperature measurement data of the particular processor core from a temperature measurement data set for the particular processor core and a temperature pattern of corresponding temperature measurement data of other processor cores from a temperature measurement data set for the particular processor core. In other words, for each processor core, the numerical value may be determined to be the same numerical value as the numerical value in the case where the processor core exists. The temperature pattern of a particular processor core may be a temperature profile of the particular processor core. For example, it may be a temperature profile when adjacent processor cores execute workloads and radiate heat, or it may be a temperature profile when a particular processor core executes workloads and is thereby stressed (i.e., actively heated by executing workloads).
[0044] For example, a numerical value for a particular processor core may describe the relationship between the temperature profile of the particular processor core and the temperature profiles of all other processor cores. For example, for a particular processor core, there may be a numerical value that describes the relationship between the temperature profile of the particular processor core and the temperature profile of one of the processor cores. For example, if there are N processor cores, there may be N×N numerical values, or N×(N-1) numerical values. For example, the relationship between the temperature profile of a particular processor core and the temperature profile of the processor core itself may be defined as 1 or as a constant.
[0045] For example, the relationship between the temperature profile of a particular processor core and the temperature profile of another processor core can be approximately described by a polynomial relationship (e.g., linear, quadratic, or cubic) or by an exponential or logarithmic relationship, etc. That is, an increase in the temperature of a particular processor core may cause a linear, quadratic, cubic, exponential, or logarithmic temperature increase in the other processor cores. In some examples, the corresponding relationship is a linear relationship.
[0046] In some examples, processing circuitry 130 may be configured to perform a regression analysis for each of the plurality of processor cores. For a particular processor core, a regression analysis may be performed between temperature measurement data from a temperature measurement dataset for the particular processor core and corresponding temperature measurement data from temperature measurement datasets for other processor cores.
[0047] For example, regression analysis can be a technique used to examine the relationship between a dependent variable and one or more independent variables. For example, regression analysis can involve determining one or more regression coefficients, which are numerical values that quantify the expected change in the dependent variable for a one-unit change in each independent variable while holding the other variables constant. For example, regression analysis can be polynomial regression, which fits a polynomial equation to the data. For example, regression analysis can be linear regression, in which the relationship is modeled as a straight line, which is a special case of polynomial regression. In this case, the regression coefficients can be referred to as linear regression coefficients. That is, the linear regression coefficients tell how much the dependent variable is expected to increase (or decrease) for each one-unit increase in the independent variable. For example, regression analysis can be logistic regression, which can be used for binary outcomes. In other words, for a particular processor core (which may be executing a workload and therefore stressed), regression analysis can be performed between the temperature measurement data of that particular processor core (the dependent variable) and the corresponding temperature measurement data of another processor core (the independent variable) to determine how the temperature measurement of the other processor core changes depending on the temperature measurement of the particular stressed processor core. This can be performed for each processor core while it is stressed (i.e., executing a workload). For example, a regression analysis between a particular processor core and itself might produce a regression coefficient of 1.
[0048] In some examples, the processing circuitry 130 can be configured to perform a linear regression analysis for each of the plurality of processor cores, wherein, for a particular processor core, the linear regression analysis is performed between temperature measurement data of the particular processor core from the temperature measurement dataset of the particular processor core and corresponding temperature measurement data of other processor cores from the temperature measurement dataset of the particular processor core. Further, the processing circuitry 130 can be configured to determine a linear regression coefficient for each of the plurality of processor cores, wherein, for a particular processor core, the linear regression coefficient for the particular processor core describes a corresponding linear relationship between the temperature measurement data of the particular processor core from the temperature measurement dataset of the particular processor core and the corresponding temperature measurement data of the other processor cores from the temperature measurement dataset of the particular processor core.
[0049] This could, for example, result in a matrix-like structure where for each processor core (when it is stressed) the linear coefficients of each other processor core are determined (see also Tables 1 and Figure 5 and Figure 6 ). The linear regression coefficient between a particular processor core and itself can be 1.
[0050] In some examples, the processing circuit system 130 can be configured to determine, for each processor core of the plurality of processors, a clustering of the plurality of processor cores into a plurality of clusters. For a particular processor core, the clustering can be based on a temperature measurement data set corresponding to the particular processor core. Clustering can be a technique for organizing a set of objects into groups or clusters based on the similarity of certain features or attributes. For example, for a particular processor core, the corresponding clustering of all the plurality of processor cores into a plurality of clusters is based on a regression coefficient determined for each processor core when the particular processor core is stressed. This clustering can be performed for each processor core in the processor cores. That is, as many clusterings as the number of processor cores can be performed. For example, from the perspective of a particular stressed processor core, the clustering of the processor cores into clusters can indicate whether another processor core is close to, far away from, or very far away from the particular stressed processor core. In other words, from the perspective of the particular stressed processor core, the clustering clusters the processor cores into similar temperature response patterns.
[0051] For example, well-known clustering algorithms such as k-means, k-medoid, hierarchical clustering, density-based spatial clustering with noise applications, spectral clustering, mean-shift clustering, Gaussian Mixture Model (GMM), and agglomerative hierarchical clustering are used for clustering. For example, the k-means algorithm performs clustering of multiple processor cores into a predetermined number of k clusters. The k clusters can be non-overlapping clusters (each processor is in only one cluster) based on the distance from the mean of the points in the cluster, which minimizes the within-cluster sum of squares (variance). For example, for a particular processor core and its corresponding regression coefficient, the k-means algorithm can iteratively assign each processor core to a cluster based on the regression coefficient. For example, the total distance between the processor cores (i.e., their regression coefficients) and their corresponding cluster centers (means) is minimized, thereby ensuring that cores with similar temperatures are grouped together (see also Figure 7). The number of clusters can be fixed, i.e., k=4 for a 1D / 2D physical layout, or k=5 for a 3D physical layout. For example, cluster A may include stressed processor cores (also referred to as stressed cores). For example, cluster B may be the cluster of processor cores closest to cluster A. However, there may not be a predetermined range for the coefficients in cluster B. In some examples, there may be a spatial limit on how many processor cores may be in a cluster. For example, cluster B represents processor cores that share surfaces with stressed cores (cluster A). If the stressed cores are viewed as cubes, then in a 3D physical layout of processor cores, there are only 6 surfaces on the cube corresponding to the maximum 6 cores in cluster B. In a 1D / 2D physical layout, the maximum number of cores in cluster B may be 4.
[0052] In yet another example, for a particular processor core, each of the plurality of clusters may include processor cores corresponding to values within a certain range for the particular processor core. For example, the values may be regression coefficients. For example, each cluster may be defined by a predetermined range, and each regression coefficient within the predetermined range may be within its corresponding cluster.
[0053] In some examples, processing circuitry 130 can be configured to determine the physical layout of the first processing circuitry based on the determined numerical values for each of the plurality of processor cores. The numerical values indicate how the temperature profile of a particular processor core behaves when the particular processor core is stressed. Because processor cores farther from the stressed core heat up more slowly than processor cores closer to the stressed core, this can be indicated by the numerical value. This can yield the relative positioning of each processor relative to the other processors. Based on this information, the physical layout can be inferred.
[0054] In some examples, processing circuitry 130 can be configured to determine a physical layout of the first processing circuitry based on the determined clustering of each processor core in the plurality of processor cores. As described above, the numerical values indicate how the temperature profile of the processor core behaves when a particular processor core is stressed. This can further result in the clustering of the processor cores into clusters that, from the perspective of a particular stressed processor core, indicate whether another core is close to, far from, or very far from the particular stressed processor core. Based on these clusters, a physical layout can be inferred that indicates the relative positioning of each processor with respect to each other processor and can be obtained for each processor core when each processor core is stressed.
[0055] In some examples, an artificial neural network (ANN) can be trained to infer the physical layout. For example, the ANN can receive as input the clustering for each processor core of the processing circuitry and the corresponding physical layout of the processing circuitry during training to perform supervised learning. After the ANN is trained, it can be used to infer the physical layout of the first processing circuitry based on the clustering for each processor core.
[0056] In some examples, the processing circuit system 130 can be configured to determine the physical layout of the first processing circuit system based on correlating the determined clusters of each processor core in the plurality of processor cores. In some examples, only some of the determined clusters of some processor cores in the plurality of processor cores can be correlated. For example, correlation can refer to a statistical technique for inferring the degree of association between the clustering results obtained from each stressed processor core. The correlation between the clustering results of each stressed processor core involves comparing the clustering outputs to see which processor cores consistently appear in similar clusters across different sets of stressed processor cores. A high degree of correlation in the clustering pattern indicates that some processor cores may be physically very close. By analyzing these correlations, the spatial arrangement of the processor cores within the first processing circuit system can be inferred - processor cores that frequently appear in the same cluster may be located close to each other. For example, correlation can be visualized and quantified using specific methods such as multidimensional scaling (MDS) or principal component analysis (PCA) to visualize and quantify the similarity between the clustering results for each processor core. By applying these techniques, data derived from k-means clustering of regression coefficients can be transformed into a spatial representation where the distance between points on the graph (representing cores) corresponds to their degree of similarity in the clustering results. This visual and quantitative analysis can show patterns, such as clusters of cores that are consistently grouped together across multiple tests, indicating that they are in close proximity or have similar thermal responses within the processor layout.
[0057] In some examples, processing circuitry 130 may be configured to determine a positional classification within the physical layout of the first processing circuitry for each of the plurality of processor cores based on at least one of the determined clusters of the plurality of processor cores. Further, processing circuitry 130 may be configured to determine the physical layout of the first processing circuitry based on the determined positional classification and the clustering of each of the plurality of processor cores. The positional classification may define the position of the processor core relative to a dominant geometry of the physical layout of the first processing circuitry. The dominant geometry of the physical layout of the first processing circuitry may define the physical layout as a line, a square, a rectangle, a cube, or the like. Thus, the positional classification may define the processor core as being close to or far from the center or outside of the dominant geometry of the physical layout of the first processing circuitry. For example, the positional classification may be a corner position, an edge position, or a center position of the processor core within the physical layout. In the case where the physical layout is a 3D layout, the positional classification may be a bottom layer position, a middle layer position, an upper layer position, or the like. In a first step of the dependency analysis, each processor core may be classified into a locality classification within the physical layout, and then in a second step of the dependency analysis, based on these locality classifications, a final physical layout of the first processing circuitry may be determined.
[0058] In some examples, processing circuitry 130 may be configured to infer the physical layout of a first processing circuitry system comprising a plurality of processor cores as follows: 1. Processing circuitry 130 selects a stressed core X (e.g., a core running a power-intensive workload) and, in a 3D layout, selects all of the cores in the second, third, and fourth clusters in a manner that maximizes thermal affinity without violating spatial rules (e.g., no more than six neighboring cores, 12 edge cores, etc.). 2. Processing circuitry 130 selects cores from those surrounding core X, starting with cores in the second cluster, then the third cluster, and then the fourth cluster, and repeats step 1 for that core to fill in the nearby cores. This process stops when all cores in the second, third, and fourth clusters have been mapped. 3. If there are unmapped cores in steps 1 and 2, processing circuitry 130 repeats steps 1 and 2 until all cores in the system have been mapped (i.e., this may result in many separate islands of mapped cores due to multiple tiles, dies, or otherwise thermally isolated cores). 4. If there are multiple islands of mapped cores, the processing circuitry 130 orients them based on the knowledge gained from the cores that fell into the 5th cluster in stage 3.
[0059] Further details and aspects are mentioned in conjunction with the examples described below. The examples shown in FIG. 1 may include those described in conjunction with the concepts proposed or described below (e.g., Figure 1B-Figure 8 ) One or more optional additional features corresponding to one or more aspects mentioned in one or more examples described.
[0060] Figure 1B The figure shows a block diagram of an example of an apparatus 200 or device 200. The apparatus 200 includes circuitry configured to provide the functionality of the apparatus 200. For example, Figure 1B The device 200 includes an interface circuit system 220, a processing circuit system 230, and (optionally) a storage circuit system 240. For example, the processing circuit system 230 may be coupled to the interface circuit system 220 and optionally to the storage circuit system 240. The circuit system of the device 200 may be similar to Figure 1A The circuit system of the device 100.
[0061] For example, processing circuitry 230 may be configured to provide the functionality of apparatus 200 in conjunction with interface circuitry 220. Interface circuitry 220 is configured to exchange information, for example, with other components internal or external to apparatus 200 and storage circuitry 240. Likewise, apparatus 200 may include apparatus configured to provide the functionality of apparatus 200.
[0062] Processing circuitry 230 is configured to control each of the components of the computer architecture block to process a corresponding workload. Processing circuitry 230 is further configured to obtain temperature measurement data from each component of the computer architecture block. The temperature measurement data is obtained while the corresponding component is processing the corresponding workload. Processing circuitry 230 is further configured to infer the physical layout of the computer architecture block based on the temperature measurement data obtained from each component of the computer architecture block.
[0063] A computer architecture block may refer to a different functional unit within a computer system, which may include one or more components designed to perform specific tasks that are integral to the overall operation and performance of the system. For example, a computer system may be connected to a processing circuit system 230 (in some examples, the computer system may be device 200) via an interface circuit system 220. A computer architecture block may be a memory that stores and retrieves data and instructions for use by a processor. In some examples, a computer architecture block may be a non-core, which may encompass various non-core elements such as memory controllers, interconnects, and peripheral controllers that support a processor core by managing data flow and connectivity. In some examples, a computer architecture block may be an AI computing unit, which may be dedicated hardware dedicated to accelerating artificial intelligence computing.
[0064] With the above Figure 1ASimilar to the description of inferring the physical layout of the first processing circuitry in , the physical layout of computer architecture blocks can be inferred.
[0065] Further details and aspects are mentioned in connection with the examples described above or below. Figure 1B The examples shown in may include those in conjunction with the concepts proposed or above (e.g., Figure 1A ) or below (for example, Figure 2A-Figure 8 ) One or more optional additional features corresponding to one or more aspects mentioned in one or more examples described.
[0066] Figure 2A The figure shows a flowchart of an example of a method 300. For example, the method 300 can be performed by an apparatus as described herein, such as the apparatus 100. The method 300 includes controlling (310) each of a plurality of processor cores of a first processing circuit system to execute a corresponding workload. The method 300 further includes obtaining (320) temperature measurement data from each of the plurality of processor cores, wherein the temperature measurement data is obtained during execution of the corresponding workload by the corresponding processor core of the plurality of processor cores. The method 300 further includes inferring (330) a physical layout of the first processing circuit system based on the temperature measurement data obtained from each of the plurality of processor cores of the first processing circuit system.
[0067] In combination with the proposed technology or above (for example, reference Figure 1A ) or below regarding Figures 2B to 8 The described one or more examples explain further details and aspects of the method 300. The method 300 may include one or more additional optional features corresponding to one or more aspects of the proposed technology, or one or more examples described above or below.
[0068] Figure 2B The figure shows a flowchart of an example of a method 400. For example, the method 400 can be performed by an apparatus as described herein, such as apparatus 200. The method 400 includes controlling (410) each component of a computer architecture block to process a corresponding workload. The method 400 further includes obtaining (420) temperature measurement data from each component of the computer architecture block. The temperature measurement data is obtained while the corresponding component processes the corresponding workload. The method 400 further includes inferring a physical layout of the computer architecture block based on the temperature measurement data obtained from each component of the computer architecture block.
[0069] In combination with the proposed technology or above (for example, reference Figure 1A – Figure 2A ) or below with respect to Figures 3 to Figure 8The described one or more examples explain further details and aspects of the method 400. The method 400 may include one or more additional optional features corresponding to one or more aspects of the proposed technology, or one or more examples described above or below.
[0070] Example
[0071] In one embodiment of the present disclosure, the concept of inferring the physical layout of a processing circuit system including multiple processor cores may include (one, some, or all of) the following four stages: Stage 1: Collect temperature data (of multiple cores) by heating one core at a time using a power-intensive workload. When a particular core is running this power-intensive workload, all other cores may be idle. Stage 2: Use the collected temperature data and perform a linear regression between the stressed core and all other cores. Stage 3: Perform a cluster analysis on the regression data to determine nearby cores for each stressed core. Stage 4: Correlate the nearby cores of each stressed core to determine the physical layout of the processing circuit system on the die. These four stages are described in more detail below.
[0072] Phase 1: Temperature Data Collection:
[0073] In the first stage of the disclosed technique for inferring the physical layout of a processing circuit system including multiple processor cores, thermal telemetry data can be collected from each processor in the system under test. Thus, for example, publicly available performance registers can be used to monitor core temperature. For example, the following procedure can be performed in this regard: 1. Collect all individual core temperatures (of all processor cores in the processing circuit system) at predetermined time intervals (e.g., every 200 milliseconds, etc.). 2. Bind a power-intensive workload to one core. This may cause the core to heat up over time. 3. Stop collecting core temperatures. 4. Repeat steps 1, 2, and 3 for each core in the processing circuit system (system).
[0074] Figure 3A The graph shows the temperature distribution over the physical layout of the first processing circuitry when the first core is stressed. Figure 3B The figure shows the temperature distribution over the physical layout of the first processing circuit system when the second core is stressed. Dark red indicates the highest temperature. From light red to purple, from dark blue to black, the temperature decreases.
[0075] Figure 4A temperature profile of a stressed processor core over time and a temperature profile of an idle processor core over time are shown. Both the stressed processor core and the idle processor core are within the same processing circuitry. The rate at which the two temperature profiles change and the temperature difference between the final temperatures of the two profiles can be inferred from the profiles and used as key metrics for determining the proximity of the cores within the physical layout of the processing circuitry.
[0076] Stage 2: Linear regression of collected temperature data:
[0077] The second stage of the disclosed technique for inferring the physical layout of a processing circuitry comprising multiple processor cores may include performing a linear regression on the temperature data collected in stage 1. For example, the following procedure may be performed in this regard: 1. Run a linear regression between the stressed core and another core of the processing circuitry and obtain coefficients representing the temperature relationship between the two cores. 2. Repeat step 1 for each core on the system and obtain N-1 coefficients, where N is the number of processor cores in the processing circuitry (i.e., the system). 3. Repeat steps 1 and 2 for N data sets (each data set collected in stage 1 and corresponding to a different stressed core).
[0078] Figure 5 The graph shows the linear regression between the collected temperature data of the stressed core and the temperature data of the idle core (see Figure 4 In this regard, Figure 5 The graph in Figure 2 illustrates the change in the temperature of the idle core as a function of the temperature of the stressed core. This graph can be described by the linear equation y = mx + b, where y is the temperature of the idle core, x is the temperature of the stressed core, b is the y-intercept, and m is the search rate (coefficient), representing the change in degrees Celsius in the idle core's temperature for every degree Celsius change in the stressed core's temperature. Therefore, the larger the coefficient m between the idle and stressed cores (closer to m = 1), the closer the two cores are. Conversely, the lower the coefficient between the idle and stressed cores (closer to m = 0), the further the two cores are apart.
[0079] After inferring N-1 regression coefficients for each of the N processor cores, an ordered list (matrix) of coefficients for each stressed core may be created, as illustrated in Table 1:
[0080]
[0081] Table 1
[0082] Stage 3: Cluster analysis based on regression data:
[0083] The third stage of the disclosed technique for inferring the physical layout of the processing circuitry including multiple processor cores may include performing a cluster analysis on the table output of stage 2. This may be done using rule-based algorithms or machine learning algorithms known to those skilled in the art, such as the K-means algorithm. For example, Figure 6 As shown in , the top 4-5 clusters can be identified.
[0084] Figure 6 The figure illustrates a cluster analysis performed on the obtained regression coefficients of one core. The cluster analysis can be performed by a K-means algorithm (e.g., K=4 or 5) based on the regression coefficients of one core obtained in stage 2. This clustering can be performed for each of the N-1 regression coefficients of the N cores as obtained in stage 2. The identified clusters can have the following definitions: Cluster 1 can be the processing core running the power-intensive workload, as illustrated by the orange dots in box 612. Cluster 2 can be the core closest to the processing core running the power-intensive workload, i.e., a neighboring core that shares a surface (with the processing core running the power-intensive workload), as illustrated by the blue dots in box 610. Cluster 3 can be the second closest processing core, i.e., an edge core that shares an edge (with the processing core running the power-intensive workload), as illustrated by the green dots in box 608. Cluster 4 may be [this may be for stacked / 3D layout only] the third closest core, i.e., the diagonal core that shares a cube corner (with the processing core running the power intensive workload), as illustrated by the cyan dots in box 606 (see also Figure 7 ). Cluster 5 may be all other cores that may be separated from the stressed core by one or more cores, as illustrated by the purple and red dots in box 604. Figure 7 and Figure 8 This is also described.
[0085] Phase 4: Correlate each stressed core with its neighboring cores to determine the physical layout:
[0086] The fourth stage of the disclosed technique for inferring the physical layout of a processing circuit system including multiple processor cores may include correlating the neighboring cores of each stressed core to determine the physical layout. This may be based on the clustering results for each stressed core obtained in stage 3. Thus, possible physical layouts (die layouts) in 1D, 2D, or 3D may be programmatically created. For example, the process of mapping the processing cores according to their clustering / coefficients may be as follows:
[0087] 1. Select stressed core X (the core running the power-intensive workload) and, in a 3D spatial layout, select all of the cores in the 2nd, 3rd, and 4th clusters in a way that maximizes thermal affinity without violating spatial rules (e.g., no more than 6 neighboring cores, 12 edge cores, etc.).
[0088] 2. Pick cores from those surrounding core X, starting with the cores in the second cluster, then the third cluster, and then the fourth cluster, and repeat step 1 for that core to fill in nearby cores. This process stops when all cores in the 2nd, 3rd, and 4th clusters have been mapped.
[0089] 3. If there are unmapped cores from steps 1 and 2, repeat steps 1 and 2 until all cores in the system have been mapped (i.e., this may result in many separate islands of mapped cores due to multiple tiles, dies, or otherwise thermally isolated cores).
[0090] 4. If there are multiple islands of mapped cores, they are oriented based on the knowledge gained from the cores that fell into the 5th cluster in stage 3.
[0091] Figure 7 The diagram illustrates an embodiment of a standard configuration of cores belonging to different clusters or dies in a 1D, 2D or 3D layout. In block 702, a 1D layout of the processing cores is illustrated, i.e., the cores may be arranged in rows. In this regard, the cores may be clustered into corner cores or middle cores based on the results of stage 3. In block 704, a 2D layout of the processing cores is illustrated, i.e., the cores may be arranged in a plane. In this regard, the cores may be clustered into corner cores or middle cores and edge cores based on the results of stage 3 (see also Figure 8 In block 706, a 3D layout of the processing cores is illustrated, i.e., the cores may be arranged in a cube. In this regard, the cores may be clustered into corner cores or middle cores and edge cores and diagonal cores based on the results of stage 3 (e.g., see cluster 4 in stage 3).
[0092] If the die layout is a 2D layout, then a correlation process such as Figure 8 As shown in , within the physical 2D layout, there may be three possible placements for any given core. Figure 8 The diagram illustrates three possible placements for any given core within the physical layout of the first processing circuitry. As shown in block 802, the core may be placed at a corner of the physical layout of the processing circuitry. As shown in block 804, the core may be placed at an edge of the physical layout of the processing circuitry, rather than at a corner. As shown in block 806, the core may be placed in the middle of the physical layout of the processing circuitry (and neither at a corner nor at an edge).
[0093] Further details and aspects are mentioned in connection with the examples described above or below. Figure 3A-Figure 8 The examples shown in may include those in conjunction with the concepts proposed or above (e.g., Figures 1A-2B ) One or more optional additional features corresponding to one or more aspects mentioned in one or more examples described.
[0094] In the following, some examples of the proposed concept are presented:
[0095] An example (e.g., Example 1) relates to a device comprising an interface circuit system, machine-readable instructions, and a processing circuit system for executing the machine-readable instructions to perform the following operations: controlling each of a plurality of processor cores of a first processing circuit system to execute a corresponding workload; obtaining temperature measurement data from each of the plurality of processor cores, wherein the temperature measurement data is obtained during execution of the corresponding workload by the corresponding processor core of the plurality of processor cores; and inferring a physical layout of the first processing circuit system based on the temperature measurement data obtained from each of the plurality of processor cores of the first processing circuit system.
[0096] Another example (e.g., Example 2) relates to the aforementioned example (e.g., Example 1) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: distributing the processing load among multiple processor cores of the first processing circuit system based on the inferred physical layout.
[0097] Another example (e.g., Example 3) relates to the aforementioned example (e.g., one of Examples 1 or 2) or to any other example, further including: for each processor core among a plurality of processor cores, the obtained temperature measurement data includes a corresponding temperature measurement data set corresponding to the corresponding processor core, wherein the temperature measurement data set corresponding to a specific processor core includes temperature measurement data of the plurality of processor cores obtained when the specific processor core executes a corresponding workload.
[0098] Another example (e.g., Example 4) relates to the aforementioned example (e.g., Example 3) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: determining a numerical value for each processor core of a plurality of processor cores, wherein, for a particular processor core, the numerical value for the particular processor core describes a corresponding relationship between a temperature pattern of temperature measurement data of the particular processor core from a temperature measurement data set of the particular processor core and a temperature pattern of corresponding temperature measurement data of other processor cores from a temperature measurement data set of the particular processor core.
[0099] Another example (eg, Example 5) relates to the aforementioned example (eg, Example 4) or to any other example, further including: the corresponding relationship is a linear relationship.
[0100] Another example (e.g., Example 6) relates to the preceding example (e.g., one of Examples 3 to 5) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: performing a regression analysis for each processor core of a plurality of processor cores, wherein, for a particular processor core, the regression analysis is performed between temperature measurement data of the particular processor core from a temperature measurement dataset of the particular processor core and corresponding temperature measurement data of other processor cores from temperature test datasets of the particular processor core.
[0101] Another example (e.g., Example 7) relates to the aforementioned example (e.g., one of Examples 3 to 6) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: performing a linear regression analysis for each processor core of a plurality of processor cores, wherein, for a specific processor core, the linear regression analysis is performed between the temperature measurement data of the specific processor core from the temperature measurement data set of the specific processor core and the corresponding temperature measurement data of other processor cores from the temperature measurement data set of the specific processor core; and determining a linear regression coefficient for each processor core of the plurality of processor cores, wherein, for a specific processor core, the linear regression coefficient for the specific processor core describes the corresponding linear relationship between the temperature measurement data of the specific processor core from the temperature measurement data set of the specific processor core and the corresponding temperature measurement data of the other processor cores from the temperature test data set of the specific processor core.
[0102] Another example (e.g., Example 8) relates to the preceding example (e.g., one of Examples 3 to 7) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: determining, for each processor core in a plurality of processor cores, clustering the plurality of processor cores into a plurality of clusters, wherein, for a particular processor core, the clustering is based on a temperature measurement data set corresponding to the particular processor core.
[0103] Another example (e.g., Example 9) relates to the aforementioned example (e.g., one of Examples 4 to 8) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: determining, for each processor core in a plurality of processor cores, clustering the plurality of processor cores into a number of clusters, wherein, for a particular processor core, the clustering is based on a numerical value determined for the particular processor core in the plurality of processor cores.
[0104] Another example (e.g., Example 10) relates to the aforementioned example (e.g., Example 9) or to any other example, further including: for a particular processor core, each cluster in the plurality of clusters includes a processor core corresponding to a numerical value within a certain range of the numerical value of the particular processor core.
[0105] Another example (e.g., Example 11) relates to the aforementioned example (e.g., one of Examples 4 to 10) or to any other example, further including: the processing circuit system is used to execute machine-readable instructions to perform the following operations: determine the physical layout of the first processing circuit system based on the determined numerical value of each processor core in the plurality of processor cores.
[0106] Another example (e.g., Example 12) relates to the aforementioned example (e.g., one of Examples 8 to 11) or to any other example, further including: the processing circuit system is used to execute machine-readable instructions to perform the following operations: determine the physical layout of the first processing circuit system based on the determined clustering of each processor core in the plurality of processor cores.
[0107] Another example (e.g., Example 13) relates to the aforementioned example (e.g., Example 12) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: determining a physical layout of a first processing circuit system based on correlating the determined clusters of each processor core in a plurality of processor cores.
[0108] Another example (e.g., Example 14) relates to the aforementioned example (e.g., one of Examples 8 to 13) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: determining a positional classification within a physical layout of a first processing circuit system for each of a plurality of processor cores based on at least one of the determined clusters of the plurality of processor cores; and determining a physical layout of the first processing circuit system based on the determined positional classification and the clustering of each of the plurality of processor cores.
[0109] Another example (e.g., Example 15) relates to a previous example (e.g., Example 14) or to any other example, further including: the positional classification within the physical layout of the first processing circuit system includes a corner position, an edge position, or a center position of the physical layout of the first processing circuit system.
[0110] Another example (e.g., Example 16) relates to the aforementioned example (e.g., one of Examples 1 to 15) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: controlling each processor core of a plurality of processor cores to execute a corresponding workload for a predetermined time.
[0111] Another example (e.g., Example 17) relates to the aforementioned example (e.g., one of Examples 1 to 16) or to any other example, further including: a processing circuit system for executing machine-readable instructions to perform the following operations: controlling each processor core of a plurality of processor cores to execute a corresponding workload until a change in the temperature of the corresponding temperature measurement data obtained of the processor core executing the workload is below a predetermined threshold.
[0112] Another example (eg, Example 18) relates to the preceding example (eg, one of Examples 1-17) or to any other example, further including: temperature measurement data is acquired at predetermined time intervals.
[0113] Another example (eg, Example 19) relates to a previous example (eg, one of Examples 1-18) or to any other example, further comprising: the physical layout comprising spatial positioning of the processor core within the first processing circuitry.
[0114] An example (e.g., Example 20) relates to an apparatus comprising an interface circuit system, machine-readable instructions, and a processing circuit system for executing the machine-readable instructions to perform the following operations: controlling each of the components of a computer architecture block to process a corresponding workload; obtaining temperature measurement data from each component of the computer architecture block, wherein the temperature measurement data is obtained during the processing of the corresponding workload by the corresponding component; and inferring a physical layout of the computer architecture block based on the temperature measurement data obtained from each component of the computer architecture block.
[0115] Another example (e.g., Example 21) relates to a previous example (e.g., Example 20) or to any other example, further including: the computer architecture block is at least one of the following: memory, non-core, AI computing unit.
[0116] An example (e.g., Example 22) relates to a method comprising: controlling each processor core of a plurality of processor cores of a first processing circuit system to execute a corresponding workload; obtaining temperature measurement data from each processor core of the plurality of processor cores, wherein the temperature measurement data is obtained during execution of the corresponding workload by the corresponding processor core of the plurality of processor cores; and inferring a physical layout of the first processing circuit system based on the temperature measurement data obtained from each processor core of the plurality of processor cores of the first processing circuit system.
[0117] Another example (eg, Example 23) relates to the preceding example (eg, Example 22) or to any other example, further comprising distributing the processing load among the plurality of processor cores of the first processing circuitry based on the inferred physical layout.
[0118] Another example (e.g., Example 24) relates to the aforementioned example (e.g., one of Examples 22 or 23) or to any other example, further including: for each processor core of a plurality of processor cores, the obtained temperature measurement data includes a corresponding temperature measurement data set corresponding to the corresponding processor core, wherein the temperature measurement data set corresponding to a particular processor core includes temperature measurement data of the plurality of processor cores obtained when the particular processor core executes a corresponding workload.
[0119] Another example (e.g., Example 25) relates to the aforementioned example (e.g., Example 24) or to any other example, further including: determining a numerical value for each processor core of a plurality of processor cores, wherein, for a particular processor core, the numerical value for the particular processor core describes a corresponding relationship between a temperature pattern of temperature measurement data of the particular processor core from a temperature measurement data set of the particular processor core and a temperature pattern of corresponding temperature measurement data of other processor cores from a temperature measurement data set of the particular processor core.
[0120] Another example (eg, Example 26) relates to the aforementioned example (eg, Example 25) or to any other example, further including: the corresponding relationship is a linear relationship.
[0121] Another example (e.g., Example 27) relates to the preceding example (e.g., one of Examples 24 to 26) or to any other example, further including: performing a regression analysis for each processor core of a plurality of processor cores, wherein, for a particular processor core, the regression analysis is performed between temperature measurement data of the particular processor core from a temperature measurement dataset of the particular processor core and corresponding temperature measurement data of other processor cores from temperature test datasets of the particular processor core.
[0122] Another example (e.g., Example 28) relates to the aforementioned example (e.g., one of Examples 24 to 27) or to any other example, further including: performing a linear regression analysis for each processor core of a plurality of processor cores, wherein, for a specific processor core, the linear regression analysis is performed between the temperature measurement data of the specific processor core from the temperature measurement dataset of the specific processor core and the corresponding temperature measurement data of the other processor cores from the temperature measurement dataset of the specific processor core; and determining a linear regression coefficient for each processor core of the plurality of processor cores, wherein, for a specific processor core, the linear regression coefficient for the specific processor core describes the corresponding linear relationship between the temperature measurement data of the specific processor core from the temperature measurement dataset of the specific processor core and the corresponding temperature measurement data of the other processor cores from the temperature test dataset of the specific processor core.
[0123] Another example (e.g., Example 29) relates to the preceding example (e.g., one of Examples 24 to 28) or to any other example, further comprising: determining, for each processor core in the plurality of processor cores, clustering the plurality of processor cores into a plurality of clusters, wherein, for a particular processor core, the clustering is based on a temperature measurement data set corresponding to the particular processor core.
[0124] Another example (e.g., Example 30) relates to the aforementioned example (e.g., one of Examples 25 to 29) or to any other example, further including: determining, for each processor core in the plurality of processor cores, clustering the plurality of processor cores into a number of clusters, wherein, for a particular processor core, the clustering is based on a determined numerical value for the particular processor core in the plurality of processor cores.
[0125] Another example (e.g., Example 31) relates to a previous example (e.g., Example 30) or to any other example, further including: for a particular processor core, each cluster in the plurality of clusters includes a processor core corresponding to a numerical value within a certain range of the numerical value of the particular processor core.
[0126] Another example (e.g., Example 32) relates to the aforementioned example (e.g., one of Examples 25 to 31) or to any other example, further including: determining a physical layout of the first processing circuit system based on the determined numerical value of each processor core in the plurality of processor cores.
[0127] Another example (e.g., Example 33) relates to the aforementioned example (e.g., one of Examples 29 to 32) or to any other example, further including: determining a physical layout of the first processing circuit system based on the determined clustering of each processor core in the plurality of processor cores.
[0128] Another example (e.g., Example 34) relates to the aforementioned example (e.g., Example 33) or to any other example, further comprising: determining a physical layout of the first processing circuit system based on correlating the determined clusters of each processor core in the plurality of processor cores.
[0129] Another example (e.g., Example 35) relates to the aforementioned example (e.g., one of Examples 29 to 34) or to any other example, further including: determining a positional classification within the physical layout of the first processing circuit system for each of the multiple processor cores based on at least one of the determined clusters of the multiple processor cores; and determining the physical layout of the first processing circuit system based on the determined positional classification and the clustering of each of the multiple processor cores.
[0130] Another example (e.g., Example 36) relates to a previous example (e.g., Example 35) or to any other example, further including: the positional classification within the physical layout of the first processing circuit system includes a corner position, an edge position, or a center position of the physical layout of the first processing circuit system.
[0131] Another example (e.g., Example 37) relates to the aforementioned example (e.g., one of Examples 22 to 36) or to any other example, further including: controlling each processor core of the plurality of processor cores to execute a corresponding workload for a predetermined time.
[0132] Another example (e.g., Example 38) relates to the aforementioned example (e.g., one of Examples 22 to 37) or to any other example, further including: controlling each processor core of a plurality of processor cores to execute a corresponding workload until a change in the temperature of the corresponding temperature measurement data obtained of the processor core executing the workload is below a predetermined threshold.
[0133] Another example (eg, Example 39) relates to the preceding example (eg, one of Examples 22 to 38) or to any other example, further including: the temperature measurement data is acquired at predetermined time intervals.
[0134] Another example (eg, Example 40) relates to a previous example (eg, one of Examples 22-39) or to any other example, further comprising: the physical layout includes spatial positioning of the processor core within the processing circuit system.
[0135] An example (e.g., Example 41) relates to a method comprising: controlling each of components of a computer architecture block to process a corresponding workload; obtaining temperature measurement data from each component of the computer architecture block, wherein the temperature measurement data is obtained during the processing of the corresponding workload by the corresponding component; and inferring a physical layout of the computer architecture block based on the temperature measurement data obtained from each component of the computer architecture block.
[0136] Another example (e.g., Example 42) relates to a previous example (e.g., Example 41) or to any other example, further including: the computer architecture block is at least one of the following: memory, non-core, AI computing unit.
[0137] An example (e.g., Example 43) relates to an apparatus comprising a processing circuit system configured to: control each of a plurality of processor cores of a first processing circuit system to execute a corresponding workload; obtain temperature measurement data from each of the plurality of processor cores, wherein the temperature measurement data is obtained during execution of the corresponding workload by the corresponding processor core of the plurality of processor cores; and infer a physical layout of the first processing circuit system based on the temperature measurement data obtained from each of the plurality of processor cores of the first processing circuit system.
[0138] An example (e.g., Example 44) relates to a device comprising a means for processing, the means for processing being used to perform the following operations: controlling each of a plurality of processor cores of a first processing circuit system to execute a corresponding workload; obtaining temperature measurement data from each of the plurality of processor cores, wherein the temperature measurement data is obtained during execution of the corresponding workload by the corresponding processor core of the plurality of processor cores; and inferring a physical layout of the first processing circuit system based on the temperature measurement data obtained from each of the plurality of processor cores of the first processing circuit system.
[0139] Another example (eg, Example 45) relates to a non-transitory machine-readable storage medium comprising program code that, when executed, causes a machine to perform any of the methods of Examples 22 to 40 or 41 to 42.
[0140] Another example (eg, Example 46) relates to a computer program having program code for performing any of the methods of Examples 22 to 40 or 41 to 42 when the computer program is executed on a computer, a processor, or a programmable hardware component.
[0141] Another example (eg, Example 47) relates to a machine-readable storage device comprising machine-readable instructions that, when executed, are used to implement a method as claimed in any pending example, or to realize an apparatus as claimed in any pending example.
[0142] Examples may further be or relate to (computer) programs, including program codes for executing one or more of the above methods when the program is executed on a computer, processor, or other programmable hardware component. Thus, the steps, operations, or processes of the different methods in the methods described above may also be performed by a programmed computer, processor, or other programmable hardware component. Examples may also encompass program storage devices, such as digital data storage media, which are machine-readable, processor-readable, or computer-readable and encode and / or include machine-executable, processor-executable, or computer-executable programs and instructions. For example, program storage devices may include or may be digital storage devices, magnetic storage media (such as disks and tapes), hard drives, or optically readable digital data storage media. Other examples may also include a computer, a processor, a control unit, a (field) programmable logic array ((F)PLA), a (field) programmable gate array ((F)PGA), a graphics processor unit (GPU), an application-specific integrated circuit (ASIC), an integrated circuit (IC), or a system-on-a-chip (SoC) system programmed to perform the steps of the method described above.
[0143] It is further understood that the disclosure of several steps, processes, operations, or functions disclosed in the specification or claims should not be interpreted as implying that these operations must be performed in the order described, unless explicitly stated in a separate use case or necessary for technical reasons. Therefore, the previous description does not limit the performance of several steps or functions to a certain order. In addition, in further examples, a single step, function, process, or operation may include and / or be decomposed into several sub-steps, sub-functions, sub-processes, or sub-operations.
[0144] If some aspects have been described in conjunction with a device or system, these aspects should also be understood as descriptions of the corresponding methods. For example, a block, device, or functional aspects of a device or system can correspond to features (such as method steps) of a corresponding method. Accordingly, aspects described in conjunction with a method should also be understood as descriptions of properties or functional features of the corresponding block, corresponding element, corresponding device, or corresponding system.
[0145] As used herein, the term "module" refers to logic for performing one or more operations consistent with the present disclosure that can be implemented using hardware components or devices, software or firmware running on a processing unit, or a combination thereof. Software and firmware can be embodied as instructions and / or data stored on a non-transient computer-readable storage medium. As used herein, the term "circuitry" can include, alone or in any combination, non-programmable (hard-wired) circuitry, programmable circuitry (such as a processing unit), state machine circuitry, and / or firmware that stores instructions that can be executed by a programmable circuitry. The modules described herein can be embodied collectively or individually as circuitry that forms part of a computing system. Thus, any one of the modules can be implemented as a circuitry. A computing system referred to as being programmed to perform a method can be programmed to perform the method via software, hardware, firmware, or a combination thereof.
[0146] Any of the disclosed methods (or portions thereof) may be implemented as computer-executable instructions or computer program products. Such instructions may enable a computing system or one or more processing units capable of executing computer-executable instructions to perform any of the disclosed methods. As used herein, the term "computer" refers to any computing system or device described or mentioned herein. Thus, the term "computer-executable instructions" refers to instructions that can be executed by any computing system or device described or mentioned herein.
[0147] Computer-executable instructions can be, for example, part of an operating system of a computing system, an application stored locally on the computing system, or a remote application accessible to the computing system (e.g., via a web browser). Any of the methods described herein can be performed by computer-executable instructions that are executed by a single computing system or by one or more networked computing systems operating in a network environment. Computer-executable instructions and updates to the computer-executable instructions can be downloaded to a computing system from a remote server.
[0148] Furthermore, it should be understood that implementation of the disclosed techniques is not limited to any particular computer language or program. For example, the disclosed techniques can be implemented by software written in C++, C#, Java, Perl, Python, JavaScript, Adobe Flash, C#, assembly language, or any other programming language. Likewise, the disclosed techniques are not limited to any particular computer system or any particular type of hardware.
[0149] Furthermore, any of the software-based examples (including, for example, computer-executable instructions for causing a computer to perform any of the disclosed methods) can be uploaded, downloaded, or remotely accessed via suitable communication means, including, for example, the Internet, the World Wide Web, an intranet, cables (including fiber optic cables), magnetic communications, electromagnetic communications (including RF, microwave, ultrasonic, and infrared communications), electronic communications, or other such communications means.
[0150] The disclosed methods, apparatus, and systems should not be construed as limiting in any way. Instead, the present disclosure is directed to all novel and non-obvious features and aspects of each disclosed example, individually and in various combinations and subcombinations with each other. The disclosed methods, apparatus, and systems are not limited to any specific aspect or feature or combination thereof, nor do the disclosed examples require any one or more specific advantages to exist or to solve any one or more specific problems.
[0151] The theories of operation, scientific principles, or other theoretical descriptions presented herein with reference to the devices or methods of the present disclosure are provided for the purpose of better understanding and are not intended to limit the scope. The devices and methods in the appended claims are not limited to those devices and methods that function in the manner described by such theories of operation.
[0152] The following claims are hereby incorporated into the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although a dependent claim in the claims refers to a specific combination with one or more other claims, other examples may include combinations of the dependent claim with the subject matter of any other dependent or independent claims. Such combinations are expressly set forth herein unless a specific combination is stated to be unintended in individual cases. Furthermore, even if a claim is not directly limited to referring to any other independent claim, features of that claim should also be included with respect to any other independent claim.
Claims
1. An apparatus comprising interface circuitry, machine-readable instructions, and processing circuitry, the processing circuitry configured to execute the machine-readable instructions to perform the following operations: controlling each processor core of the plurality of processor cores of the first processing circuitry to execute a corresponding workload; Temperature measurement data is obtained from each processor core of the plurality of processor cores, wherein The temperature measurement data is obtained during execution of a corresponding workload by a corresponding processor core among the plurality of processor cores; as well as A physical layout of the first processing circuitry is inferred based on temperature measurement data obtained from each of the plurality of processor cores of the first processing circuitry.
2. The device according to claim 1, wherein The processing circuitry is configured to execute the machine-readable instructions to distribute a processing load among the plurality of processor cores of the first processing circuitry based on the inferred physical layout.
3. The device according to claim 1, wherein For each processor core of the plurality of processor cores, the obtained temperature measurement data includes a corresponding temperature measurement data set corresponding to the corresponding processor core, The temperature measurement data set corresponding to the specific processor core includes temperature measurement data of the multiple processor cores obtained when the specific processor core executes a corresponding workload.
4. The device according to claim 3, wherein The processing circuit system is configured to execute the machine-readable instructions to perform the following operations: determining a value for each processor core of the plurality of processor cores, In which, for a specific processor core, the numerical value for the specific processor core describes the corresponding relationship between the temperature pattern of the temperature measurement data of the specific processor core from the temperature measurement data set of the specific processor core and the temperature pattern of the corresponding temperature measurement data of other processor cores from the temperature measurement data set of the specific processor core.
5. The device according to claim 4, wherein The corresponding relationship is a linear relationship.
6. The device according to claim 3, wherein The processing circuit system is configured to execute the machine-readable instructions to perform the following operations: performing a regression analysis on each processor core of the plurality of processor cores, Wherein, for a specific processor core, the regression analysis is performed between the temperature measurement data of the specific processor core from the temperature measurement data set of the specific processor core and the corresponding temperature measurement data of other processor cores from the temperature measurement data set of the specific processor core.
7. The device according to claim 3, wherein The processing circuit system is configured to execute the machine-readable instructions to perform the following operations: performing a linear regression analysis for each processor core of the plurality of processor cores, wherein, for a particular processor core, the linear regression analysis is performed between temperature measurement data of the particular processor core from the temperature measurement data set of the particular processor core and corresponding temperature measurement data of other processor cores from the temperature measurement data set of the particular processor core; and A linear regression coefficient is determined for each processor core of the plurality of processor cores, wherein, for a particular processor core, the linear regression coefficient for the particular processor core describes a respective linear relationship between temperature measurement data of the particular processor core from the temperature measurement dataset of the particular processor core and respective temperature measurement data of other processor cores from the temperature measurement dataset of the particular processor core.
8. The device according to claim 3, wherein The processing circuitry is configured to execute the machine-readable instructions to: determine, for each processor core of the plurality of processor cores, a clustering of the plurality of processor cores into a plurality of clusters; Wherein, for a specific processor core, the clustering is based on a temperature measurement data set corresponding to the specific processor core.
9. The device according to claim 4, wherein The processing circuitry is configured to execute the machine-readable instructions to: determine, for each processor core of the plurality of processor cores, a clustering of the plurality of processor cores into a plurality of clusters; Wherein, for a specific processor core, the clustering is based on a value determined for the specific processor core among the plurality of processor cores.
10. The device according to claim 9, wherein For a particular processor core, each cluster in the plurality of clusters includes a processor core corresponding to a value within a certain range of the value of the particular processor core.
11. The device according to claim 4, wherein The processing circuitry is to execute the machine-readable instructions to determine the physical layout of the first processing circuitry based on the determined value of each processor core in the plurality of processor cores.
12. The device according to claim 8, wherein The processing circuitry is to execute the machine-readable instructions to determine the physical layout of the first processing circuitry based on the determined clustering of each processor core in the plurality of processor cores.
13. The device according to claim 12, wherein The processing circuitry is to execute the machine-readable instructions to determine the physical layout of the first processing circuitry based on correlating the determined clustering of each processor core in the plurality of processor cores.
14. The device according to claim 8, wherein The processing circuit system is configured to execute the machine-readable instructions to perform the following operations: determining, for each of the plurality of processor cores, a positional classification within the physical layout of the first processing circuitry based on at least one of the determined clusters of the plurality of processor cores; as well as The physical layout of the first processing circuitry is determined based on the determined locality classification and the clustering of each processor core in the plurality of processor cores.
15. The device according to claim 14, wherein The positional classification within the physical layout of the first processing circuitry comprises a corner position, an edge position, or a center position of the physical layout of the first processing circuitry.
16. The device according to claim 1, wherein The processing circuit system is configured to execute the machine-readable instructions to perform the following operations: controlling each processor core of the plurality of processor cores to execute a corresponding workload for a predetermined time.
17. The device according to claim 1, wherein The processing circuit system is configured to execute the machine-readable instructions to perform the following operations: controlling each processor core of the plurality of processor cores to execute a corresponding workload until a change in temperature of corresponding temperature measurement data obtained for the processor core executing the workload is lower than a predetermined threshold.
18. The device according to claim 1, wherein The physical layout includes the spatial positioning of a processor core within the first processing circuitry.
19. An apparatus comprising interface circuitry, machine-readable instructions, and processing circuitry, the processing circuitry configured to execute the machine-readable instructions to: Each of the components of the control computer architecture block processes a corresponding workload; Temperature measurement data is obtained from each component of the computer architecture block, wherein The temperature measurement data is obtained while the corresponding component is processing the corresponding workload; and A physical layout of the computer architecture block is inferred based on temperature measurement data obtained from each component of the computer architecture block.
20. A method comprising: controlling each processor core of the plurality of processor cores of the first processing circuitry to execute a corresponding workload; obtaining temperature measurement data from each of the plurality of processor cores, wherein the temperature measurement data is obtained during execution of a corresponding workload by a corresponding processor core of the plurality of processor cores; as well as A physical layout of the first processing circuitry is inferred based on temperature measurement data obtained from each of the plurality of processor cores of the first processing circuitry.
21. The method according to claim 20, further comprising: A processing load is distributed among the plurality of processor cores of the first processing circuitry based on the inferred physical layout.
22. The method according to claim 20, wherein For each processor core of the plurality of processor cores, the obtained temperature measurement data includes a corresponding temperature measurement data set corresponding to the corresponding processor core, wherein the temperature measurement data set corresponding to a particular processor core includes temperature measurement data of the plurality of processor cores obtained when the particular processor core executes a corresponding workload.
23. The method of claim 20, further comprising: determining a value for each processor core of the plurality of processor cores, In which, for a specific processor core, the numerical value for the specific processor core describes the corresponding relationship between the temperature pattern of the temperature measurement data of the specific processor core from the temperature measurement data set of the specific processor core and the temperature pattern of the corresponding temperature measurement data of other processor cores from the temperature measurement data set of the specific processor core.
24. A non-transitory machine-readable storage medium comprising a program code, which, when executed, causes a machine to perform the method according to claim 20.
25. A computer program product comprising a program code for performing the method according to claim 20 when the computer program is executed on a computer, a processor or a programmable hardware component.