Cache memory mounted arithmetic operation device

By introducing an operation management unit into the three-dimensional stacking chip, selectively running the computing unit and cache memory and ensuring partial overlap, the problem of the heat of the computing unit affecting the cache memory is solved, and higher computing power and cache capacity are achieved.

JP2025072119APending Publication Date: 2025-05-09FUJITSU LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023182658
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In a three-dimensional stacking chip, when the computing unit and the cache memory are stacked, the heat generated by the computing unit may cause the cache memory to fail.

Method used

By setting up an operation management unit between the computing unit and the cache memory, the computing unit and the cache memory are selectively run according to the calculation intensity requirements and ensuring that they overlap at least partially on the plan view to reduce the impact of caloric on the cache memory.

Benefits of technology

It effectively reduces the impact of heat generated by computing units on cache memory operations, improves the thermal management capabilities of the system, and allows the number of computing units and cache memory to be increased in a given space, thereby improving computing power and cache capacity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072119000001_ABST
    Figure 2025072119000001_ABST
Patent Text Reader

Abstract

To mitigate influence to operation of a cache memory from heat evolution of an arithmetic element in a three-dimensional lamination die in which the arithmetic element and the cache memory are laminated.SOLUTION: A cache memory mounted arithmetic operation device 1 includes: a first semiconductor die including a plurality of arithmetic elements 10; a second semiconductor die including a plurality of cache memories 20 being a pair with either one of the plurality of arithmetic elements 10 and laminated on the first semiconductor die; and an operation management unit 31 for managing the operation of the first semiconductor die and the second semiconductor die. The arithmetic element 10 and the cache memory 20 being a pair with each other are at least partially overlapped in a plan view. According to arithmetic strength required to the plurality of arithmetic elements 10, the operation management unit 31 selectively allows either one of the arithmetic element 10 and the cache memory 20 being the pair with the arithmetic element 10 to operate.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a processor incorporating a cache memory. [Background technology]

[0002] There is known a computing device including a three-dimensional stacked die in which a plurality of semiconductor dies are stacked. For example, a three-dimensional stacked die in which a memory die and a logic die for memory latency control are stacked has been proposed.

[0003] The arithmetic device includes an arithmetic unit that executes arithmetic operations and a cache memory that stores data. In order to increase the number of arithmetic units called cores and the number of cache memories to improve the arithmetic performance and memory capacity, it is advantageous to adopt a stacked structure in which the arithmetic units and the cache memories are stacked. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] US Patent Application Publication No. 2023 / 0125009 Summary of the Invention [Problem to be solved by the invention]

[0005] However, when the arithmetic unit and the cache memory are stacked, there is a risk that the cache memory will become inoperable due to heat generated by the arithmetic unit.

[0006] One aspect of the present invention aims to reduce the effect of heat generation from a computing unit on the operation of a cache memory in a three-dimensional stacked die in which a computing unit and a cache memory are stacked. [Means for solving the problem]

[0007] In one aspect, a cache memory-equipped computing device includes a first semiconductor die including a plurality of computing units, and a second semiconductor die. The second semiconductor die includes a plurality of cache memories that are paired with any of the plurality of computing units, and is stacked on the first semiconductor die. The cache memory-equipped computing device includes an operation management unit that manages the operation of the first semiconductor die and the second semiconductor die. The paired computing unit and the cache memory at least partially overlap in a plan view. The operation management unit selectively operates either the computing unit or the cache memory that is paired with the computing unit according to the computing strength required of the plurality of computing units. Effect of the Invention

[0008] According to one aspect, in a three-dimensional stacked die in which a computing unit and a cache memory are stacked, the effect of heat generated by the computing unit on the operation of the cache memory can be reduced. [Brief description of the drawings]

[0009] [Figure 1] FIG. 2 is a diagram illustrating an example of a stacked structure of a computing unit and a cache memory. [Diagram 2] FIG. 2 is a side view illustrating an example of a computing device according to an embodiment. [Diagram 3] 3 is a side view showing an example of a three-dimensional stacked die in the computing device shown in FIG. 2. [Figure 4] 3 is a top view illustrating an example of a logic die and a memory die in the computing device illustrated in FIG. 2. [Diagram 5] FIG. 11 is a diagram illustrating an example of the relationship between calculation performance and calculation intensity. [Figure 6] FIG. 13 is a diagram illustrating an example of a process for selectively operating a computing unit and a cache memory. [Figure 7] 3 is a top view showing another example of the logic die and the memory die in the arithmetic device shown in FIG. 2. [Figure 8] FIG. 1 is a circuit diagram of a first embodiment of a computing device. [Figure 9]FIG. 2 is a sequence diagram of a first embodiment of the arithmetic device. [Figure 10] FIG. 11 is a circuit diagram of a second embodiment of the arithmetic device. [Figure 11] FIG. 11 is a sequence diagram of a second embodiment of the arithmetic device. [Figure 12] 12 is a flowchart showing an example of an operating core number adjustment process in FIG. 11. [Figure 13] FIG. 11 is a circuit diagram of a third embodiment of the arithmetic device. [Figure 14] FIG. 13 is a sequence diagram of a calculation device according to a fourth embodiment of the present invention. [Figure 15] 15 is a flowchart showing an example of an operating core number determination process in FIG. 14. [Figure 16] FIG. 13 is a circuit diagram of a fifth embodiment of the arithmetic device. [Figure 17] FIG. 13 is a sequence diagram of a calculation device according to a fifth embodiment of the present invention. [Figure 18] 18 is a flowchart showing an example of a timer process in FIG. 17. [Figure 19] 18 is a flowchart showing a process for ending the timer process in FIG. 17. [Figure 20] FIG. 13 is a top view showing a modification of the logic die and the memory die in the computing device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] [A] Related technologies 1 is a diagram showing an example of a stacked structure of a computing unit 10 (i.e., 10a and 10b in the diagram) and a cache memory 20 (i.e., 20a and 20b in the diagram). The amount of heat generated by the computing unit 10 is greater than the amount of heat generated by the cache memory 20. Therefore, as shown in (a) of the diagram, it is difficult to adopt a three-dimensional stacked die structure in which multiple computing units 10a and 10b are stacked.

[0011] As shown in FIG. 1(b), even in a three-dimensional stacked die in which a computing unit 10a and a cache memory 20b are stacked, a heat dissipation design is required to prevent circuit malfunctions of the cache memory 20b due to temperature rise caused by heat generation from the computing unit 10a.

[0012] In order to reduce the effect of heat generation from the arithmetic unit 10a, a stacked structure is adopted in which the arithmetic unit 10 and the cache memory 20 are arranged so as not to overlap in a plan view, as shown in (c) of Fig. 1. However, avoiding stacking the arithmetic unit 10 and the cache memory 20 as shown in (c) of Fig. 1 places a constraint on increasing the number of arithmetic units 10 and cache memories 20 per unit area of ​​a die. This therefore hinders improvements in the arithmetic performance of the arithmetic device and improvements in the capacity of the cache memory 20.

[0013] In the embodiment of the present invention, even in a stacked structure in which the arithmetic unit 10 and the cache memory 20 are stacked as in FIG. 1(b), the effect of heat generation from the arithmetic unit 10 on the operation of the cache memory can be reduced.

[0014] [B] Embodiment An embodiment will be described below with reference to the drawings. However, the embodiment described below is merely an example, and is not intended to exclude the application of various modified examples and techniques not explicitly stated in the embodiment. In other words, the present embodiment can be modified in various ways without departing from the spirit of the present invention. In addition, each figure does not necessarily include only the components shown in the figure, but may include other functions, etc.

[0015] In the following drawings, the same reference numerals denote the same parts, and therefore the description thereof will be omitted. In this specification, the upper and lower surfaces of the logic die and memory die, which are semiconductor dies, may be parallel to the XY plane. The X-axis direction and the Y-axis direction are directions perpendicular to each other, and the Z-axis direction is a direction perpendicular to the XY plane. In this specification, a plan view means a case where the semiconductor die is viewed in the Z-axis direction.

[0016] 2 is a side view showing an example of the arithmetic device 1 in the embodiment. The arithmetic device 1 includes a logic die 2 and a memory die 3. The arithmetic device 1 is an example of a arithmetic device equipped with a cache memory.

[0017] The logic die 2 and the memory die 3 are semiconductor dies made of silicon or compound semiconductors. The semiconductor dies are sometimes called semiconductor chips. The logic die 2 is an example of a first semiconductor die, and the memory die 3 is an example of a second semiconductor die stacked on the first semiconductor die.

[0018] The logic die 2 and the memory die 3 are stacked by a connecting structure (not shown) to form a three-dimensional stacked die 4.

[0019] The computing device 1 may include an interposer 5. The interposer 5 is an example of a substrate. The interposer 5 electrically connects the three-dimensional stacked die 4 to a printed circuit board (not shown). The interposer 5 may be a silicon interposer or an organic interposer. Instead of the interposer 5, the three-dimensional stacked die 4 may be disposed on another semiconductor substrate or an organic substrate.

[0020] The connection between the logic die 2 and the memory die 3 may be a direct Cu-Cu connection, a connection using a through-silicon stacking (TSV), or a connection using a solder-based microbump technology. Similarly, the connection between the interposer 5 and the three-dimensional stacked die 4 may be made using various connection technologies such as a solder-based microbump technology.

[0021] 2, the interposer 5, the memory die 3, and the logic die 2 are stacked in the Z direction in the order of the interposer 5, the memory die 3, and the logic die 2. In this case, the memory die 3 is stacked adjacent to the interposer 5. However, the computing device 1 is not limited to this case. The interposer 5, the memory die 3, and the logic die 2 may be stacked in the Z direction in the order of the interposer 5, the logic die 2, and the memory die 3. In this case, the logic die 2 is stacked adjacent to the interposer 5.

[0022] Not only the three-dimensional stacked die 4 but also a main memory 6 and other circuits may be arranged on the interposer 5. For example, a die provided with a CPU (Central Processing Unit) may be arranged on the interposer 5.

[0023] Fig. 3 is a side view showing an example of the logic die 2 and the memory die 3 in the computing device 1 shown in Fig. 2. Fig. 4 is a top view showing an example of the logic die 2 and the memory die 3 in the computing device 1 shown in Fig. 2.

[0024] 3 and 4, the logic die 2 includes a plurality of arithmetic units 10. In FIG. 3, a plurality of arithmetic units 10-1 to 10-4 (operating units #1 to #4 in the figure) are shown. In FIG. 4, the logic die 2 includes, as the plurality of arithmetic units 10, Core 0,0 From Core 3,3 The number of the arithmetic units 10 is not limited to this case. The number of the arithmetic units 10 may be two or more.

[0025] As shown in Fig. 4, a plurality of arithmetic units 10 may be formed in an arithmetic unit region 7 of a logic die 2. The arithmetic unit region 7 may be, for example, rectangular. In Fig. 4, there is one arithmetic unit region 7, but unlike the case of Fig. 4, the arithmetic unit region 7 may include a plurality of regions.

[0026] Each computing unit 10 is a processor core (in FIG. 4, “Core i,j(i, j are integers equal to or greater than 0)"). The processor core is a computing unit that functions independently inside the arithmetic device 1 (processor). The processor core may have a logic circuit for interpreting and executing an instruction sequence. Each processor core may include a first level cache memory (L1 cache). The logic die 2 may have a shared circuit shared by multiple computing units 10. The shared circuit may be provided with an interface circuit for inputting and outputting data to and from the outside, and a second level cache memory (L2 cache). However, the internal configuration of the computing unit 10 is not limited to this case. The internal configuration of the computing unit 10 itself is similar to that of a conventional computing unit, so it will not be illustrated and will not be described in detail.

[0027] 3 and 4, the memory die 3 includes a plurality of cache memories 20. In FIG. 3, a plurality of cache memories 20-1 to 20-4 (indicated as caches #1 to #4 in the figure) are shown. In FIG. 4, the memory die 3 includes LLCs as the plurality of cache memories 20. 0,0 From LLC 3,3 The number of cache memories 20 is not limited to this case. The number of cache memories 20 may be two or more.

[0028] As shown in FIG. 4, multiple cache memories 20 may be formed in a memory region 8 of a memory die 3. The memory region 8 may correspond to the arithmetic unit region 7. For example, the memory region 8 may be rectangular. In FIG. 4, there is one memory region 8, but unlike FIG. 4, the memory region 8 may include multiple regions.

[0029] In general, the cache memory in the computing device may include a primary cache (L1 (level 1) cache), a secondary cache (L2 (level 2) cache), a tertiary cache (L3 (level 3) cache), etc., in order of proximity to the computing unit 10. The cache memory closer to the computing unit 10 may be faster and have a smaller capacity, and the read / write speed may be slower and the capacity may be larger as it is farther from the computing unit 10. In one example, the cache memory 20 of this example may be an LLC (Last Level Cache). The LLC may be substantially a tertiary cache (L3 cache). However, depending on the architecture, the LLC may be a quaternary cache (L4 (level 4) cache), etc.

[0030] All elements, such as processor cores, on the logic die 2 and memory die 3 may be able to access the cache memory 20.

[0031] Each cache memory 20 included in the memory die 3 may also be called a memory array. The memory die 3 includes a plurality of memory arrays. Each cache memory 20 is controlled by a memory selection signal (not shown). Each cache memory 20 (memory array) includes a plurality of memory cells (not shown). In one example, each memory cell may be a static random access memory (SRAM) cell. The memory die 3 may have a shared bus wiring (not shown) that transmits signals to each memory cell, and a driver circuit (not shown) connected to the shared bus wiring. The internal configuration of the cache memory 20 itself is similar to that of a conventional cache memory, so it is not illustrated and will not be described in detail.

[0032] Each cache memory 20 is paired with one of the multiple arithmetic units 10. i,j and LLC i,j3, the arithmetic unit 10-1 and the cache memory 20-1 form a pair 30-1. Similarly, the arithmetic units 10-2, 10-3, and 10-4 form pairs 30-2, 30-3, and 30-4 with the cache memories 20-2, 20-3, and 20-4, respectively.

[0033] The pair 30 of the computing unit 10 and the cache memory 20 at least partially overlap in a plan view. As shown in Fig. 3, the computing unit 10-1, the computing units 10-2, 10-3, and 10-4 at least partially overlap with the cache memories 20-1, 20-2, 20-3, and 20-4, respectively, in a plan view.

[0034] A first region occupied by each computing unit 10 in a plan view corresponds to a second region occupied by its counterpart cache memory 20 in a plan view. The first region may be larger than the second region, or may be smaller than the second region, or may have the same area as the second region. One first region may overlap with one second region, and one first region may not overlap with multiple second regions. One second region may overlap with one first region, and one second region may not overlap with multiple first regions.

[0035] The arithmetic unit 10 and the cache memory 20 paired with the arithmetic unit 10 are controlled so that either one of them operates selectively. An operation management unit 31 (see FIG. 6, etc., described later) selectively operates the arithmetic unit 10 and the cache memory 20.

[0036] In Figure 4, the subscripts are the same. i,j and LLC i,j LLC shall not be used in conjunction with, and shall be used exclusively with, i,j When is enabled, Core i,j is disabled (disabled: paused), and the Core i,j When is enabled, LLC i,j will be disabled (disabled: no reading or writing).

[0037] In Fig. 3, elements that selectively operate among the computing unit 10 and cache memory 20 that form the pair 30 are hatched. For example, in the state of Fig. 3, among the computing unit 10-1 and cache memory 20-1 that form the pair 30-1, the computing unit 10-1 operates. Among the computing unit 10-2 and cache memory 20-2 that form the pair 30-2, the computing unit 10-2 operates. Among the computing unit 10-3 and cache memory 20-3 that form the pair 30-3, the cache memory 20-3 operates. Among the computing unit 10-4 and cache memory 20-4 that form the pair 30-4, the cache memory 20-4 operates.

[0038] 3, the arrows show the flow of heat in a schematic manner. The amount of heat generated from the arithmetic unit 10 is greater than the amount of heat generated from the cache memory 20. The thickness of the arrows shows the amount of heat generated in a schematic manner.

[0039] When the computing units 10-1 and 10-2 are operating, the temperatures of the cache memories 20-1 and 20-2 that are closest to them and paired with them 30 may become higher than a predetermined value. However, since the computing device 1 does not use the cache memories 20-1 and 20-2 when the computing units 10-1 and 10-2 are operating, the occurrence of problems with circuit malfunctions in the cache memories 20-1 and 20-2 is suppressed.

[0040] When the cache memories 20-3 and 20-4 are in operation, the nearest arithmetic units 10-3 and 10-4 that are paired with them 30 are unused, and the arithmetic units 10-3 and 10-4 do not generate heat. Therefore, the temperature of the cache memories 20-3 and 20-4 does not reach a predetermined value. As a result, the occurrence of problems such as circuit malfunctions is suppressed.

[0041] The arithmetic device 1 (particularly, an operation management unit 31 in FIG. 6 described later) switches between a pair 30 of a arithmetic unit 10 and a cache memory 20 to be operated. The operation management unit 31 selectively operates either the arithmetic unit 10 or the cache memory 20 paired with the arithmetic unit 10 according to the operation intensity required of the multiple arithmetic units 10.

[0042] Fig. 5 is a diagram showing an example of the relationship between computing performance and computing intensity, in which the horizontal axis indicates computing intensity (Flop / Byte) and the vertical axis indicates computing performance (Flop / s (second)).

[0043] The calculation intensity indicates the number of floating-point calculations executed per byte of data transfer. The calculation intensity corresponds to the workload, which is the magnitude of the load on the calculation device 1. The higher the workload, the higher the calculation intensity. On the other hand, the calculation performance indicates the number of floating-point calculations that can be executed per unit time (1 second).

[0044] The higher the computation intensity of a process, the longer the time required to compute data transferred from memory (i.e., computation time). A high computation intensity region where the computation time is longer than the time required to transfer data to and from memory (i.e., data reading, etc.) (i.e., data transfer time) is called a computation bottleneck region.

[0045] In the computation bottleneck region, the computation performance is limited by the peak computation performance determined by the number and performance of the computing units 10. In the computation bottleneck region (region of high computation intensity), the computation performance is improved by increasing the number of computing units 10, etc., thereby enhancing the overall capacity of the computing units 10.

[0046] On the other hand, the lower the computation intensity of a process, the shorter the computation time. A low computation intensity region where the computation time is equal to or shorter than the data transfer time is called a memory bottleneck region.

[0047] In the memory bottleneck region, the computing performance is limited by the memory bandwidth, etc. Memory bandwidth means the amount of data that can be transferred per second (Byte / s (second)), and is also called memory performance or memory band. Memory bandwidth differs depending on the type of memory. The L1 cache has the highest memory bandwidth, followed by the L2 cache, L3 cache, and main memory 6 (DRAM: Dynamic Random Access Memory), in that order.

[0048] When the cache memory 20 is an L3 cache, if the amount of data used in a calculation exceeds the total capacity of the cache memories 20, data will be read and written to the main memory 6. Therefore, the memory bandwidth, which is the rate limiting factor, becomes lower, and as a result, the calculation performance may decrease.

[0049] In a memory bottleneck region (region of low computational intensity), the total memory capacity of the multiple cache memories 20 can be increased by increasing the number of cache memories 20. As a result, the frequency of reading and writing data in the main memory 6 is reduced, thereby suppressing a decrease in the rate-limiting memory bandwidth. Therefore, a decrease in computational performance is suppressed. In other words, in a memory bottleneck region, the computational performance is improved by increasing the number of cache memories 20.

[0050] 5, the operation intensity is divided into two regions, a memory bottleneck region (region of low operation intensity) and an operation bottleneck region (region of high operation intensity), but the operation intensity may be classified into three or more regions. In one example, the operation intensity may be classified into a low operation intensity region where the operation intensity is less than the first threshold, a middle region where the first threshold ≦ operation intensity ≦ the second threshold, and a high operation intensity region where the second threshold ≦ operation intensity.

[0051] Fig. 6 is a diagram showing an example of a process for selectively operating the arithmetic unit 10 and the cache memory 20. In Fig. 6, the portion marked "Core" indicates that the arithmetic unit 10 is operating, and the portion marked "LLC" indicates that the cache memory 20 is operating.

[0052] The operation management unit 31 manages the operation of the three-dimensional stacked die 4. In other words, the operation management unit 31 manages the operation of the logic die 2 (i.e., the first semiconductor die) and the memory die 3 (i.e., the second semiconductor die). The operation management unit 31 selectively operates either the arithmetic units 10 or the cache memory 20 paired with the arithmetic units 10 according to the arithmetic strength required of the multiple arithmetic units 10.

[0053] The operation management unit 31 may be realized as a function of a CPU (not shown) provided on the logic die 2 or another die, or may be realized by a dedicated circuit provided on at least one of the logic die 2 and the memory die 3.

[0054] 6A is an example of an initial operating state of the three-dimensional stacked die 4. The operation management unit 31 may set the ratio (operation ratio) of operating the computing unit 10 (represented as Core in FIG. 6) at the outer edge of the three-dimensional stacked die 4 in the XY plane to be higher than the operation ratio of the cache memory 20 (represented as LLC in FIG. 6) at the center. The center may be an area within a predetermined distance from the center of gravity on the upper surface of a rectangle parallel to the XY plane of the three-dimensional stacked die 4, and the outer edge may be an area outside the range of the distance.

[0055] The outer edge dissipates heat more easily than the center. Therefore, it is advantageous in terms of heat dissipation design to operate the arithmetic units 10, which generate more heat than the cache memory 20, at the outer edge. However, the distribution and number of the arithmetic units 10 to be operated in the initial operating state on the XY plane are not limited to those shown in FIG. 6(A).

[0056] Fig. 6B is an example of an operating state of the three-dimensional stacked die 4 in a memory bottleneck region. Fig. 6C is an example of an operating state of the three-dimensional stacked die 4 in a computation bottleneck region.

[0057] As shown in Fig. 6(B), in a memory bottleneck region where the calculation intensity (in other words, the workload) is low, the operation management unit 31 reduces the number of arithmetic units 10 in an operating state and increases the number of cache memories 20 in an operating state. As shown in Fig. 6(C), in a calculation bottleneck region where the calculation intensity is high, the operation management unit 31 increases the number of arithmetic units 10 in an operating state and decreases the number of cache memories 20 in an operating state.

[0058] 6B and 6C, the operation ratio of the arithmetic units 10 at the outer edge may be higher than the operation ratio of the arithmetic units 10 at the center. The distribution and number of the arithmetic units 10 to be operated in the memory bottleneck region and the operation bottleneck region on the XY plane are not limited to those in FIG. 6B and 6C.

[0059] 6 shows three states, Fig. 6(A), Fig. 6(B), and Fig. 6(C), but the operation management unit 31 may switch the number of arithmetic units 10 and the number of cache memories 20 in an operating state between two states, or between four or more states. In one example, the operation management unit 31 may switch the number of arithmetic units 10 and the number of cache memories 20 in an operating state between two states, as shown in Fig. 6(B) and Fig. 6(C).

[0060] The operation management unit 31 may selectively operate either one of the arithmetic units 10 and the cache memory 20 paired with the arithmetic unit 10 in accordance with the arithmetic intensity required for the multiple arithmetic units 10. The operation control of the arithmetic units 10 and the cache memory 20 in accordance with the arithmetic intensity may include a case where either the arithmetic unit 10 or the cache memory 20 of the pair 30 is operated based on the detection results such as the usage rate (operation rate) of the arithmetic unit 10 and the cache memory 20. The operation control may also include a case where a list of the arithmetic units 10 to be operated (i.e., operating cores) is designated in advance by a program in accordance with the level of arithmetic intensity predicted due to the processing contents, regardless of the detection results.

[0061] Fig. 7 is a top view showing another example of the logic die 2a and the memory die 3a in the arithmetic device shown in Fig. 2. The logic die 2a and the memory die 3a are not limited to the configuration shown in Fig. 4. The logic die 2a may be provided with not only the arithmetic unit 10 but also another logic control circuit 9a. The logic control circuit 9a may include, for example, an operation management unit 31 and may include a CPU (such as the CPU 32 in Figs. 8, 10, 13, and 16 described later).

[0062] The memory die 3a may be provided with a memory control circuit 9b in addition to the cache memory 20. The memory control circuit 9b may include an LLC control unit (such as an LLC control unit 33 shown in FIGS. 8, 10, 13, 16, etc., described later) that controls the operation of the cache memory 20, and may also include an operation management unit 31 and a CPU.

[0063] 7, logic die 2a may be provided with a plurality of arithmetic unit regions 7a, 7b, 7c, 7d, . . . 7l, etc. Note that arithmetic unit region 7 may be provided in a portion of the XY plane of logic die 2a, and arithmetic unit region 7 does not have to be rectangular.

[0064] In the memory die 3a, a plurality of memory regions 8a, 8b, etc. may be provided as shown in Fig. 7. Note that the memory regions 8a, 8b may be provided in a part of the XY plane of the memory die 3a, and the memory regions 8a, 8b do not have to be rectangular.

[0065] 7, a first area occupied by each computing unit 10 in a plan view may be smaller than a second area occupied by a cache memory 20 that is paired with the computing unit 10 in a plan view. In this case, when the computing units 10 are in operation, the temperature of the cache memory 20 that is paired with the computing unit 10 and is adjacent thereto becomes higher than a predetermined value. However, since the operation management unit 31 does not operate the adjacent cache memory 20 when the computing unit 10 is in operation, the occurrence of problems with circuit malfunctions is suppressed.

[0066] 7, a computing unit 10 that does not have a cache memory 20 as a counterpart 30 may be provided. In one example, the computing unit 10-4 may perform processing in place of the CPU 32 in FIGS. 8, 10, 13, and 16, which will be described later. In this case, the computing unit 10-4 may operate independently of the operation of the cache memory 20.

[0067] As described above, according to the computing device 1 of this embodiment, the operation management unit 31 operates the computing unit 10 and the cache memory 20 according to the computation intensity. In the three-dimensional stacked die 4 in which the computing unit 10 and the cache memory 20 are stacked, the effect of heat generation from the computing unit 10 on the operation of the cache memory 20 can be reduced.

[0068] As a method for operating the arithmetic unit 10 and the cache memory 20 according to the arithmetic intensity, various embodiments are possible, as will be described below.

[0069] [B-1] First embodiment [B-1-1] Configuration 8 is a circuit diagram of a first embodiment of the arithmetic device 1. In the first embodiment, the arithmetic units 10 to be operated (i.e., operating cores) are designated in advance by a program. In other words, in the first embodiment, the operation management unit 31 determines the arithmetic units 10 or cache memories 20 to be operated in the multiple pairs 30 based on a list 22 that specifies in advance the arithmetic units 10 or cache memories 20 to be operated. The user can determine the operating cores in advance through a program.

[0070] The arithmetic device 1 includes a plurality of arithmetic units 10, a plurality of cache memories 20, and an operation management unit 31. The arithmetic device 1 may further include a CPU 32, an LLC control unit 33, and a main memory 6.

[0071] In this example, the computing unit 10 is a Core i,j (i=0, 1, 2, 3, j=0, 1, 2, 3) are provided, and the LLC is used as the cache memory 20. i,j(i=0, 1, 2, 3, j=0, 1, 2, 3) are provided. Note that i and j are not limited to the case in this example and may be any integer.

[0072] The CPU 32 may be provided on the logic die 2 or on another die (not shown). The CPU 32 may execute control of a plurality of arithmetic units 10. The CPU 32 may obtain a list 22 that specifies in advance the arithmetic units 10 or cache memories 20 to be operated, and transmit the list 22 to the operation management unit 31.

[0073] In one example, the multiple arithmetic units 10 may function as a hardware accelerator that serves to increase the arithmetic processing speed. In this case, the CPU 32 may control the hardware accelerator that is composed of the multiple arithmetic units 10. However, at least one of the multiple arithmetic units 10 may control the remaining arithmetic units 10 instead of the CPU 32.

[0074] The operation management unit 31 outputs i,j The output terminal EN i,j The core consists of multiple processors 10 (Core i,j The enable / disable input terminal EN of the EN pin outputs either the enable signal or the disable signal. i,j The number of the arithmetic units 10 corresponds to the number of the arithmetic units 10. Based on information acquired from the CPU 32, the operation management unit 31 may transmit an enable signal to the input terminal EN of the arithmetic unit 10 to be operated, and transmit a disable signal to the input terminal EN of the arithmetic unit 10 to be suspended.

[0075] Output terminal EN of operation management unit 31 i,j The enable signal and the disable signal output from the cache memory 20 (LLC 20) are fed to the cache memory 20 (LLC 20) via a NOT gate circuit 21 (inverter circuit). i,j The NOT gate circuit 21 outputs a state opposite to the input. 0,0When an enable signal is given to the LLC that constitutes pair 30, 0,0 The disable signal is given to Core 0,0 When a disable signal is given to LLC 0,0 The enable signal is given to the other cores. i,j (i=0,1,2,3, j=0,1,2,3) and LLC i,j (i=0,1,2,3, j=0,1,2,3) i,j If an enable signal is given to an LLC with the same subscript, i,j The disable signal is given to Core i,j If a disable signal is given to an LLC with the same subscript, i,j An enable signal is given to.

[0076] When the NOT gate circuit 21 (inverter circuit) is used, the arithmetic device 1 of this embodiment can avoid the circuit configuration from becoming complicated by exclusively operating the arithmetic unit 10 and the cache memory 20 that form a pair 30. However, this is not limited to the case of this embodiment, and the operation management unit 31 may have both an output terminal for the arithmetic unit 10 and an output terminal for the cache memory 20.

[0077] 8 controls a plurality of cache memories 20. Therefore, instead of inputting enable / disable individually to the input terminal EN of each cache memory 20, an enable / disable signal for each cache memory 20 may be input to the input terminal of the LLC control unit 33. In this case, the LLC control unit 33 controls the cache memories 20 (LLC i,j ) has a number of input terminals corresponding to the number of

[0078] The LLC control unit 33 is communicably connected to each arithmetic unit 10, each cache memory 20, the main memory 6, and the CPU 32. The LLC control unit 33 may group multiple cache memories 20 into one set, and associate a value calculated from a memory address in a certain procedure with the set. Data read from the main memory is stored in one of the cache memories 20 included in the set corresponding to the address. In this way, the multiple cache memories 20 may be integrated into one set and operate as one cache memory as a whole. Specifically, the i×j cache memories 20 may operate as one cache memory of an (i×j)-way set associative system.

[0079] However, this is not limited to the above case, and a direct mapping method may be adopted in which each cache memory 20 is uniquely determined from the memory address and multiple cache memories 20 are used individually.

[0080] [B-1-2] Operation 9 is a sequence diagram in the first embodiment of the arithmetic device 1. The CPU 32 acquires the list 22 that specifies in advance the arithmetic units 10 or cache memories 20 to be operated, and passes the list 22 to the operation management unit 31 (step S1).

[0081] The list 22 may be specified by a computer program. The list 22 may include numbers or subscripts as in FIG. 4 that identify the operators 10 to be operated on of pairs 30 as in FIG.

[0082] In one example, two or more lists 22 are prepared, such as a list for a computation bottleneck region (region with high computation intensity) and a list for a memory bottleneck region (region with low computation intensity). In the program, the list 22 to be adopted may be specified according to the content of the computation process.

[0083] The programmer (user) knows the content of the computation process. Therefore, the programmer can predict the location where the computation intensity will be high and the location where the computation intensity will be low in the computation process. In one example, in the location where the computation intensity is predicted to be high, a list 22 of operating cores for a computation bottleneck area (area with high computation intensity) (e.g., an operating core list corresponding to FIG. 6(C)) may be specified in advance in the program. In the location where the computation intensity is predicted to be low, a list 22 for a memory bottleneck area (area with low computation intensity) (e.g., a list corresponding to FIG. 6(B)) may be specified in advance in the program. Note that three or more lists 22 may be prepared in advance according to the computation intensity. In this case, a list 22 may be specified from the three or more lists 22 according to the predicted level of the computation intensity.

[0084] The operation management unit 31 transmits an enable signal to the arithmetic unit 10 (i.e., the operating core) whose number is in the list 22 (step S2). The operation management unit 31 transmits a disable signal to the cache memory 20 that is paired with the arithmetic unit 10 to which the enable signal is to be transmitted (step S3). The operation management unit 31 may use the output of the NOT gate circuit 21 (inverter circuit) as a disable signal to the cache memory 20 by inputting an enable signal to the NOT gate circuit 21.

[0085] The operation management unit 31 transmits a disable signal to the arithmetic unit 10 that does not have a number in the list 22 (i.e., a non-operating core) (step S4). The operation management unit 31 transmits an enable signal to the cache memory 20 that is paired with the arithmetic unit 10 to which the disable signal is to be transmitted (step S5). The operation management unit 31 may use the output of the NOT gate circuit 21 as an enable signal to the cache memory 20 by inputting a disable signal to the NOT gate circuit 21.

[0086] The operation management unit 31 receives a completion notification regarding the change of the operation state from each arithmetic unit 10 and each cache memory 20 (steps S6, S7). When the operation management unit 31 receives the completion notification from each arithmetic unit 10 and each cache memory 20, it transmits the completion notification to the CPU 32 (step S8).

[0087] The CPU 32 instructs the start of arithmetic processing based on the program contents (step S9). The arithmetic unit 10 and the cache memory 20 execute arithmetic processing and data read / write, etc. The CPU 32 receives a notification of the end of the arithmetic processing from the arithmetic unit 10 (step S10).

[0088] The computing device 1 of the first embodiment can reduce the impact of heat generation from the computing unit 10 on the operation of the cache memory 20 in a three-dimensional stacked die 4 in which the computing unit 10 and the cache memory 20 are stacked. The operation management unit 31 determines the computing unit 10 and the cache memory 20 to be selected as targets to be operated according to the operation intensity based on a list 22 that specifies in advance. Therefore, physical measurement of the usage rate of the computing unit 10 related to the operation intensity can be omitted, and the processing time and notification data amount for acquiring and reflecting the measurement results can be reduced. The programmer can grasp the state in which the operation state of a specific computing unit 10 or cache memory 20 is switched.

[0089] [B-2] Second embodiment [B-2-1] Configuration 10 is a circuit diagram of a second embodiment of the arithmetic device 1. The arithmetic device 1 of the second embodiment acquires information on the utilization rates of the multiple arithmetic units 10 and the utilization rate of the memory by the monitor 11 and the monitor 23. The operation management unit 31 operates the arithmetic units 10 or the cache memory 20 according to a comparison result between first information on the utilization rates of the multiple arithmetic units 10 and second information on the utilization rate of the memory. The second embodiment is suitably applied when similar arithmetic processing is repeatedly executed multiple times.

[0090] In the second embodiment, the output of the enable signal and the disable signal from the operation management unit 31 is similar to that in the first embodiment, so in FIG. 10, the indication of the terminals relating to the enable signal and the disable signal is omitted.

[0091] The arithmetic device 1 includes a plurality of arithmetic units 10, a plurality of cache memories 20, an operation management unit 31, a CPU 32, an LLC control unit 33, and a main memory 6, as well as an operating core number adjustment unit 34, a monitor 11, and a monitor 23.

[0092] The CPU 32 may obtain the list 22 of operating cores in the initial state based on a program or the like, and transmit the list 22 to the operation management unit 31. Except for this, the configurations of the multiple arithmetic units 10, the multiple cache memories 20, the CPU 32, the LLC control unit 33, and the main memory 6 are the same as those in the first embodiment.

[0093] The monitor 11 acquires first information 24 related to the usage rates of the multiple arithmetic units 10. The monitor 23 acquires second information 25 related to the usage rates of the multiple cache memories 20. The monitor 11 is an example of a first acquisition unit, and the monitor 23 is an example of a second acquisition unit.

[0094] The usage rate is also called an operation rate. i,j ) per unit time, or the number of instructions waiting to be executed in each arithmetic unit 10. The first information 24 may be the total number of instruction executions per unit time in a plurality of arithmetic units 10, or the total number of instructions waiting to be executed in a plurality of arithmetic units 10. However, the first information 24 is not limited to these cases and may be information relating to the utilization rate of the arithmetic units 10.

[0095] For example, the second information 25 may be a memory usage rate, a cache miss rate, the number of cache misses, or a busy rate in the entire plurality of cache memories 20. The cache miss rate may be a ratio of LLC cache misses to the number of loads and stores. As the memory usage rate increases, the cache miss rate also increases. The second information 25 may be a memory utilization rate in each cache memory 20, etc. However, the second information 25 is not limited to these cases and may be information related to the usage rate of the cache memory 20.

[0096] The operating core number adjustment unit 34 may be one of the functions of the operation management unit 31. The operation management unit 31 and the operating core number adjustment unit 34 may be realized as a function of a CPU 32 provided on the logic die 2 or another die, or may be realized by a dedicated circuit provided on at least one of the logic die 2 and the memory die 3.

[0097] The operating core number adjustment unit 34 acquires the first information 24 and the second information 25 from the monitor 11 and the monitor 23. The operating core number adjustment unit 34 compares the first information 24 and the second information 25. The operating core number adjustment unit 34 adjusts an increase or decrease in the number of arithmetic units 10 (referred to as operating cores) to be operated among the multiple arithmetic units 10 in the multiple pairs 30 according to the comparison result between the first information 24 and the second information 25. Since the arithmetic units 10 and the cache memory 20 operate exclusively, it can also be said that the operating core number adjustment unit 34 adjusts an increase or decrease in the number of cache memories 20 (referred to as operating memories) to be operated among the multiple cache memories 20 in the multiple pairs 30 according to the comparison result.

[0098] Before measurement, the active core number adjustment unit 34 instructs the monitors 11 and 23 to reset.

[0099] The operation management unit 31 selectively operates one of the arithmetic units 10 and the cache memory 20 that form a pair 30, depending on the comparison result between the first information 24 and the second information 25, based on an instruction from the operating core number adjustment unit 34. In other words, the operation management unit 31 increases or decreases the number of arithmetic units 10 (referred to as the number of operating cores) that are operated among the multiple arithmetic units 10 in the multiple pairs 30, depending on the comparison result. Since the arithmetic units 10 and the cache memory 20 operate exclusively, it can also be said that the operation management unit 31 increases or decreases the number of cache memories 20 (referred to as the number of operating memories) that are operated among the multiple cache memories 20 in the multiple pairs 30, depending on the comparison result.

[0100] [B-2-2] Operation Fig. 11 is a sequence diagram in the second embodiment of the arithmetic device 1. The processes in steps S11 to S16 in Fig. 11 relate to designation of the arithmetic units 10 and cache memories 20 to be operated in the initial state. For example, the CPU 32 instructs the operating core number adjustment unit 34 on the number of operating cores so as to achieve the initial state as shown in Fig. 6(A) (step S11).

[0101] The operating core number adjustment unit 34 instructs the operation management unit 31 on the number of operating cores (step S12). The operation management unit 31 creates the list 22 based on the instructed number of operating cores (step S13). For example, a list corresponding to the state of FIG. 6(A) may be prepared in advance when the number of operating cores is 12, a list corresponding to the state of FIG. 6(B) when the number of operating cores is 8, and a list corresponding to the state of FIG. 6(C) when the number of operating cores is 14. Specifically, numbers identifying the operating cores may be provided in advance as the list 22 according to the number of operating cores. The operation management unit 31 may select the list 22 corresponding to the instructed number of operating cores.

[0102] The process of step S14 is the same as the processes of steps S2 to S7 in Fig. 9, and therefore description thereof will be omitted. The operation management unit 31 sends a completion notification to the operating core number adjustment unit 34 (step S15). The operating core number adjustment unit 34 sends a completion notification to the CPU 32 (step S16).

[0103] The CPU 32 instructs the operation management unit 31 to reset the monitors 11 and 23 (step S17). The operation management unit 31 instructs the monitors 11 and 23 to reset (steps S18, S19). This enables the monitors 11 and 23 to newly measure and acquire the first information 24 and the second information 25.

[0104] The CPU 32 starts the arithmetic processing based on the program contents (step S20). The arithmetic unit 10 and the cache memory 20 execute the arithmetic processing and data read / write, etc. The CPU 32 receives a notification of the end of the arithmetic processing from the arithmetic unit 10 (step S21).

[0105] The monitor 11 and the monitor 23 may acquire the first information 24 and the second information 25, respectively, in the calculation process started in step S18. The calculation process (step S20) that is started first after the process for the initial state (steps S11 to S16) is completed is an example of the first calculation process.

[0106] The CPU 32 instructs the operating core number adjustment unit 34 to perform the operating core number adjustment process (step S22). The operating core number adjustment unit 34 executes the operating core number adjustment process (step S23). The operating core number adjustment process will be described later. The operating core number adjustment unit 34 instructs the operation management unit 31 to increase or decrease the number of operating cores (step S24).

[0107] The operation management unit 31 determines the number of operating cores based on the instructed increase or decrease in the number of operating cores, and creates the list 22 based on the number of operating cores (step S25).

[0108] The process of step S26 is the same as the processes of steps S2 to S7 in Fig. 9, and therefore description thereof will be omitted. The operation management unit 31 notifies the operating core number adjustment unit 34 of the completion (step S27). The operating core number adjustment unit 34 notifies the CPU 32 of the completion (step S28).

[0109] The process returns to step S17. A new calculation process is started (step S18). The second calculation process is an example of a second calculation process executed by one of the multiple calculators 10 after the first calculation process. In the second calculation process, the operation management unit 31 may selectively operate either the calculator 10 or the cache memory 20 depending on a comparison result between the first information 24 and the second information 25 acquired in the first calculation process.

[0110] Thereafter, in the repeated arithmetic processing, the previous arithmetic processing may be set as the first arithmetic processing, and the current arithmetic processing may be set as the second arithmetic processing.

[0111] 12 is a flowchart illustrating an example of the operation core number adjustment process in FIG. 11. FIG. 12 may be an example of the process of step S23 in FIG. 11. In the process of FIG. 12, Core i,j For example, let us assume that the cache memory is 20 and the LLC is i,j An explanation will be given using an example.

[0112] The operating core number adjustment unit 34 adjusts the number of operating cores. i,j Monitor 11 to each Core i,j Usage rate of UC ij (Step S100). The operating core number adjustment unit 34 obtains the utilization rate UC of the arithmetic unit 10 by ij The active core number adjustment unit 34 calculates the average of the utilization rates UC ij The sum of multiple usage rates UC ij Alternatively, the top m and bottom n may be deleted and the remaining average value may be calculated.

[0113] The operating core number adjustment unit 34 is an LLC i,j The usage rate UL of the cache memory 20 (LLC) is obtained from the monitor 23 of the LLC control unit 33 that controls the cache memory 20 (LLC) (step S101). The usage rate UL only needs to correspond to UC. i,j may be the average of each LLCi,j may be the sum of multiple utilization LLC i,j Alternatively, the top m and bottom n may be deleted and the average value of the remaining values ​​may be calculated.

[0114] The operating core number adjustment unit 34 determines whether the absolute value of the difference |UC-UL| between the utilization rate UC of the arithmetic unit 10 and the utilization rate UL of the cache memory 20 is less than the threshold value Vth (step S102). If the absolute value of the difference |UC-UL| is less than the threshold value Vth (see the YES route in step S102), the number of operating cores is not increased or decreased, and the process proceeds to step S103. In step S103, the operating core number adjustment unit 34 resets the values ​​of the monitors 11, 23 (step S103).

[0115] On the other hand, if |UC-UL| is equal to or greater than the threshold value Vth (see the NO route in step S102) and the usage rate UL of the cache memory 20 is greater than the usage rate UC of the arithmetic unit 10 (see the YES route in step S104), the process proceeds to step S105. In step S105, the operating core number adjustment unit 34 instructs the operation management unit 31 to decrease the number of operating cores by one. Note that the arithmetic unit 10 (core) and the cache memory 20 that form the pair 30 operate exclusively, so the process in step S105 corresponds to instructing the operation management unit 31 to increase the number of cache memories 20 to be operated by one.

[0116] If |UC-UL| is equal to or greater than the threshold value Vth (see the NO route from step S102) and the usage rate UL of the cache memory 20 is equal to or less than the usage rate UC of the calculator 10 (see the NO route from step S104), the process proceeds to step S106. In step S106, the operating core number adjustment unit 34 instructs the operation management unit 31 to increase the number of operating cores by one. Note that the calculator 10 (core) and cache memory 20 that form a pair 30 operate exclusively, so the process of step S106 corresponds to instructing the operation management unit 31 to decrease the number of cache memories 20 to be operated by one.

[0117] After executing step S105 or step S106, the operating core number adjustment unit 34 resets the values ​​of the monitors 11, 23 (step S103). After the process of step S103, the process ends.

[0118] The computing device 1 of the second embodiment can reduce the impact of heat generated by the computing unit 10 on the operation of the cache memory 20 in a three-dimensional stacked die 4 in which the computing unit 10 and the cache memory 20 are stacked. The computing unit and cache memory to be operated are selected using a comparison result between the utilization rate UC of the computing unit 10 and the utilization rate UL of the cache memory 20 obtained in the previous computing process among the repeated computing processes. In other words, the operation management unit 31 selectively operates a specific computing unit 10 and cache memory 20 so as to achieve the selected number of operating cores in the current computation.

[0119] [B-3] Third embodiment FIG. 13 is a circuit diagram of the third embodiment of the arithmetic device 1. In the arithmetic device 1 of the third embodiment, the monitor 11 acquires the utilization rates of the multiple arithmetic units 10. The operation management unit 31 operates the arithmetic units 10 or the cache memory 20 according to the acquisition result of the utilization rates of the multiple arithmetic units 10. In the third embodiment, the configuration in which the monitor 23 acquires the utilization rates of the multiple cache memories 20 in the second embodiment and the related processing are omitted. The other configurations and processing are the same as those in the second embodiment. Therefore, detailed explanations are omitted. If the utilization rate UL of the arithmetic unit 10 is less than the first threshold, an instruction to decrease the number of operating cores by one may be issued, if UL is equal to or greater than the first threshold and less than the second threshold, the number of operating cores may not be changed, and if UL is equal to or greater than the second threshold, an instruction to increase the number of operating cores by one may be issued. Detailed explanations are omitted.

[0120] Based on the result of the utilization rate of the arithmetic unit 10 acquired by the monitor 11, it is possible to selectively operate a specific arithmetic unit 10 and the cache memory 20 according to the intensity of the calculations.

[0121] [B-4] Fourth embodiment The configuration of the arithmetic unit in the fourth embodiment may be the same as that in the second or third embodiment, and therefore repeated explanation will be omitted.

[0122] 14 is a sequence diagram in the fourth embodiment of the arithmetic device 1. The processes from step S31 to step S36 are similar to the processes from step S11 to step S16 in the second embodiment in FIG.

[0123] In steps S37 to S41 in FIG. 14, the arithmetic process (steps S40, S41) executed as the first arithmetic process is different from that in the second embodiment. In the second embodiment, the first arithmetic process and the second arithmetic process are also arithmetic processes to be processed. On the other hand, in the fourth embodiment, the first arithmetic process is a tuning arithmetic process that is started in steps S40 and S41 in FIG. 14. The tuning arithmetic process is an example of an adjustment arithmetic process that includes fewer instructions than the arithmetic process to be processed. The processes in steps S37 to S39 are similar to those in the second embodiment.

[0124] The CPU 32 instructs the operating core number adjustment unit 34 to perform an operating core number determination process (step S42). The operating core number adjustment unit 34 executes the operating core number determination process (step S43). The operating core number determination process will be described later. The operating core number adjustment unit 34 notifies the operation management unit 31 of the determined number of operating cores (step S44).

[0125] The processing from steps S45 to S48 is similar to the processing from steps S25 to S28 in FIG.

[0126] The CPU 32 starts the main arithmetic processing of the processing target based on the program contents (step S49). The arithmetic unit 10 and the cache memory 20 execute the arithmetic processing and data read / write, etc. The CPU 32 receives a notification of the end of the main arithmetic processing of the processing target from the arithmetic unit 10 (step S50).

[0127] The main calculation process to be performed (steps S49 and S50) is an example of the second calculation process.

[0128] 15 is a flowchart illustrating an example of the process of determining the number of operating cores in FIG. 14. FIG. 15 may be an example of the process of step S43 in FIG. 14. In the process of FIG. 15, Core i,j For example, let us assume that the cache memory is 20 and the LLC is i,j An explanation will be given using an example.

[0129] The processes in steps S110 and S111 are similar to those in steps S100 and S101 in FIG. 12, and therefore will not be described repeatedly.

[0130] If the total number of pairs 30 is P, then active core number adjustment unit 34 calculates the number of active cores by the formula Number of active cores=P×Core utilization rate UC / (Core utilization rate UC+LLC utilization rate) (step S112). Core utilization rate UC / (Core utilization rate UC+LLC utilization rate) means the ratio of arithmetic units 10 that are activated among the multiple arithmetic units 10 in the multiple pairs 30.

[0131] The operation management unit 31 and the number of operating cores adjustment unit 34 create a list 22 of operating cores based on the result of the determination of the number of operating cores by the number of operating cores adjustment unit 34, and selectively operate either the arithmetic unit 10 or the cache memory 20 in the pair 30 based on the list 22. The operation management unit 31 may cooperate with the number of operating cores adjustment unit 34. The operation management unit 31 controls the ratio of arithmetic units 10 to be operated among the multiple arithmetic units 10 in the multiple pairs 30, depending on the result of comparison between first information 24 which is the utilization rate UC of the arithmetic units and second information 25 which is the utilization rate of the cache memory 20.

[0132] The computing device 1 of the fourth embodiment can reduce the impact of heat generation from the computing unit 10 on the operation of the cache memory 20 in a three-dimensional stacked die 4 in which the computing unit 10 and the cache memory 20 are stacked. The operation management unit 31 selects the number of operating cores using a comparison result between first information on the utilization rate UC of the computing unit 10 and second information on the utilization rate UL of the cache memory 20, both of which are acquired in the tuning calculation process. In particular, the operation management unit 31 can control the ratio of the computing units 10 to be operated among the multiple computing units 10 in the multiple pairs 30, even when it is not possible to gradually adjust the number of operating cores to an appropriate number according to the multiple calculation processes.

[0133] [B-5] Fifth embodiment [B-5-1] Configuration 16 is a circuit diagram of a fifth embodiment of the arithmetic device 1. In the arithmetic device 1 of the fifth embodiment, the monitor 11 and the monitor 23 acquire the usage rate of the arithmetic device 10 and the usage rate of the cache memory 20 at predetermined time intervals during a calculation process executed by any one of the multiple arithmetic devices 10.

[0134] The arithmetic device 1 is provided with a switching timer unit 35 in addition to the configuration of the second embodiment shown in Fig. 10 and the third embodiment shown in Fig. 13. The switching timer unit 35 may be a timer constituted by an electronic circuit.

[0135] The switching timer unit 35 may be capable of communicating with the CPU 32 and the number of active cores adjustment unit 34. Upon receiving a switching timer start instruction (i.e., a start command) from the CPU 32, the computing unit 10, or the like, the switching timer unit 35 operates the number of active cores adjustment unit 34 at regular intervals using a timer. The switching timer start instruction may include information about the switching period. The switching period may be set in advance by a user via a program, etc.

[0136] The configuration of the arithmetic device 1 in the fifth embodiment may be similar to that of Fig. 10 and Fig. 13, except for including the switching timer unit 35. Therefore, repeated description will be omitted.

[0137] [B-5-2] Operation 17 is a sequence diagram in the fifth embodiment of the arithmetic device 1. The processes in steps S61 to S65 in FIG. 17 are similar to the processes in steps S31 to S in FIG.

[0138] The CPU 32 transmits a switching timer start instruction to the switching timer unit 35 (step S66). The switching timer unit 35 starts a switching process for switching the number of active cores, etc. at a predetermined time interval (step S67). The switching timer unit 35 starts measuring a period (step S68).

[0139] The switching timer unit 35 instructs the operating core number adjustment unit 34 to execute a monitor reset (step S69). Note that the subsequent steps S70 to S72 are the same processes as steps S17 to S19 in FIG.

[0140] In the fifth embodiment, the usage rate of the arithmetic unit 10 and the usage rate of the cache memory 20 are acquired at predetermined time intervals during the arithmetic processing of the processing target that is started in step S72.

[0141] When a predetermined period of time has elapsed since the start of measurement in step S63 (step S73), switching timer unit 35 notifies number of active cores adjustment unit 34 that the period has elapsed (step S74). The time measured by switching timer unit 35 may be reset.

[0142] The processing in steps S75 to S78 corresponds to the processing in steps S23 to S27 in Fig. 11. Therefore, repeated explanation will be omitted.

[0143] However, in step S77 (step S4), the arithmetic unit 10 (Core i,j ) waits for the currently running process to finish and then goes into a disabled (disabled) state. The computing unit 10 prevents itself from pausing in the middle of a process, thereby reducing the impact on the computing process.

[0144] When operating core number adjustment unit 34 receives a completion notification regarding the increase or decrease in operation of operating cores and operating cache memory (step S78), the receipt of the completion notification is used as a trigger to notify switching timer unit 35 of the start of a timer. In other words, operating core number adjustment unit 34 notifies switching timer unit 35 of the completion notification. Switching timer unit 35 starts measuring a period (step S80).

[0145] The switching timer unit 35 instructs the operating core number adjustment unit 34 to execute a monitor reset (step S81). Steps S82 and S83 are the same processes as steps S18 and S19 in FIG.

[0146] After step S83 is completed, the process returns to step S73. A loop process is executed in which steps S73 to S83 are repeated until the switching timer unit 35 receives a switching timer completion instruction from the CPU 32 (step S84). In the example shown in Fig. 17, when the switching timer unit 35 is notified of the completion of the increase or decrease in the number of operations for the operating cores and the operating cache memories, the switching timer unit 35 starts measurement (step S80).

[0147] The switching timer unit 35 waits for a predetermined period of time corresponding to the switching cycle to elapse (step S73 after returning), and then increases or decreases the number of operating cores and the number of operating memories. Furthermore, when the switching timer unit 35 is notified of the completion of the increase or decrease in the number of operating cores and the number of operating memories (step S79), the switching timer unit 35 starts measuring the period (step S80). The switching timer unit 35 waits for a predetermined period of time corresponding to the switching cycle to elapse (step S73 after returning), and then increases or decreases the number of operating cores and the number of operating memories.

[0148] Fig. 18 is a flowchart showing an example of the timer processing in Fig. 17. Fig. 18 is a flowchart showing an example of the timer processing in steps S66 to S79 ​​in Fig. 17.

[0149] The switching timer unit 35 receives the switching timer start instruction (step S120).

[0150] The switching timer unit 35 starts measuring time by a measurement timer for a period corresponding to the switching cycle (step S121). When the switching timer unit 35 starts measuring time for the period corresponding to the switching cycle (step S121), the switching timer unit 35 notifies the active core number adjustment unit 34 of a monitor value reset instruction (step S122).

[0151] The switching timer unit 35 waits until the switching cycle comes (step S123). When the period corresponding to the switching cycle has elapsed, the switching timer unit 35 notifies the operating core number adjustment unit 34 of the start of the adjustment operation for the number of operating cores (step S124). The value of the measurement period is reset.

[0152] When the switching timer unit 35 receives a completion notification from the active core number adjustment unit 34 (step S125), the switching timer unit 35 newly starts measuring time by the measurement timer for a period corresponding to the switching cycle (step S121).

[0153] Fig. 19 is a flowchart showing a process for ending the timer process in Fig. 17. Fig. 19 is a flowchart showing an example of the processes in steps S84 and S85 in Fig. 17.

[0154] When the switching timer unit 35 receives a switching timer end signal (end instruction) from the CPU 32 (step S130), the switching timer unit 35 ends the process of periodic timing for the period corresponding to the switching cycle (step S131).

[0155] The computing device 1 of the fifth embodiment can reduce the impact of heat generation from the computing unit 10 on the operation of the cache memory 20 in a three-dimensional stacked die 4 in which the computing unit 10 and the cache memory 20 are stacked. The number of operations is adjusted in the computing process to be processed for the computing unit 10 and the cache memory 20 selected according to the computing intensity. In particular, since the number of operations is periodically adjusted using the switching timer unit 35, the operation management unit 31 can selectively operate a specific computing unit 10 and cache memory 20 according to the computing intensity.

[0156] [B-6] Modified version FIG. 20 is a top view showing a modification of the logic die and the memory die in the arithmetic device 1. In FIG.

[0157] For example, in the above description, one arithmetic unit 10 (core) and one cache memory 20 are paired (grouped) 30, but the disclosed technology is not limited to this. As shown in Fig. 20, N arithmetic units 10 (core) and one cache memory 20 may be paired (grouped) 30. Here, N is an integer equal to or greater than 2.

[0158] FIG. 20 shows a case where four (N=4) arithmetic units (Cores) 10a, 10b, 10c, and 10d and one cache memory 20 form one pair 30 (group).

[0159] The N computing units 10a to 10d and one LLC at least partially overlap each other in a plan view. th The operation management unit 31 may operate the cache memory 20 that is paired with the operation units 10a to 10d only when the operation units 10a to 10d are in the off state. thThe case where the number of the arithmetic units 10a to 10d is more than one may be defined as the ON state of the arithmetic units 10a to 10d. When the arithmetic units 10a to 10d are in the ON state, the operation management unit 31 may suspend the cache memory 20 that is the counterpart 30.

[0160] Furthermore, N arithmetic units 10 and M cache memories 20 may constitute one pair 30. Here, N and M are integers equal to or greater than 2. In this case, the number of operations of the M cache memories 20 constituting one pair 30 (group) may be greater than or equal to the threshold m th When the number of cache memories 20 in operation is less than the threshold m, the cache memory 20 is turned off, and only when the cache memories 20 are turned off, the arithmetic unit 10 that is paired with the cache memory 30 is operated. th When the number of cache memories 20 is more than one, the cache memory 20 is turned on, and when the cache memory 20 is in the on state, the arithmetic unit 10 that is the counterpart 30 may be put into a halt.

[0161] When a plurality of arithmetic units 10 form a group, the plurality of arithmetic units 10 may belong to a plurality of groups. Similarly, when a plurality of cache memories 20 form a group, the plurality of cache memories 20 may belong to a plurality of groups.

[0162] [C] Effect According to the arithmetic device 1 in the above-described embodiment, for example, the following advantageous effects can be achieved.

[0163] The arithmetic device 1 includes a logic die 2 including a plurality of arithmetic units 10, and a memory die 3 including a plurality of cache memories 20 that form a pair 30 with any one of the plurality of arithmetic units 10 and that is stacked on the logic die 2. The arithmetic device 1 includes an operation management unit 31 that manages the operation of the logic die 2 and the memory die 3. The arithmetic units 10 and the cache memories 20 that form the pair 30 at least partially overlap each other in a plan view. The operation management unit 31 selectively operates either the arithmetic unit 10 or the cache memory 20 that forms the pair 30 with the arithmetic unit 10 according to the operation strength required of the plurality of arithmetic units 10.

[0164] As a result, in the three-dimensional stacked die 4 in which the arithmetic units 10 and the cache memories 20 are stacked, the effect of heat generated by the arithmetic units 10 on the operation of the cache memories 20 can be reduced.

[0165] In particular, it is possible to stack the computing unit 10 (i.e., core) that performs the computation to be processed and dissipates a large amount of heat, and the cache memory 20. This makes it possible to increase the number of cores and the number of cache memories 20 that can be mounted in a given space, thereby realizing an increase in the capacity of the cache memory 20 in the computing device 1 and an increase in the computing power of the computing unit 10.

[0166] Since the pair 30 of the arithmetic unit and cache memory can be used exclusively, the heat problem in the three-dimensional stacked die 4 can be avoided.

[0167] The operation management unit 31 determines the arithmetic units 10 or cache memories 20 to be operated in a plurality of pairs, based on a list 22 that specifies in advance the arithmetic units 10 or cache memories 20 to be operated.

[0168] This makes it possible to omit physical measurement of the operation intensity, thereby reducing the processing time and notification data volume required to acquire and reflect the measurement results. A programmer can grasp the state in which the operation state of a specific arithmetic unit 10 or cache memory 20 is switched.

[0169] The computing device 1 further includes a monitor 11 that acquires first information 24 on utilization rates UC of the multiple computing elements 10. The operation management unit 31 selectively operates one of the computing elements 10 and the cache memory 20 that form a pair 30, according to the first information 24 on utilization rates UC of the computing elements 10.

[0170] This makes it possible to select an operation that is appropriate for the computation intensity based on the utilization rate actually measured by the monitor 11.

[0171] The arithmetic device 1 further includes a monitor 23 that acquires second information 25 related to utilization rates UL of the plurality of cache memories 20. The operation management unit 31 selectively operates one of the arithmetic unit 10 and the cache memory 20 that form a pair 30, depending on a comparison result between the first information 24 related to the utilization rate UC of the arithmetic unit 10 and the second information 25 related to the utilization rate UL of the cache memory 20.

[0172] This allows selective operation of either the calculator 10 or the cache memory 20, taking into consideration both the state of the calculator 10 and the state of the cache memory 20 as to whether the state is a calculation bottleneck region or a memory bottleneck region.

[0173] In a first arithmetic process executed by any one of the plurality of arithmetic units 10, the monitor 11 and the monitor 23 respectively acquire the first information 24 and the second information 25. In a second arithmetic process executed by any one of the plurality of arithmetic units 10 after the first arithmetic process, the operation management unit 31 selectively operates either the arithmetic unit 10 or the cache memory 20 depending on a comparison result between the first information 24 and the second information 25.

[0174] This makes it possible to select the number of arithmetic units 10 to be operated in the current calculation using the comparison result between the usage rate UC of the arithmetic units 10 and the usage rate UL of the cache memory 20 obtained in the previous calculation.

[0175] The operation management unit 31 controls the ratio of the arithmetic units 10 to be operated among the plurality of arithmetic units 10 in the plurality of pairs 30 according to the comparison result between the first information 24 and the second information 25 .

[0176] Thereby, the ratio of the arithmetic units 10 to be operated is controlled according to the comparison result between the first information 24 and the second information 25. Therefore, the adjustment time is shortened compared to the control in which the number of operating arithmetic units 10 is gradually increased or decreased. Also, the amount of information communication data regarding the measurement results by the monitors 11 and 23 is reduced compared to the control in which the number of operating arithmetic units 10 is gradually increased or decreased.

[0177] In a first arithmetic process executed by one of the plurality of arithmetic units 10, the monitor 11 and the monitor 23 acquire the first information 24 and the second information 25, respectively. In a second arithmetic process executed by one of the plurality of arithmetic units 10 after the first arithmetic process, the operation management unit 31 selectively operates either the arithmetic unit 10 or the cache memory 20 according to a comparison result between the first information 24 and the second information 25.

[0178] As a result, the first information 24 and the second information 25 are acquired in the first arithmetic processing, and in the next second arithmetic processing, either the arithmetic unit 10 or the cache memory 20 can be selectively operated. Therefore, the selective operation of the arithmetic unit 10 and the cache memory 20 can be executed in accordance with the progress of a plurality of arithmetic processing.

[0179] The second arithmetic process is a arithmetic process to be processed, and the first arithmetic process is an adjustment arithmetic process that is executed prior to the arithmetic process to be processed and includes a smaller number of instructions than the second arithmetic process.

[0180] This reduces the effect of the acquisition of the first information 24 and the second information 25 on the arithmetic processing of the processing target, and prevents the performance of the arithmetic processing of the processing mode from being reduced.

[0181] The monitor 11 and the monitor 23 acquire the first information 24 and the second information 25 at predetermined time intervals during the arithmetic processing executed by any of the multiple arithmetic units 10.

[0182] As a result, even when one arithmetic process to be processed continues, it is possible to start obtaining the first information 24 and the second information 25 without waiting for the completion of one arithmetic process.

[0183] The multiple cache memories 20 operate as a single cache memory in total.

[0184] As a result, even when the cache memories 20 are inactive, the entire cache memory operates as a single cache memory, making it easier to control.

[0185] [D] Other The disclosed technology is not limited to the above-described embodiment, and various modifications can be made without departing from the spirit of the present embodiment. Each configuration and each process of the present embodiment can be selected as necessary, or can be combined as appropriate.

[0186] [E] Notes Regarding the above embodiment, the following supplementary notes are further disclosed.

[0187] (Appendix 1) a first semiconductor die including a plurality of computing units; a second semiconductor die including a plurality of cache memories paired with any one of the plurality of arithmetic units and stacked on the first semiconductor die; an operation management unit that manages operations of the first semiconductor die and the second semiconductor die; Equipped with the pair of computing units and the pair of cache memories at least partially overlap each other in a plan view, The operation management unit selectively operates either the computing unit or the cache memory paired with the computing unit according to the computing intensity required of the computing units.

[0188] (Appendix 2) the operation management unit determines the arithmetic unit or the cache memory to be operated in the plurality of pairs based on a list that specifies in advance the arithmetic unit or the cache memory to be operated; 2. A processor incorporating a cache memory as recited in claim 1.

[0189] (Appendix 3) a first acquisition unit that acquires first information regarding a usage rate of the plurality of arithmetic units; the operation management unit selectively operates one of the pair of the arithmetic unit and the cache memory in response to the first information; 3. A processor incorporating a cache memory according to claim 1 or 2.

[0190] (Appendix 4) a second acquisition unit that acquires second information regarding usage rates of the plurality of cache memories; the operation management unit selectively operates one of the pair of the arithmetic unit and the cache memory according to a comparison result between the first information and the second information; 4. A processor equipped with a cache memory as described in claim 3.

[0191] (Appendix 5) the operation management unit controls a ratio of the arithmetic units to be operated among the plurality of arithmetic units in the plurality of pairs according to a comparison result between the first information and the second information. 5. A cache memory-equipped computing device as described in appendix 4.

[0192] (Appendix 6) In a first arithmetic process executed by any one of the plurality of arithmetic units, the first acquisition unit and the second acquisition unit acquire the first information and the second information, respectively; In a second arithmetic operation executed by any one of the plurality of arithmetic units after the first arithmetic operation, the operation management unit selectively operates either one of the arithmetic unit and the cache memory according to a comparison result between the first information and the second information. 5. A cache memory-equipped computing device as described in appendix 4.

[0193] (Appendix 7) the second arithmetic processing is arithmetic processing of a processing target, The first arithmetic processing is an adjustment arithmetic processing that is executed prior to the arithmetic processing to be processed and includes a smaller number of instructions than the second arithmetic processing. 7. A cache memory-equipped computing device as described in appendix 6.

[0194] (Appendix 8) the first acquisition unit and the second acquisition unit acquire the first information and the second information at a predetermined time interval during a calculation process executed by any one of the plurality of calculation units. 5. A cache memory-equipped computing device as described in appendix 4.

[0195] (Appendix 9) The plurality of cache memories operate as a single cache memory as a whole. 2. A processor incorporating a cache memory as recited in claim 1. [Explanation of symbols]

[0196] 1: Arithmetic device 2,2a:Logic die 3,3a:Memory die 4: 3D stacked die 5: Interposer 6: Main memory 7: Arithmetic unit area 8: Memory area 9a: Logic control circuit 9b: Memory control circuit 10: Arithmetic unit 11: Monitor 20: Cache memory 21: NOT gate circuit 22: List 23: Monitor 24 :1st information 25:Second information 30: vs. 31: Operation management section 32 :CPU 33 :LLC control section 34: Operating core number adjustment unit 35: Switching timer section

Claims

1. a first semiconductor die including a plurality of computing units; a second semiconductor die including a plurality of cache memories paired with any one of the plurality of arithmetic units and stacked on the first semiconductor die; an operation management unit that manages operations of the first semiconductor die and the second semiconductor die; Equipped with the pair of computing units and the pair of cache memories at least partially overlap each other in a plan view, the operation management unit selectively operates either one of the plurality of arithmetic units or the cache memory paired with the arithmetic unit according to an arithmetic strength required for the plurality of arithmetic units; A computing device equipped with cache memory.

2. the operation management unit determines the arithmetic unit or the cache memory to be operated in the plurality of pairs based on a list that specifies in advance the arithmetic unit or the cache memory to be operated; 2. A processor incorporating a cache memory according to claim 1.

3. a first acquisition unit that acquires first information regarding a utilization rate of the plurality of arithmetic units; the operation management unit selectively operates one of the pair of the arithmetic unit and the cache memory in response to the first information; 3. A processor incorporating a cache memory according to claim 1.

4. a second acquisition unit that acquires second information regarding usage rates of the plurality of cache memories; the operation management unit selectively operates one of the pair of the arithmetic unit and the cache memory according to a comparison result between the first information and the second information; 4. A processor incorporating a cache memory according to claim 3.

5. the operation management unit controls a ratio of the arithmetic units to be operated among the plurality of arithmetic units in the plurality of pairs according to a comparison result between the first information and the second information.

5. A processor incorporating a cache memory according to claim 4.

6. In a first arithmetic process executed by any one of the plurality of arithmetic units, the first acquisition unit and the second acquisition unit acquire the first information and the second information, respectively; In a second arithmetic process executed by any one of the plurality of arithmetic units after the first arithmetic process, the operation management unit selectively operates either one of the arithmetic unit and the cache memory according to a comparison result between the first information and the second information.

5. A processor incorporating a cache memory according to claim 4.

7. the second arithmetic processing is arithmetic processing of a processing target, The first arithmetic processing is an adjustment arithmetic processing that is executed prior to the arithmetic processing to be processed and includes a smaller number of instructions than the second arithmetic processing.

7. A processor incorporating a cache memory according to claim 6.

8. the first acquisition unit and the second acquisition unit acquire the first information and the second information at a predetermined time interval during a calculation process executed by any one of the plurality of calculation units.

5. A processor incorporating a cache memory according to claim 4.

9. The plurality of cache memories operate as a single cache memory as a whole.

2. A processor incorporating a cache memory according to claim 1.

Citation Information

Patent Citations

  • Computer system, memory device and memory control method based on wafer-on-wafer architecture

    US20230125009A1