Accelerator architecture system and interface core particle

Through independent interface core and accelerator core architecture, the challenges of the accelerator interconnection protocol are solved, larger interconnection and higher bandwidth are achieved, development costs are reduced, interconnection performance and flexibility are improved, and the rapid development model segmentation method is adapted to the rapidly developing model segmentation method.

CN120542371APending Publication Date: 2025-08-26METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510612506.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing accelerator architecture faces challenges to interconnect protocols when expanding the scale of interconnection and compatible with the rapidly developing model segmentation method. The high-speed interconnection interface relies on the accelerator process, affecting computing power performance and development costs.

Method used

It adopts an independent interface core and accelerator core architecture. The interface core includes a variety of high-speed interconnection interfaces. The accelerator core is connected through the core interconnection interface, reducing the area occupied by the interconnection interface on the accelerator chip, and supporting iterative updates of multiple interconnection protocols.

Benefits of technology

It has achieved a rapid development model segmentation without affecting the performance of computing power, reduced development costs, improved interconnection performance and flexibility without dependent on accelerator process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542371A_ABST
    Figure CN120542371A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of chip design, in particular to an accelerator architecture system and an interface core grain, the system comprises an interface core grain and an accelerator core grain, and the interface core grain comprises N first core grain interconnection interfaces and M types of high-speed interconnection interfaces; the mth high-speed interconnection interface comprises K (m) high-speed interconnection interfaces; each first core particle interconnection interface is connected with F high-speed interconnection interfaces; the accelerator core particle comprises Q second core particle interconnection interfaces; wherein the ith first core particle interconnection interface of the interface core particle is connected with the jth second core particle interconnection interface of the accelerator core particle. The area occupied by the core particle interconnection interface is much smaller than that of the high-speed interconnection interface. Due to the introduction of the framework, long-distance high-bandwidth transmission can be ensured, and the area of an accelerator chip can be saved; meanwhile, when the high-speed interconnection interface is iteratively updated, only the interface core particles need to be updated, and the whole accelerator core particles do not need to be updated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chip design, and in particular to an accelerator architecture system and an interface core particle. Background Art

[0002] The three core pillars of large-scale models—computing power, algorithms, and data—have collectively driven the rapid development of AI. These pillars rely on the hardware support of accelerators' computing cores and storage capacity. Especially under the physical limitations of lithography machines, the computing and storage capabilities of a single accelerator are severely limited. Accelerator clusters with a certain interconnected scale are needed to provide greater storage capacity to enable large-scale training or inference applications.

[0003] To achieve larger-scale, higher-bandwidth GPU clusters, it's often necessary to add more interconnect capabilities and bandwidth to the accelerators. This consumes significant accelerator real estate, impacting the accelerator's computing performance. Furthermore, as model algorithms evolve, model partitioning strategies are constantly being updated. These different partitioning strategies place varying demands on the accelerator's cluster scale, interconnection methods, and transmission semantics. These requirements typically manifest themselves in the selection of different high-speed interconnect protocols to meet the requirements of the supernode domain for accelerator scale, switching methods (point-to-point or full switching), interconnection bandwidth, and communication latency. These demands pose significant challenges to the planning and development cycle of accelerator interconnection protocols. Furthermore, achieving higher-bandwidth interconnection scale requires high-speed interconnect interfaces, which often require specific process technology support.

[0004] From this, we can see that how to propose a new accelerator architecture that can expand the interconnection scale to a greater extent without affecting computing power performance, be compatible with the rapidly developing model segmentation methods and interconnection protocols, and realize a high-speed interconnection interface that is independent of the accelerator process, thereby improving the interconnection performance, interconnection flexibility and scalability of the entire system, has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] In response to the above technical problems, the technical solution adopted by the present invention is: an accelerator architecture system, the system comprising: an interface chip, including N first chip interconnection interfaces and M high-speed interconnection interfaces, N≥1, M≥1; the m-th high-speed interconnection interface includes K(m) high-speed interconnection interfaces, K(m)≥1, 1≤m≤M; each first chip interconnection interface is connected to F high-speed interconnection interfaces, 1≤F≤R, R is the maximum number of high-speed interconnection interfaces connected to the current first chip interconnection interface; an accelerator chip, including Q second chip interconnection interfaces, Q≥1; wherein the i-th first chip interconnection interface of the interface chip is connected to the j-th second chip interconnection interface of the accelerator chip, wherein 1≤i≤N, 1≤j≤Q.

[0006] In addition, the present invention also provides an interface core particle, which includes: N first core particle interconnection interfaces, N≥1; M types of high-speed interconnection interfaces, M≥1; the mth type of high-speed interconnection interface includes K(m) high-speed interconnection interfaces, K(m)≥1, 1≤m≤M; wherein each first core particle interconnection interface is connected to F high-speed interconnection interfaces, 1≤F≤R, and R is the maximum number of high-speed interconnection interfaces that the current first core particle interconnection interface can access.

[0007] The present invention has at least the following beneficial effects:

[0008] An embodiment of the present invention provides an accelerator architecture system and an interface core particle. By independently dividing the high-speed interconnect interface into a separate interface core particle and connecting the interface core particle and the accelerator core particle through the core particle interconnect interface, the area occupied by the core particle interconnect interface is much lower than that of the high-speed interconnect interface, and the delay on the core particle interconnect interface is much lower than that of the high-speed interconnect interface, achieving the purpose of both ensuring long-distance transmission and saving accelerator chip area. At the same time, when the high-speed interconnect interface is iteratively updated, only the interface core particle needs to be updated, without updating the entire accelerator core particle, which significantly reduces the accelerator development cost and improves the system upgrade speed. The interface core particle is compatible with the rapidly developing model segmentation method and interconnection protocol and is not dependent on the accelerator process. In addition, the number and types of high-speed interconnect interfaces that the accelerator can realize through the interface core particle are more diverse, making the cluster formed by the accelerator interconnected through the high-speed interconnect interface larger, the cluster's storage capacity is also larger, and the switch forms that can be connected are also more diverse, improving the interconnection performance, interconnection flexibility and scalability of the entire system. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 A schematic diagram of an accelerator architecture system provided by an embodiment of the present invention;

[0011] Figure 2 A schematic structural diagram of an interface core particle provided by an embodiment of the present invention;

[0012] Figure 3 A schematic diagram of a single chip interconnect interface connected to four high-speed interconnect interfaces provided by an embodiment of the present invention;

[0013] Figure 4A schematic diagram of a die interconnection interface including two different package types in the same interface die according to an embodiment of the present invention;

[0014] Figure 5 A schematic diagram of a package layout of a standard or advanced package provided by an embodiment of the present invention;

[0015] Figure 6 A schematic diagram of another advanced packaging layout provided by an embodiment of the present invention;

[0016] Figure 7 A schematic diagram of a packaging layout of a third advanced package provided by an embodiment of the present invention;

[0017] Figure 8 A schematic diagram of a package layout of a standard or advanced package provided by an embodiment of the present invention;

[0018] Figure 9 A schematic diagram of another package layout of a standard or advanced package provided by an embodiment of the present invention;

[0019] Figure 10 A schematic diagram of a package layout of a standard package provided by an embodiment of the present invention;

[0020] Figure 11 A schematic diagram of a package layout for a standard or advanced package of a shared interface chip provided by an embodiment of the present invention;

[0021] Figure 12 A schematic structural diagram of another interface core particle provided by an embodiment of the present invention;

[0022] Figure 13 A schematic diagram of another accelerator architecture system provided by an embodiment of the present invention;

[0023] Figure 14 A schematic diagram of another accelerator architecture system provided by an embodiment of the present invention;

[0024] Figure 15 A schematic diagram of forming a heterogeneous cluster through switches according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meanings as commonly understood by those skilled in the art.

[0027] Example 1

[0028] See also Figure 1 An embodiment of the present invention provides an accelerator architecture system, which includes: at least one interface core particle and at least one accelerator core particle.

[0029] It should be noted that the interface core and the accelerator core are two independent core dies. The interface core is specifically responsible for expanding the high-speed interconnect interface. The accelerator core is a core specifically used for data processing. This is different from the traditional manufacturing method of integrating the high-speed interconnect interface on the accelerator core. The use of this independent core approach can better perform iterative updates of the high-speed interconnect interface. Because the high-speed interconnect interface is traditionally integrated on the accelerator core, once the high-speed interconnect interface is iteratively updated with new technology, the entire accelerator core may need to be re-taped, which is costly and time-consuming. The independent interface core provided in the embodiment of the present invention can effectively solve the problem of iterative updates.

[0030] In one embodiment, the accelerator is a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), or an AI processor (Artificial Intelligence Processor). Other types of accelerators that are specifically designed to improve the performance of specific computing tasks and work in conjunction with a general-purpose CPU also fall within the scope of protection of the present invention. In one embodiment, the AI ​​processor is a tensor processing unit (TPU) or a neural network processor (NPU), etc. Other types of AI processors also fall within the scope of protection of the present invention.

[0031] For further information, see Figure 2 , the interface core particle includes N first core particle interconnection interfaces and M types of high-speed interconnection interfaces, N≥1, M≥1; the mth type of high-speed interconnection interface includes K(m) high-speed interconnection interfaces, K(m)≥1, 1≤m≤M; each first core particle interconnection interface is connected to F high-speed interconnection interfaces, 1≤F≤R, R is the maximum number of high-speed interconnection interfaces that the current first core particle interconnection interface can access. It should be noted that in order to more clearly show the interconnection relationship between a single core particle interconnection interface and a high-speed interconnection interface, Figure 2Only one of the chip interconnection interfaces is shown to be interconnected with F high-speed interconnection interfaces, and the interconnection relationship between the other chip interconnection interfaces and the high-speed interconnection interfaces is not shown.

[0032] In one embodiment, the i-th first chip interconnect interface of the interface chip is connected using a short-distance high-density transmission protocol or a short-distance transmission protocol, 1≤i≤N. It should be noted that when the first chip interconnect interface and the second chip interconnect interface adopt advanced packaging scenarios, a short-distance high-density transmission protocol is adopted. When the first chip interconnect interface and the second chip interconnect interface adopt standard packaging scenarios, a short-distance transmission protocol is adopted. The interface chip can adapt to the rapidly developing interconnect interface protocol, and the adapted protocol is flexible.

[0033] In one embodiment, the short-distance, high-density transmission protocol is UCIe (Universal Chiplet Interconnect Express), BOW (Bunch of Wires), HBI (High Bandwidth Interface), etc. Other types of short-distance, high-density transmission protocols also fall within the scope of protection of the present invention, for example, the transmission distance is less than or equal to 50 mm, the unidirectional transmission bandwidth per unit length is greater than or equal to 32 Gb / s / mm, or the unidirectional transmission bandwidth per unit area is greater than or equal to 16 GB / s / mm 2 transmission protocol.

[0034] In one embodiment, the short-distance transmission protocol may also be UCIe (Universal Chiplet Interconnect Express). Other types of short-distance transmission protocols also fall within the protection scope of the present invention.

[0035] Among them, since the core-grain interconnect interface is aimed at short-distance transmission and covers a transmission distance of millimeters, the hardware complexity of its physical layer and protocol layer is much lower than that of the high-speed interconnect interface for long-distance transmission, and the chip area occupied by the core-grain interconnect interface is also much smaller than that of the high-speed interconnect interface. The transmission distance covered by the high-speed interconnect interface is from several centimeters to several meters, and its hardware includes one or more modules such as parallel-to-serial conversion, equalizer, differential pair signal transmission, and forward error correction (FEC). By integrating the core-grain interconnect interface in the accelerator core, the chip area occupied by the interconnection is reduced, and the chip area originally occupied by the high-speed interconnect interface is released to allow for more integrated computing units, which can improve the computing power of the accelerator.

[0036] The high-speed interconnection interface supports direct interconnection with the high-speed interconnection interfaces of other accelerator core systems through physical connection lines, and also supports indirect interconnection with the high-speed interconnection interfaces of other accelerator core systems by accessing switches of different protocol types, such as Ethernet switches, UALink switches, PCIe switches, and CXL switches.

[0037] In one embodiment, the high-speed interconnection interface adopts the Ethernet Roce protocol (RDMA over Converged Ethernet), the UAlink protocol (Ultra Accelerator Link), the PCIe protocol (Peripheral Component Interconnect Express), the UEC protocol (Ultra Ethernet Consortium), the OSIA protocol (Open System Interconnection Architecture) and other protocols that may be used for the Internet of Things, communications or data exchange (collectively referred to as XLink protocols). High-speed interconnection interfaces using other protocols also fall within the scope of protection of the present invention.

[0038] The types of high-speed interconnect interfaces are categorized by the type of protocol used. High-speed interconnect interfaces that use the same protocol are classified as the same type of high-speed interconnect interface, while high-speed interconnect interfaces that use different protocols are classified as different types of high-speed interconnect interfaces. For example, if one high-speed interconnect interface uses the RoCe protocol and the other uses the PCIe protocol, the high-speed interconnect interface that uses the RoCe protocol is one type of high-speed interconnect interface, while the high-speed interconnect interface that uses the PCIe protocol is another type of high-speed interconnect interface.

[0039] In one embodiment, the value of R satisfies the bandwidth constraint: the bandwidth of the first chiplet interconnect interface is greater than the bandwidth of F high-speed interconnect interfaces; wherein, when the i-th first chiplet interconnect interface is connected to F i When there are high-speed interconnection interfaces, the following conditions must be met: i *dw i >∑ j=1 Fi (hs i,j *hw i,j ), where hs i,j is the number of channels of the jth high-speed interconnect interface connected to the i-th chiplet interconnect interface, hw i,j is the bandwidth of a single channel in the jth high-speed interconnect interface connected to the i-th chiplet interconnect interface; ds i is the number of channels of the i-th chiplet interconnection interface, dwi is the bandwidth of a single channel in the i-th chiplet interconnect interface.

[0040] As an example, see Figure 3 , F i =4. Each interface chiplet includes the i-th first chiplet interconnect interface and four high-speed interconnect interfaces. The i-th first chiplet interconnect interface has 64 channels, a single-channel bandwidth of 32 Gbps, and a total bandwidth of 2048 Gbps. Each high-speed interconnect interface has four channels, a single-channel bandwidth of 112 Gbps, and a total bandwidth of 1792 Gbps. The total bandwidth of 2048 Gbps for the i-th first chiplet interconnect interface is greater than the total bandwidth of 1792 Gbps for the four high-speed interconnect interfaces.

[0041] As an example, F i =2. Each interface chiplet includes the i-th first chiplet interconnect interface and two high-speed interconnect interfaces. The i-th first chiplet interconnect interface has 32 channels, a single-channel bandwidth of 32 Gbps, and a total bandwidth of 1024 Gbps. Each high-speed interconnect interface has 4 channels, a single-channel bandwidth of 112 Gbps, and a total bandwidth of 896 Gbps. The total bandwidth of 1024 Gbps for the i-th first chiplet interconnect interface is greater than the total bandwidth of 896 Gbps for the four high-speed interconnect interfaces.

[0042] In one embodiment, the interface core comprises the i-th core interconnect interface and the t-th core interconnect interface; wherein, ds i and ds t Different, ds t is the number of channels of the t-th first chiplet interconnection interface, where 1≤t≤N and i≠t; and / or, dw i and dw t Different, dw t is the bandwidth of a single channel in the t-th first chiplet interconnection interface.

[0043] It should be noted that within the same interface chiplet, the number of channels and bandwidth of the chiplet interconnect interface and the high-speed interconnect interface can be flexibly configured. Specifically, the number of channels in different first chiplet interconnect interfaces can be the same or different. The bandwidth of a single channel in different first chiplet interconnect interfaces can be the same or different. The number of channels in different high-speed interconnect interfaces can be the same or different. The bandwidth of a single channel in different high-speed interconnect interfaces can be the same or different.

[0044] In one embodiment, the j-th high-speed interconnect interface is connected to at least one first chip interconnect interface.

[0045] In one embodiment, the F high-speed interconnect interfaces include at least one type of high-speed interconnect interface. It should be noted that not only can the same interface chip include M types of high-speed interconnect interfaces, but the F high-speed interconnect interfaces connected to the same chip interconnect interface can also be of different types. For example, when F = 2, the i-th first chip interconnect interface can be connected to two high-speed interconnect interfaces of the same type, or it can be connected to two high-speed interconnect interfaces of different types.

[0046] In one embodiment, the interface core particle includes N1 first core particle interconnection interfaces using advanced packaging and N2 first core particle interconnection interfaces using standard packaging, N1 ≥ 0, N2 ≥ 0, and N1 + N2 = N. The use of other types of packaging technologies applicable to the core particle level also falls within the scope of protection of the present invention. That is, the first core particle interconnection interfaces in the interface core particle can be all using advanced packaging or standard packaging, or can be compatible with two or more packaging forms. As an example, please refer to Figure 4 ,exist Figure 4 FIG shows a die interconnect interface including two different package types in the same interface die, such as Figure 4 As shown, one interface chip includes one first chip interconnect interface using advanced packaging and two first chip interconnect interfaces using standard packaging, and shows a connection relationship between the chip interconnect interface and the high-speed interconnect interface.

[0047] Furthermore, the accelerator core comprises Q second core interconnection interfaces, where Q≥1.

[0048] It should be noted that the definition and implementation of the first chiplet interconnection interface integrated in the interface chiplet are applicable to the second chiplet interconnection interface integrated in the accelerator chiplet, and will not be repeated here.

[0049] In one embodiment, the accelerator core does not integrate a high-speed interconnect interface. Instead, the core primarily includes a computing unit and a storage unit. Compared to accelerators with integrated high-speed interconnect interfaces, this significantly reduces the chip area occupied by the high-speed interconnect interface. The freed-up chip area can be used to design more computing units, thereby increasing the chip's computing power.

[0050] Furthermore, the i-th first chiplet interconnection interface of the interface chiplet and the j-th second chiplet interconnection interface of the accelerator chiplet are interconnected, where 1≤i≤N, 1≤j≤Q.

[0051] It should be noted that the core interconnect interface integrated in the accelerator occupies a smaller area, has a higher access bandwidth, and can access different types of high-speed interconnect interfaces. In other words, the accelerator can access a greater number and types of high-speed interconnect interfaces through the interface core, making the cluster formed by the interconnection between accelerators larger, and thus the storage capacity on the cluster larger.

[0052] In one embodiment, the connection between the i-th first chiplet interconnect interface and the j-th second chiplet interconnect interface is one of the following: physical layer interconnection based on a silicon dielectric layer; or physical layer interconnection based on an organic substrate or a printed circuit board. Other types of connection methods applicable to the chiplet level also fall within the scope of protection of the present invention.

[0053] The physical layer interconnection between the i-th first chiplet interconnect interface and the j-th second chiplet interconnect interface is achieved using a silicon dielectric layer. It should be noted that the silicon dielectric layer enables high-density interconnection and high transmission speeds. Compared to the latency of high-speed interconnect interfaces or large-scale clusters, the latency introduced by the chiplet interconnect interface is very small, generally on the order of ~10ns. In an advanced packaging-based chiplet interconnect implemented using the UCIe protocol, the driving distance of the chiplet interconnect interface is approximately 2mm.

[0054] The physical layer interconnection between the i-th first chiplet interconnection interface and the j-th second chiplet interconnection interface is based on an organic substrate or printed circuit board. Similarly, high-density interconnection can be achieved based on an organic substrate or printed circuit board, with high transmission speed. Compared to the latency on high-speed interconnection interfaces or the latency on large-scale clusters, the latency introduced by the chiplet interconnection interface is also very small, generally on the order of ~10ns. Under a standard package-based chiplet interconnection implemented by the UCIe protocol, the driving distance of the chiplet interconnection interface is approximately within 30mm.

[0055] In one embodiment, the i-th first chiplet interconnection interface of the interface chiplet and the j-th second chiplet interconnection interface of the accelerator chiplet are both connected using a short-distance high-density transmission protocol or a short-distance transmission protocol.

[0056] In one embodiment, when the accelerator core adopts advanced packaging, the interface core is interconnected with the second core interconnect interface of the accelerator core via a first core interconnect interface of a standard or advanced packaging. In another embodiment, when the accelerator core adopts standard packaging, the interface core is interconnected with the second core interconnect interface of the accelerator core via a first core interconnect interface of a standard packaging.

[0057] In one embodiment, the advanced packaging connects the interface chip to the accelerator chip through three to five silicon dielectric layers. The first chip interconnect interface of the interface chip contacts the silicon dielectric layer via small bumps with a spacing of 45-55 μm, and interconnection is achieved through metal traces on the silicon dielectric. In the accelerator chip's advanced packaging configuration, when the accelerator chip is configured with video memory, the interface chip is located on the same side as the video memory chip, either between two memory chips or at one end.

[0058] In one embodiment, for a standard package, the interface chip's first chip interconnect interface contacts the organic substrate via small bumps with a pitch of 100-150 μm, and then connects to the accelerator chip's second interconnect interface via physical layer traces on the substrate. The interface chip is located on the opposite side of the memory chip.

[0059] In one embodiment, the standard package interconnects the interface chip and the accelerator chip via an organic substrate. The first chip interconnect interface contacts the organic substrate via bumps spaced 100 to 150 μm apart, and interconnection is achieved via metal traces on the organic substrate. When the accelerator chip is in a standard package, the location of the second chip interconnect interface within the accelerator chip determines the optional location of the interface chip. The interface chip can be connected to one or more locations on the long or short sides of the accelerator; a single interface chip can also be connected to two accelerator chips simultaneously.

[0060] In one embodiment, when the interface chip and the video memory are located on the same side of the accelerator chip, the interface chip, the video memory, and the accelerator chip use the same packaging method. When the interface chip and the video memory are located on different sides of the accelerator chip, the packaging method of the video memory and the accelerator chip is the same as or different from the packaging method of the interface chip. In one embodiment, when the interface chip and the video memory are located on different sides of the accelerator chip, the accelerator chip, the video memory, and the interface chip all use standard packaging; or the video memory and the accelerator chip use advanced packaging, while the interface chip uses advanced packaging or standard packaging. As an example, when the interface chip and the video memory are located on the same side of the accelerator chip, if the video memory is HBM using advanced packaging, the interface chip also uses advanced packaging; if the video memory is DDR using standard packaging, the interface chip also uses standard packaging. As another example, when the interface chip and the video memory are located on different sides of the accelerator chip, if the video memory is HBM using advanced packaging, the interface chip also uses advanced packaging or standard packaging.

[0061] By distributing the accelerator chips and the interface chips, various types of packaging layouts can be formed. The packaging layouts of the accelerator chips and the interface chips include at least the following four types.

[0062] In one embodiment, when the accelerator core is configured with video memory on one side and U interface cores are configured on the other side, where U ≥ 1, and the accelerator core and video memory are packaged in advanced packaging, each interface core is interconnected with the second core interconnect interface of the accelerator core via a first core interconnect interface in standard or advanced packaging. As an example, see Figure 5 , U=2, two video memories are configured on one side of the accelerator core particle, and two interface core particles are configured on the other side. The first core particle interconnection interface of each interface core particle using standard or advanced packaging is interconnected with the second core particle interconnection interface of the accelerator core particle, and a filler chip (dummy die) is also configured between the two core particles. It should be noted that the position of the video memory controller in the accelerator core particle determines the packaging position of the video memory, and the position of the second core particle interconnection interface determines the optional packaging position of the interface core particle. The size of the interface core particle determines whether a filler chip is required between different core particles. The filler chip is introduced to solve the packaging stress problem, and the cost is generally negligible.

[0063] In one embodiment, when X memory and Y interface chiplets are configured on the same side of the accelerator chiplet, where X≥1, Y≥1, and the accelerator chiplet and memory are packaged in advanced packaging, each interface chiplet is interconnected with the second chiplet interconnection interface of the accelerator chiplet via the first chiplet interconnection interface in the advanced packaging. As an example, see Figure 6 , two video memories are configured on one side of the accelerator chiplet, an interface chiplet is configured between the two video memories on the same side, and the first chiplet interconnect interface of each interface chiplet using advanced packaging is interconnected with the second chiplet interconnect interface of the accelerator chiplet. As another example, two video memories are configured on both sides of the accelerator chiplet, an interface chiplet is configured between the two video memories on the same side, and the first chiplet interconnect interface of each interface chiplet using advanced packaging is interconnected with the second chiplet interconnect interface of the accelerator chiplet. As an example, please refer to Figure 7 The system includes two accelerator cores arranged side by side. Each accelerator core is configured with two memory chips and two interface chips on the same side of the two accelerator cores. The two accelerator cores share a single interface chip, located at either end of each accelerator's two memory chips. Each interface chip is interconnected with the accelerator core's second chip interconnect interface using a first chip interconnect interface in an advanced package.

[0064] In one embodiment, when L interface chips are configured on one side of the accelerator chip or L interface chips are configured on both sides, L ≥ 1, and the accelerator chip adopts advanced packaging, each interface chip is interconnected with a second chip interconnect interface of the accelerator chip via a first chip interconnect interface adopting standard or advanced packaging. As an example, please refer to Figure 8 and Figure 9 , Figure 8 Two interface cores are configured on one side (short side) of the accelerator. Figure 9 It is shown that two interface chips are configured on the other side (long side) of the accelerator, and the first chip interconnection interface of the interface chip and the second chip interconnection interface of the accelerator chip adopt standard or advanced packaging interconnection.

[0065] In one embodiment, when two accelerator cores share L interface cores, L ≥ 1, and the accelerator cores use standard packaging, each interface core includes a first first core interconnect interface and a second first core interconnect interface using standard packaging, wherein the first first core interconnect interface is interconnected with the second core interconnect interface of the first accelerator core, and the second first core interconnect interface is interconnected with the second chip interconnect interface of the second accelerator core. As an example, please refer to Figure 10 , Figure 10 The system includes two groups of accelerator core particles, which share a group of interface core particles. The two accelerators and the two interface core particles all use standard packaging.

[0066] In another embodiment, when two accelerator cores share L interface cores, L ≥ 1, and the accelerator cores use advanced packaging, each interface core includes a first first core interconnect interface and a second first core interconnect interface using advanced or standard packaging, wherein the first first core interconnect interface is interconnected with the second core interconnect interface of the first accelerator core, and the second first core interconnect interface is interconnected with the second chip interconnect interface of the second accelerator core. As an example, please refer to Figure 11 , Figure 11 It includes two groups of accelerator core particles, which share a set of interface core particles. Each accelerator core particle and its surrounding video memory adopt advanced packaging, while the two interface core particles can adopt either advanced packaging or standard packaging.

[0067] It should be noted that the same accelerator does not need to be re-produced. By simply packaging interface cores that support different protocols or using different packaging strategies, it can simultaneously support multiple types of high-speed interconnect interfaces through the same or different interface cores. For example, the same accelerator can be packaged with interface cores that use the first type of high-speed interconnect interface, with interface cores that use the second type of high-speed interconnect interface, or with interface cores that use both the first and second types of high-speed interconnect interfaces.

[0068] As an example, see Figure 12 ,exist Figure 12The interface core includes a first high-speed interconnect interface, a second high-speed interconnect interface, a third high-speed interconnect interface, and a fourth high-speed interconnect interface, omitting the connection relationship between the core interconnect interface and the high-speed interconnect interface. The first, second, third, and fourth high-speed interconnect interfaces can all use the PCIe protocol or the UAlink protocol. Alternatively, the first and second high-speed interconnect interfaces can use the PCIe protocol, while the third and fourth high-speed interconnect interfaces can use the UAlink protocol. Please refer to Figure 13 and Figure 14 , when Figure 12 The interface between the core particle and the accelerator may form Figure 13 and Figure 14 The architecture system. Figure 13 Without changing the accelerator core, it can be integrated with the interface core of the high-speed interconnect interface that adopts the PCIe protocol, or with the interface core of the high-speed interconnect interface that adopts the UAlink protocol, or with the interface core of the high-speed interconnect interface that adopts both types of protocols. Figure 14 The two interface cores can both be interface cores for high-speed interconnect interfaces using the PCIe protocol, or they can both be interface cores for high-speed interconnect interfaces using the UAlink protocol. Alternatively, one interface core can be an interface core for a high-speed interconnect interface using the PCIe protocol and the other interface core can be an interface core for a high-speed interconnect interface using the UAlink protocol. Alternatively, both interface cores can be interface cores for high-speed interconnect interfaces using two different types of protocols. An accelerator core equipped with interface cores of different protocols implements an accelerator system using different high-speed interconnect protocols.

[0069] In one embodiment, the system further comprises M types of switches, M≥1; each switch is connected to a corresponding type of high-speed interconnect interface in the interface core particle to implement a homogeneous or heterogeneous cluster. Different switch systems are constructed by replacing interface core particles based on different protocols. Interface core particles can quickly adapt to the accelerator and switch system cluster construction with specific high-speed interconnect protocol requirements. As an example, please refer to Figure 15 All switches use the PCIe protocol, all use the UALink protocol, or some use the PCIe protocol and some use the UALink protocol. The high-speed interconnect interfaces connected to each type of switch use matching protocols.

[0070] In one embodiment, the accelerator core particles and the interface core particles that are packaged together have the same or different processes, and different processes are connected to achieve docking. Among them, the process is a measure of the chip manufacturing process precision, which is the minimum feature size that can be achieved in the chip manufacturing process. The smaller the value, the higher the integration and the lower the power consumption. For example, the process of the packaged core particles is the same: a 7nm accelerator core particle can be packaged together with a 7nm interface core particle. The process of the packaged core particles is different: a 7nm accelerator core particle can be packaged together with a 28nm interface core particle, or a 28nm accelerator core particle can be packaged together with a 7nm interface core particle. In traditional accelerators with integrated high-speed interconnect interfaces, the rate limit of its high-speed interconnect interface is limited by the process technology of the accelerator. The interface core particles provided by the present invention can make the rate limit of the high-speed interconnect interface no longer limited by the process technology of the accelerator. The interface core particles can adopt more advanced processes to greatly improve the rate of the high-speed interconnect interface. Furthermore, under the illumination limits of a lithography machine, the process used by traditional accelerators with integrated high-speed interconnect interfaces can only achieve computing power A and high-speed interconnect bandwidth B. By using the interface chiplet based on a more advanced process technology provided by the present invention, under the illumination limits of the same lithography machine, the accelerator chiplet combined with the interface chiplet can achieve a computing power of a*A and an interconnect bandwidth of b*B, where a>1 and b>1, significantly improving the performance of the accelerator system. In other words, while maintaining the same computing power A and high-speed interconnect bandwidth B, the area of ​​the GPU can be significantly reduced, thereby improving the yield and reducing the cost of the accelerator.

[0071] In summary, an embodiment of the present invention provides a GPU chip architecture system, which includes an interface chip and an accelerator chip. The interface chip includes a first chip interconnection interface and a high-speed interconnection interface. The accelerator chip includes a second chip interconnection interface. The interface chip and the accelerator are interconnected through the first chip interconnection interface and the second chip interconnection interface. The area occupied by the chip interconnection interface is much smaller than that of the high-speed interconnection interface, and the delay introduced by the chip interconnection interface is much smaller than that of the entire accelerator chip cluster system, thereby achieving the purpose of ensuring long-distance transmission while saving chip area and taking into account the rapid iterative update of the high-speed interconnection interface.

[0072] Based on the same inventive concept as that of the first embodiment, the present invention further provides a second embodiment.

[0073] Example 2

[0074] Please refer again Figure 2An embodiment of the present invention provides an interface chip, comprising: N first chip interconnect interfaces and M types of high-speed interconnect interfaces. Where N ≥ 1, M ≥ 1; the mth type of high-speed interconnect interface comprises K(m) high-speed interconnect interfaces, K(m) ≥ 1, 1 ≤ m ≤ M; each first chip interconnect interface connects to F high-speed interconnect interfaces, 1 ≤ F ≤ R, where R is the maximum number of high-speed interconnect interfaces that can be connected to the current first chip interconnect interface.

[0075] It should be noted that the relevant definitions and implementation methods of the interface core particles in Example 1 are completely applicable to Example 2 and will not be repeated here.

[0076] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0077] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. An accelerator architecture system, characterized in that: The system includes at least one interface core particle and at least one accelerator core particle; The interface chip includes N first chip interconnection interfaces and M types of high-speed interconnection interfaces, N≥1, M≥1; the mth type of high-speed interconnection interface includes K(m) high-speed interconnection interfaces, K(m)≥1, 1≤m≤M; each first chip interconnection interface is connected to F high-speed interconnection interfaces, 1≤F≤R, R is the maximum number of high-speed interconnection interfaces that can be connected to the current first chip interconnection interface; The accelerator core comprises Q second core interconnection interfaces, where Q≥1; The i-th first chip interconnection interface of the interface chip is connected to the j-th second chip interconnection interface of the accelerator chip, where 1≤i≤N, 1≤j≤Q.

2. The system according to claim 1, wherein: The i-th first chip interconnection interface of the interface chip and the j-th second chip interconnection interface of the accelerator chip are connected using a short-distance high-density transmission protocol or a short-distance transmission protocol.

3. The system according to claim 2, characterized in that The connection mode between the i-th first chip interconnection interface and the j-th second chip interconnection interface is one of the following connection modes: physical layer interconnection based on a silicon dielectric layer; physical layer interconnection based on an organic substrate or a printed circuit board.

4. The system according to claim 1, wherein: When the accelerator chip adopts advanced packaging, the interface chip is interconnected with the second chip interconnect interface of the accelerator chip through a first chip interconnect interface adopting advanced packaging or standard packaging; Alternatively, when the accelerator chip adopts a standard package, the interface chip is interconnected with the second chip interconnection interface of the accelerator chip through a first chip interconnection interface adopting the standard package.

5. The system according to claim 1, wherein: When the interface chip and the video memory are located on the same side of the accelerator chip, the interface chip, the video memory and the accelerator chip use the same packaging method; when the interface chip and the video memory are located on different sides of the accelerator chip, the packaging method of the video memory and the accelerator chip is the same as or different from the packaging method of the interface chip.

6. The system according to claim 1, wherein: The packaging forms of the accelerator core particle and the interface core particle include at least the following four types: When one side of the accelerator chip is configured with video memory and the other side is configured with U interface chiplets, U ≥ 1, and the accelerator chiplet and video memory adopt advanced packaging, each interface chiplet is interconnected with the second chiplet interconnection interface of the accelerator chiplet through a first chiplet interconnection interface adopting advanced or standard packaging; When X video memories and Y interface chiplets are configured on the same side of the accelerator chiplet, where X≥1 and Y≥1, and the accelerator chiplet and video memories are packaged in advanced packaging, each interface chiplet is interconnected with a second chiplet interconnection interface of the accelerator chiplet via a first chiplet interconnection interface in the advanced packaging; When L interface chiplets are configured on one side of the accelerator chiplet or L interface chiplets are configured on both sides, L≥1, and the accelerator chiplet adopts a standard package, each interface chiplet is interconnected with a second chiplet interconnection interface of the accelerator chiplet through a first chiplet interconnection interface adopting the standard package; When two accelerator core particles share L interface core particles, L≥1, and the accelerator core particles adopt standard packaging, each interface core particle includes a first first core particle interconnection interface and a second first chip interconnection interface adopting standard packaging, wherein the first first core particle interconnection interface is interconnected with the second core particle interconnection interface of the first accelerator core particle, and the second first core particle interconnection interface is interconnected with the second chip interconnection interface of the second accelerator core particle. When two accelerator core particles share L interface core particles, L≥1, and the accelerator core particles adopt advanced packaging, each interface core particle includes a first first core particle interconnection interface and a second first chip interconnection interface adopting advanced or standard packaging, wherein the first first core particle interconnection interface is interconnected with the second core particle interconnection interface of the first accelerator core particle, and the second first core particle interconnection interface is interconnected with the second chip interconnection interface of the second accelerator core particle.

7. The system according to any one of claims 4 to 6, characterized in that: The accelerator chip and the interface chip packaged together may have the same or different process steps.

8. The system according to claim 1, wherein: The system further includes M types of switches, M≥1; each switch is connected to a high-speed interconnect interface of a corresponding type in the interface core particle.

9. An interface core particle, characterized in that: The interface core particle includes: N first chiplet interconnection interfaces, N ≥ 1; M types of high-speed interconnect interfaces, M ≥ 1; the m-th type of high-speed interconnect interface includes K(m) high-speed interconnect interfaces, K(m) ≥ 1, 1 ≤ m ≤ M; Each first chip interconnect interface is connected to F high-speed interconnect interfaces, 1≤F≤R, and R is the maximum number of high-speed interconnect interfaces that can be connected to the current first chip interconnect interface.

10. The core particle according to claim 9, characterized in that The interface chip includes N1 first chip interconnection interfaces using advanced packaging and N2 first chip interconnection interfaces using standard packaging, N1≥0, N2≥0, and N1+N2=N.

11. The core particle according to claim 9, characterized in that The value of R satisfies the bandwidth constraint: the bandwidth of the first chiplet interconnection interface is greater than the bandwidth of F high-speed interconnection interfaces; wherein, when the i-th first chiplet interconnection interface is connected to F i When there are high-speed interconnection interfaces, the following conditions must be met: i *dw i >∑ j=1 Fi (hs i,j *hw i,j ), where hs i,j is the number of channels of the jth high-speed interconnect interface connected to the i-th chiplet interconnect interface, hw i,j is the bandwidth of a single channel in the jth high-speed interconnect interface connected to the i-th chiplet interconnect interface; ds i is the number of channels of the i-th chiplet interconnection interface, dw i is the bandwidth of a single channel in the i-th chiplet interconnect interface.

12. The core particle according to claim 9, characterized in that The interface core particle includes the i-th core particle interconnection interface and the t-th core particle interconnection interface; Among them, ds i and ds t Different, ds t is the number of channels of the t-th first chiplet interconnection interface, where 1≤t≤N and i≠t; and / or, dw i and dw t Different, dw t is the bandwidth of a single channel in the t-th first chiplet interconnection interface.

13. The core particle according to claim 9, characterized in that The F high-speed interconnection interfaces include at least one type of high-speed interconnection interface.