Chip, chip communication method and apparatus
Patent Information
- Application Number
- CN202611056622.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2026-09-01
AI Technical Summary
[0005]本申请提供了一种芯片、芯片通信方法及装置,以解决上述不同类型的计算单元的裸片之间的通信效率较低的问题
[0034]第四方面,提供一种电子设备,包括如第一方面及第一方面中任一实现方式所述的芯片。
Smart Images

Figure CN122673147A_ABST
Abstract
Description
[0001] This application is a divisional application. The original application has the application number 202510334828.9 and the original application date is March 19, 2025. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This application relates to the field of chip communication technology, and in particular to a chip, a chip communication method and apparatus. Background Technology
[0003] General-purpose server processors typically employ a homogeneous computing architecture, integrating any number of general-purpose computing cores supporting x86, Advanced Reduced Instruction Set Computer (ARM), and RISC-V instruction sets into a single processor. With the increasing prevalence of scenarios such as large-scale model inference in artificial intelligence (AI), recommendation content inference, retrieval-augmented generation (RAG), and vector databases, general-purpose processors suffer from limitations in floating-point vector / matrix computing power and memory access bandwidth, which cannot meet the demands of these scenarios. Therefore, domain-specific architecture (DSA) accelerators exist to address these needs. DSA is an architecture design that combines different types of computing units, such as central processing units (CPUs), graphics processing units (GPUs), neural network processing units (NPUs), field-programmable gate arrays (FPGAs), and digital signal processing (DSP) chips, to perform tasks. Its aim is to leverage the strengths of different hardware units to optimize performance, power consumption, and computational efficiency.
[0004] When different types of computing units are packaged in a single chip, the computing cores of these different types of computing units, such as the dies, need to communicate to execute task calls. Communication between dies of different types of computing units is inefficient. Summary of the Invention
[0005] This application provides a chip, a chip communication method, and an apparatus to solve the problem of low communication efficiency between the dies of the aforementioned different types of computing units.
[0006] Firstly, a chip is provided. The chip includes a first die, a second die, and static random-access memory (SRAM). The first die is a general-purpose computing die, and the second die is a DSA die, or vice versa. The first die and the second die are respectively connected to the SRAM. The SRAM includes a first partition and a second partition. The first partition includes data structures for a first asynchronous task queue and a second asynchronous task queue. The first asynchronous task queue is used to manage tasks on the first die, and the second asynchronous task queue is used to manage tasks on the second die. The second partition is a last-level cache (LLC) shared by the first and second dies. In this chip, the first die and the second die are used to access data in the first partition based on the same visible physical address, thereby enabling inter-die communication based on the first asynchronous task queue and the second asynchronous queue.
[0007] Based on the aforementioned chip, general-purpose computing dies and DSA dies can access data in the first partition of SRAM based on the same visible physical address. This supports low-latency memory read / write operations between different dies using unified instructions, enabling inter-die communication based on hardware queues (such as a first asynchronous task queue and a second asynchronous task queue). Traditional chips use SRAM as a shared cache among multiple computing cores, and the address of this cache is invisible to user programs, making it impossible to perform locking, queuing, or other operations based on this cache. However, the chip provided in this application exposes the visible physical address of the first partition of SRAM to different dies, allowing programs on different dies to directly read and write the memory of the first partition to implement locking, queuing, and other operations. This supports inter-die communication based on hardware queues in lock-free scenarios, improving the communication efficiency between dies of different types of computing units.
[0008] In this application's embodiments, a bare die refers to the most basic component of an integrated circuit (IC) chip, which is an exposed chip directly fabricated on a silicon wafer using semiconductor manufacturing processes. In computer systems, the bare die is a critical component within the processor. As a possible example, different bare dies, such as a first bare die and a second bare die, may be packaged together based on three-dimensional integrated circuit (3DIC) technology.
[0009] For example, the core of a general-purpose computing unit, such as a general-purpose computing die, refers to an integrated circuit specifically responsible for performing computational tasks. It includes core processor elements such as the arithmetic and logic unit (ALU), control unit, cache, and possibly multiple processing cores. General-purpose computing dies employ a general-purpose architecture and typically support multiple instruction sets (e.g., x86, ARM, RISC-V). The design of general-purpose computing dies prioritizes versatility, enabling them to perform various types of computational tasks and possessing high flexibility. This general-purpose computing die can be a CPU die.
[0010] For example, the hardware design of the core of heterogeneous computing units, such as DSA dies, is customized according to the needs of specific tasks. This includes specially designed computing units, memory hierarchies, and data flow paths. DSA dies are specifically designed for a particular domain or application, such as artificial intelligence (AI), machine learning (ML), data acceleration, encryption, video processing, and network communication. The design of DSA dies focuses on performance optimization within a specific domain, rather than general-purpose computing. Therefore, DSA dies can be customized in terms of hardware and functionality as needed, designed entirely for specific tasks, and typically lack general-purpose computing capabilities.
[0011] As one possible implementation, the inter-die communication in this embodiment may include the following steps: a first die inserts a first task into a first asynchronous task queue; a second die retrieves the first task from the first asynchronous task queue; the second die executes the first task and writes the first task result into a first partition; the first task result is stored in the memory space corresponding to a first memory address in the first partition; the second die sends a notification message to the first die; the notification message includes the first memory address; the first die reads the first task result from the first partition based on the first memory address.
[0012] As one possible implementation, the first partition is a scratchpad memory (SPM).
[0013] In this way, using the first partition of SRAM as a note memory allows for addressing based on visible physical addresses, compared to multi-level caches, and the data access speed is faster. This enables dies with different architectures to communicate asynchronously via task queues based on the note memory, thus improving the efficiency of inter-die communication.
[0014] As one possible implementation, the first asynchronous task queue and the second asynchronous task queue are hardware atomic circular queues.
[0015] Optionally, the hardware atomic circular queue uses the storage space of the first partition as the queue lock address, and the data blocks of the first partition to store the queue data.
[0016] Thus, based on the atomic operations of the hardware atomic circular queue and the characteristics of the circular queue, the overhead of race conditions and locks in different architecture die environments is reduced, further improving the communication efficiency between dies.
[0017] As one possible implementation, the first partition is set to a non-cacheable state.
[0018] Thus, compared to the traditional multi-level cache where the general computing die or DSA die serially accesses the first-level cache (L1 cache) / second-level cache (L2 cache), the general computing die or DSA die can directly access the data in the first partition based on load (LD) / store (STR) instructions, reducing the latency of data access.
[0019] As one possible implementation, the first and second dies are used with atomic instructions of the same architecture to access data in the first partition based on a visible uniform physical address.
[0020] Thus, based on atomic instructions, the operation of shared data in the first partition by dies of different architectures can be prevented from violating data consistency. Furthermore, it can ensure safe synchronization between threads corresponding to dies of different architectures without using traditional locking mechanisms, reducing the operation latency caused by data consistency checks and lock contention, and further improving the communication efficiency between dies.
[0021] As one possible implementation, the chip also includes a bottom die, in which the first die, the second die, and the SRAM are packaged together.
[0022] In this way, the first and second dies are co-packaged on the underlying die, resulting in higher chip integration and shorter connection paths between different dies, thus ensuring efficient communication between them.
[0023] As one possible implementation, the chip also includes on-chip high-bandwidth memory, which is encapsulated in the underlying die and communicates with the first and second dies through the underlying die.
[0024] Optionally, the first and second dies are used to access data in the first partition and the on-chip high bandwidth memory (HBM) based on a visible unified physical address.
[0025] In this way, while achieving low-latency communication based on the first partition of SRAM, the chip can also support operations requiring large memory capacity, such as recommendation inference, vector database inference, or large language model (LLM) inference, based on the high-bandwidth on-chip memory, thus expanding the chip's applicable scenarios. On the other hand, compared to the serialization of operators or kernels by the general-purpose computing core of the general-purpose computing unit, followed by deserialization by the general-purpose computing core of the DSA computing unit before calling the operators or kernels for computation, communication between the bare dies of different types of computing units can be achieved through hardware queues or the distribution of operators or kernels via the high-bandwidth on-chip memory. This avoids the additional computational resource costs of serialization and deserialization, as well as the latency of operator or kernel distribution.
[0026] As one possible implementation, the chip also includes off-chip memory and input / output (IO) dies, with the first die and the second die connected to the off-chip memory via the IO dies.
[0027] Optionally, the first and second dies are used to access data from the first partition, on-chip high-bandwidth memory, and off-chip memory based on a visible unified physical address.
[0028] Optionally, the on-chip high-bandwidth memory is used to store memory pages whose memory page access frequency is greater than or equal to a preset threshold, and the off-chip memory is used to store memory pages whose memory page access frequency is less than the preset threshold.
[0029] In this way, the first partition of SRAM, on-chip high-bandwidth memory and off-chip memory can form a multi-level hierarchical memory to meet the different needs of different types of workloads for computing power, bandwidth and capacity, and further improve the applicability of the chip.
[0030] Secondly, a chip communication method is provided. This method is applied to a chip, which includes a first die, a second die, and a static random access memory (SRAM). The first die is a general-purpose computing die, and the second die is a domain-specific architecture (DSA) die, or the first die is a DSA die and the second die is a general-purpose computing die. The first die and the second die are respectively connected to the SRAM. The SRAM includes a first partition, which includes a data structure for a first asynchronous task queue. The first asynchronous task queue is used to manage tasks on the first die. The method includes: the first die inserting a first task into the first asynchronous task queue; the second die retrieving the first task from the first asynchronous task queue; the second die executing the first task and writing the first task result into the first partition; the first task result being stored in the memory space corresponding to a first memory address in the first partition; the second die sending a notification message to the first die; the notification message including the first memory address; and the first die reading the first task result from the first partition based on the first memory address.
[0031] As one possible implementation, the first die, the second die, and the SRAM are packaged together on the underlying die. The chip also includes on-chip high-bandwidth memory and input / output (I / O) dies packaged together on the underlying die. The first die and the second die are connected to external memory via the I / O dies. The method further includes: determining the memory page access frequencies of the on-chip high-bandwidth memory and the external memory; storing memory pages with memory page access frequencies greater than or equal to a preset threshold in the on-chip high-bandwidth memory; and storing memory pages with memory page access frequencies less than the preset threshold in the external memory.
[0032] The above-described chip communication method is applied to the chip described in the first aspect and any implementation thereof. For details and beneficial effects, please refer to the description in the first aspect, which will not be repeated here.
[0033] Thirdly, a chip communication device is provided, comprising a chip as described in the first aspect and any implementation thereof. As an example, the chip communication device includes a processor and a memory, the processor comprising the chip as described in the first aspect and any implementation thereof, and the memory storing instructions to be executed by the processor.
[0034] Fourthly, an electronic device is provided, comprising a chip as described in the first aspect and any implementation thereof.
[0035] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.
[0036] The following description includes more specific details about the implementation methods provided for the above aspects. Attached Figure Description
[0037] Figure 1 A schematic diagram of the structure of a network system provided in an embodiment of this application; Figure 2 This application provides a schematic diagram of the structure of a server according to an embodiment of the present application. Figure 3 A schematic diagram of a chip structure provided in this application embodiment. Figure 1 ; Figure 4 This is a schematic diagram of an SRAM structure provided in an embodiment of this application; Figure 5 A schematic diagram illustrating the principle of a hardware atomic circular queue provided in an embodiment of this application; Figure 6 A schematic diagram of a chip structure provided in this application embodiment. Figure 2 ; Figure 7 A schematic diagram illustrating the principle of a multi-level memory provided in an embodiment of this application; Figure 8 A schematic flowchart illustrating a chip communication method provided in an embodiment of this application; Figure 9 A flowchart illustrating a memory page allocation step provided in an embodiment of this application; Figure 10 This is a schematic diagram of a chip communication device provided in an embodiment of this application; Figure 11 A schematic diagram of the structure of an electronic device provided in this application embodiment. Figure 1 ; Figure 12 A schematic diagram of the structure of an electronic device provided in this application embodiment. Figure 2 . Detailed Implementation
[0038] The technical solutions involved in this application may be applied not only to current chip technologies or chip devices, but also to future chip technologies or chip devices, or to processors, computer systems, or servers including chips. The terminology used in the embodiments section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. A brief introduction to some concepts that may be involved in this application is given below.
[0039] General-purpose computing is designed for a wide range of tasks, emphasizing flexibility, and typically relies on a single type of hardware, such as a CPU. Taking the CPU as an example, it is designed as a general-purpose computing platform capable of handling various types of tasks, such as mathematical operations, control flow, and data access. General-purpose computing units typically support complex control flows and diverse instructions (such as branch prediction and out-of-order execution).
[0040] Domain-specific computing (DSA) is designed for domain-specific tasks, emphasizing extreme performance and efficiency. These architectures are optimized for specific types of computing tasks, providing higher performance and efficiency than general-purpose computing hardware. DSA typically employs dedicated hardware, such as FPGAs, ASICs, or specific processor units, to perform specific computing tasks, such as deep learning inference, image processing, and encryption. DSA computing units usually have dedicated instructions, with hardware logic optimized for specific operations (such as matrix multiplication and image rendering).
[0041] Notepad memory is a small, high-speed type of memory typically used in processors or computer systems to store temporary data for quick retrieval and processing. Notepad memory is usually a small, fast storage area integrated directly within the chip. Unlike register files, notepad memory can store larger blocks of data and can be shared by multiple threads. Furthermore, notepad memory has the following characteristics: because it is typically directly connected to the processor or core, it is generally faster than traditional memory; its small capacity results in lower overhead for data storage and retrieval, contributing to higher access speeds; and it is used to store intermediate computational results, reducing data access latency.
[0042] A hardware queue typically refers to a data queue stored in hardware resources, used to store data packets or tasks for hardware modules to read. Hardware queues usually have dedicated hardware support, providing fast access performance. A hardware atomic circular queue is a special type of hardware queue that not only possesses the characteristics of a circular queue (i.e., the queue space is used cyclically when the head pointer and tail pointer meet), but also supports atomic operations. This means that when operating on the queue in a concurrent environment, no race conditions will occur, ensuring data correctness and consistency.
[0043] Atomic instructions are hardware-level instructions in computer systems used to ensure the atomicity of shared data operations in multi-threaded or multi-core environments. Their core characteristic is the indivisibility of the operation—that is, the operation either executes completely or not at all, without any intermediate state being interfered with by other threads. For example, an atomic addition operation involves three steps: "reading the value → modifying the value → writing the value back," but other threads cannot insert operations between these three stages. Atomic instructions do not rely on traditional locks (such as mutexes), can directly manipulate shared resources, reduce thread blocking and context switching overhead, and improve parallel efficiency.
[0044] In semiconductor manufacturing, chips (or wafers) need to be packaged after production to better connect to the outside world and protect internal circuitry. The bottom die refers to the die located at the very bottom of a multi-chip package or 3D package structure. The bottom die can be connected to other dies, such as the top die, using technologies like microbumps and through-silicon vias (TSVs). In 3D packaging, multiple dies are stacked together, and the bottom die is typically the chip at the bottom layer, responsible for handling specific functions or interfacing with the external environment.
[0045] SRAM stands for Static Random Access Memory, widely used for storing temporary data and as cache memory. In a chip, it is primarily used in areas requiring high-speed access, such as CPU caches (L1 cache, L2 cache, L3 cache, last-level cache, etc.), to store recently accessed data, accelerating data retrieval and reducing main memory access time. For example, the last-level cache shared by multiple cores in a chip (taking L3 cache as an example) is built based on SRAM.
[0046] This application provides a chip including a first die, a second die, and SRAM. The first die is a general-purpose computing die, and the second die is a DSA die, or vice versa. The first die and the second die are respectively connected to the SRAM. The SRAM includes a first partition and a second partition. The first partition includes a data structure for a first asynchronous task queue and a second asynchronous task queue. The first asynchronous task queue is used to manage tasks on the first die, and the second asynchronous task queue is used to manage tasks on the second die. The second partition is an LLC shared by the first die and the second die. In this chip, the first die and the second die are used to access data in the first partition based on the same visible physical address, thereby realizing inter-die communication based on the first asynchronous task queue and the second asynchronous queue.
[0047] Based on the chip described above, the first partition of the chip's SRAM exposes visible physical addresses to different dies, allowing programs on different dies to directly read and write the memory of the first partition to implement operations such as locks and queues. This enables different dies to communicate with each other based on hardware queues in lock-free scenarios, improving the communication efficiency between dies of different types of computing units.
[0048] The application scenarios of the embodiments of this application will be described below with reference to the accompanying drawings.
[0049] The chip provided in this application can be used in servers or racks in Internet data centers and enterprise data center scenarios. Servers or racks containing the chip provided in this application can utilize the chip's built-in high floating-point computing power and high bandwidth to support new applications, including AI inference, or use large AI models to help improve the performance or functionality of traditional applications.
[0050] Figure 1 This is a schematic diagram of the structure of a network system provided in an embodiment of this application.
[0051] This network system can belong to a data center network topology, an interconnection between multiple data centers, or a wide area network. The service scenarios of the network system can include distributed machine learning training, distributed storage, high-performance computing, containerization, and other high-performance service scenarios. The communication protocols of the network system can be remote direct memory access (RDMA) protocols, transmission control protocols (TCP), such as infinite bandwidth and RoCEv2 RDMA protocols.
[0052] For example, network system 100 includes multiple servers and multiple network devices.
[0053] Multiple servers are used to support high-performance services with different communication requirements, such as AI training, AI inference, and storage. Figure 1 As shown, network system 100 includes multiple server groups ( Figure 1 Only server groups 101, 102, 103, and 104 are shown, but the number of server groups is not limited to four. Each server group includes one or more servers. Figure 1 Only three servers are shown, but it is not limited to three servers. For example, server group 101 includes servers 1-3, server group 102 includes servers 4-6, server group 103 includes servers 7-9, and server group 104 includes servers 10-12. The servers in each server group are connected to network devices via network interface cards (NICs).
[0054] The server involved in this application embodiment can be a server for providing cloud computing services, that is, a server that provides corresponding resources to clients. For example, the resources provided by the server to the client can be storage resources, which can be storage systems, such as file storage systems, block storage systems, or object storage systems, or combinations of the above storage systems. The resources provided by the server to the client can also be computing resources, such as virtual machines used for AI training, AI inference, and other businesses. This application embodiment does not limit the type of server or the resources provided by the server.
[0055] Multiple network devices can be routers, switches, gateways, and other devices with data exchange and transmission functions. For example... Figure 1 As shown, network system 100 includes multiple network devices ( Figure 1 Only network devices 105, 106, 107, and 108 are shown, but there are no restrictions on the number of network devices. Network devices can be located at different levels in network system 100.
[0056] For example, in network system 100, each server in server group 101 and server group 102 is communicatively connected to network device 105 via a network interface card (NIC), and each server in server group 103 and server group 104 is communicatively connected to network device 106 via a NIC. Network device 105 is communicatively connected to network devices 107 and 108, respectively, and network device 106 is communicatively connected to network devices 107 and 108, respectively.
[0057] It should be understood that Figure 1 This is a simplified diagram for ease of understanding only. The network system 100 may also include other network devices, servers, and / or other devices, and the connection relationships between nodes may also vary. Figure 1 It was not drawn in the middle.
[0058] Figure 2 This is a schematic diagram of the structure of a server provided in an embodiment of this application. The server can be... Figure 1 Any of the servers shown. Please refer to... Figure 2 The server 200 includes components such as a memory 210, a chip 220, a baseboard management controller (BMC) 230, and a network interface card 240. Those skilled in the art will understand that... Figure 2 The server structure shown in the figure does not constitute a limitation on the server. The server provided in the embodiments of this application may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0059] The following is combined with Figure 2 A detailed introduction to each component of server 200: Memory 210 can be used to store operating system (OS) instructions, data instructions, and data generated by memory chip 220 when executing various functional applications of server 200 by running the instructions of the OS. The data instructions are operation instructions used to indicate method flows. Optionally, memory 210 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Figure 2 As shown, the memory 210 may include at least one memory card. Figure 2 Taking three memory cards as an example, labeled memory card 211, memory card 212, and memory card 213 respectively. Specifically, a memory card can be any storage medium that can serve as memory, such as a memory board or a memory module. A memory card can refer to a single memory module or a memory board, or it can refer to a collection of multiple memory modules or a collection of multiple memory boards. It should be noted that... Figure 2 This is merely an example of a memory in the embodiments of this application, and the embodiments of this application do not limit the number or form of memory.
[0060] The OS involved in the embodiments of this application is the most basic system software running on a server. For example, the operating system can be a Windows operating system, a Linux operating system, or a VMware operating system, etc., and there is no limitation herein.
[0061] Chip 220 is the control center of server 200. It connects various components through various interfaces and lines. By running or executing the instructions of the OS stored in memory 210 and calling the data stored in memory 210, it performs various functions of server 200 and processes data, thereby realizing various services based on server 200.
[0062] The chip 220 in this embodiment includes a first die, a second die, and SRAM. The first die is a general-purpose computing die, and the second die is a DSA die, or vice versa. The first die and the second die are respectively connected to the SRAM. The SRAM includes a first partition and a second partition. The first partition includes a data structure for a first asynchronous task queue and a second asynchronous task queue. The first asynchronous task queue is used to manage tasks on the first die, and the second asynchronous task queue is used to manage tasks on the second die. The second partition is an LLC shared by the first die and the second die. In this chip 220, the first die and the second die are used to access data in the first partition based on the same visible physical address, so as to realize inter-die communication based on the first asynchronous task queue and the second asynchronous queue.
[0063] Please refer to the specific structure and function of the aforementioned chip 220. Figure 3 , Figure 4 The details and related descriptions will not be repeated here.
[0064] The BMC230 is used to monitor and control the hardware of server 200. For example, it can monitor information such as the temperature and voltage of server 200 and make corresponding adjustments to ensure that server 200 is in a normal operating state. When server 200 is in an abnormal state, it can also be restarted by resetting. The BMC230 can also record various hardware information and logs to provide information to users of server 200 and to help locate faults. It should be noted that the BMC230 is an independent device; it does not depend on other hardware in server 200 (such as processor 220 or memory 210) nor on the operating system. However, the BMC230 can interact with the operating system, such as... Figure 2 As shown, the BMC230 includes a BMC storage medium 231 and a BMC processor 232. The BMC processor 232 interacts with the OS, enabling better management. The BMC storage medium 231 stores the instructions required for the BMC230 to run and the data generated during runtime. This BMC storage medium 231 can be flash memory, random access memory (RAM), or read-only memory (ROM), etc., and is not limited here.
[0065] Network interface card 240 is used to establish physical connections with other devices. Network interface card 240 has at least one network port, for example, in... Figure 2 The server 200 includes network interface cards (NICs 1, 2, and 3). At least one of these NICs is connected to other devices via a cable to enable data transmission between the server 200 and other devices. Optionally, NIC 240 can be integrated on a circuit board or inserted through a high-speed serial computer expansion bus standard (PCIE) slot or a PCI slot in the server 200; no limitation is imposed here.
[0066] Next, combine Figure 3 The chip provided in this application is described in detail.
[0067] Figure 3 A schematic diagram of a chip structure provided in this application embodiment. Figure 1 .like Figure 3 As shown, chip 300 can be the above-mentioned Figure 2 The chip 220 includes a first die 301, a second die 302, and an SRAM 303.
[0068] As one possible implementation, the first die 301 is a general-purpose computing die, and the second die 302 is a DSA die.
[0069] As another possible implementation, the first die 301 is a DSA die, and the second die 302 is a general computing die.
[0070] The first die 301 and the second die 302 are respectively connected to the SRAM 303. For example, the first die 301, the second die 302, and the SRAM 303 are encapsulated in the underlying die 304, and the first die 301 and the second die 302 are communicatively connected to the SRAM 303 through the network on chip (NOC) of the underlying die 304.
[0071] Network-on-chip (NIC) is a network architecture used to enable communication between various components in a system-on-chip (SoC). It is a high-efficiency communication architecture used to connect multiple processor cores (such as multiple dies), memory units, peripherals, and other hardware modules, supporting high-bandwidth, low-latency data transmission.
[0072] SRAM303 and on-chip network can be integrated into the underlying die 304 to improve the data access speed of the first die 301 and the second die 302 to SRAM303.
[0073] The storage space of SRAM303 can be divided into a first partition and a second partition. The first partition serves as a note-taking memory shared by the first die 301 and the second die 302, while the second partition serves as a last cache shared by the first die 301 and the second die 302.
[0074] This application does not limit the number of general-purpose computing dies or DSA dies in chip 300. For example, in possible embodiments of this application, chip 300 may also include a third die, a fourth die, a fifth die, and any number of dies, any of which may be a general-purpose computing die or a DSA die. All dies are encapsulated in a bottom die 304, and are communicatively connected to SRAM 303 through the bottom die 304 (such as an on-chip network), and share the second region of SRAM 303. Each die has a corresponding asynchronous task queue in the first region of SRAM 303.
[0075] Next, combine Figure 4 A detailed description of SRAM303 is provided.
[0076] Figure 4 This is a schematic diagram of an SRAM structure provided in an embodiment of this application. Figure 4 As shown, SRAM303 includes a first partition 401 and a second partition 402.
[0077] The first die 301 and the second die 302 can use the same architecture of atomic instructions to access data in the first partition 401 based on the visible unified physical address.
[0078] Optionally, the first die 301 only supports basic vector instructions and does not integrate instruction-driven matrix computation acceleration, such as scalable matrix extension (SME), within the core.
[0079] Atomic instructions are implemented through hardware support (such as CPU instruction sets) or software emulation (such as locks), and their behavior must conform to the visibility and ordering constraints defined by memory semantics. Memory semantics categorizes operations on atomic instructions of different atomic types into three main types: load, store, and read-modify-write. Loading can include memory_order_relaxed, memory_order_release, or memory_order_seq_cst memory semantic order. Storening can include memory_order_relaxed, memory_order_consume, memory_order_acquire, or memory_order_seq_cst memory semantic order. Read-modify-write can include memory_order_relaxed, memory_order_consume, memory_order_acquire, memory_order_release, memory_order_acq_rel, or memory_order_seq_cst memory semantic order. In this context, atomic instructions that conform to any memory semantic order can also be called instructions of that memory semantic type. For example, atomic instructions that conform to the memory semantic order of loading can be called load instructions, and atomic instructions that conform to the memory semantic order of storing can be called store instructions, etc.
[0080] The first partition 401 includes the data structures of the first asynchronous task queue and the second asynchronous task queue, that is, the first asynchronous task queue and the second asynchronous task queue are hardware queues.
[0081] As one possible implementation, the first and second asynchronous task queues are hardware atomic circular queues. Hardware atomic circular queues rely on hardware-supported atomic instructions, such as compare and swap (CAS) or load linked / store conditional (LL / SC) instructions.
[0082] The first asynchronous task queue and the second asynchronous task queue in the first partition 401 have similar structures. Each queue in the first partition 401 corresponds to a hardware atomic circular queue data structure. This hardware atomic circular queue data structure includes array storage, a head pointer, and a tail pointer. The array storage uses contiguous memory space (array) in the first partition 401 to simulate a circular buffer. The head pointer (front or head) points to the next element to be read in the queue. The tail pointer (rear or tail) points to the next empty position to be written to in the queue. The characteristic of this hardware atomic circular queue is its circular nature; when the tail pointer reaches the end of the queue, it wraps back to the beginning of the queue.
[0083] For example, such as Figure 5 As shown, the first partition 401 is used to support a hardware atomic circular queue. The storage space of the first partition 401 is used as the queue lock address, and the data blocks of the first partition store the queue data. That is, one data block corresponds to one queue data to achieve array storage. The application is based on software code, such as atomic instructions to perform lock acquisition, update queue head / tail operations, or instructions to enqueue / dequeue task data to perform task data enqueue / dequeue operations.
[0084] For details regarding the communication between the first die 301 and the second die 302 in chip 300 based on the first asynchronous task queue and the second asynchronous task queue of the first partition 401, please refer to [link / reference needed]. Figure 8 The details and related content will not be elaborated here.
[0085] As one possible implementation, the address space of the first partition 401 is set to non-cacheable. Non-cacheable means that data cannot be stored in the cache; each access to this memory must directly reach physical memory. Setting the first partition 401 as a notepad storage as non-cacheable means that data in this memory area will not be cached (i.e., it will not be stored in CPU caches such as L1 cache, L2 cache, or the last level cache), and each access will directly read from or write to this memory. In this way, the first die 301 and the second die 302 avoid the latency of serially searching the L1 and L2 caches when accessing data in the first partition 401, ensuring a low load-to-use latency for a single access based on atomic instructions such as memory ld / str instructions.
[0086] When the second partition 402 serves as the last-level cache shared by general-purpose computing dies and DSA dies such as the first die 301 and the second die 302, the first die 301 and the second die 302 can use a lookup to retrieve specific data from the second partition 402. A lookup refers to checking whether the requested data is already stored in the cache.
[0087] The last level of cache typically refers to the third level cache (L3 cache) in the CPU cache hierarchy. It is the last high-speed cache between the CPU and main memory, used to coordinate the speed differences between lower-level caches (L1, L2) and main memory, reducing CPU latency in accessing memory. The L3 cache is usually designed to be shared by multiple CPU cores. For example, in a multi-core processor, all cores share the same L3 cache to coordinate data consistency between different cores and reduce memory access conflicts.
[0088] For example, when a processor or program needs data, it makes a request to the cache. The system checks whether the data is already stored in the cache; this process is called a lookup and typically uses specific algorithms (such as hash tables, direct mappings, group associations, etc.) to check if the data is in the cache. If the data is already in the cache, the system returns the data directly from the cache, which usually speeds up the access process. If the data is not in the cache, the system retrieves the data from slower storage (such as main memory) and then loads it into the cache for later access.
[0089] As can be seen, while the first partition 401 exposes a visible unified physical address to the first die 301 and the second die 302, the physical address of the second partition 402 is not visible to the first die 301 and the second die 302.
[0090] Based on the aforementioned chip, general-purpose computing dies and DSA dies can access data in the first partition of SRAM based on the same visible physical address. This supports low-latency memory read / write operations between different dies using unified instructions, enabling inter-die communication based on hardware queues (such as a first asynchronous task queue and a second asynchronous task queue). Traditional chips use SRAM as a shared cache among multiple computing cores, and the address of this cache is invisible to user programs, making it impossible to perform locking, queuing, or other operations based on this cache. However, the chip provided in this application exposes the visible physical address of the first partition of SRAM to different dies, allowing programs on different dies to directly read and write the memory of the first partition to implement locking, queuing, and other operations. This supports inter-die communication based on hardware queues in lock-free scenarios, improving the communication efficiency between dies of different types of computing units.
[0091] The above text combined Figures 3-5The chip provided in this application is described. This chip 300 is constructed based on a die and SRAM, and high-speed communication between dies is achieved through a hardware atomic circular queue supported by the first partition of the SRAM. In some possible application scenarios of chip 300, in addition to high-speed communication between dies, chip 300 may also need to meet the differentiated requirements of different types of workloads for computing power, bandwidth, and capacity. Therefore, embodiments of this application also provide a chip that integrates large-capacity programmable SRAM and high-bandwidth memory by adding high-bandwidth memory, off-chip memory, etc., to chip 300, and combines it with off-chip memory to form a multi-level hierarchical memory, meeting the differentiated requirements of different types of workloads for computing power, bandwidth, and capacity.
[0092] Next, combine Figure 6 Another chip 300 provided in the embodiments of this application will be described in detail.
[0093] Figure 6 A schematic diagram of a chip structure provided in this application embodiment. Figure 2 .like Figure 6 As shown, in Figure 3 Based on the structure of the chip 300 shown, the chip 300 also includes a substrate 305, a high-bandwidth memory 306, and a high-bandwidth memory 307.
[0094] Substrate 305 can be a package substrate (SUB). A package substrate is a basic material used to support semiconductor chips and provide electrical connections. The main function of the package substrate is to connect the semiconductor chip to external circuits and provide physical protection and thermal management for the chip. It is an important part of the semiconductor packaging structure.
[0095] The substrate 305 may serve the following functions: electrical connection, responsible for transmitting electrical signals on the chip to external circuits through solder balls, pins or other connection methods; thermal management, helping to dissipate heat and prevent the chip from being damaged due to overheating; physical protection, providing physical protection so that the chip is not affected by external physical impacts and contamination; and mechanical support, providing stable support for the chip so that it is not affected by external stress during operation.
[0096] This application does not limit the material of the substrate 305. For example, the substrate 305 can be a ceramic substrate, an organic substrate, a metal substrate, a glass substrate, etc.
[0097] High-bandwidth memory 306 and high-bandwidth memory 307 are encapsulated together with the underlying die 304 on substrate 305. High-bandwidth memory 306 and high-bandwidth memory 307 are on-chip high-bandwidth memory and are communicatively connected to the first die 301 and the second die 302 through substrate 305 and underlying die 304.
[0098] The high-bandwidth memory in the chip is a high-performance dynamic random access memory (DRAM) based on 3D stacking technology, which can meet the application scenarios that require high bandwidth, low latency and large capacity.
[0099] Because the high-bandwidth memory 306 and high-bandwidth memory 307 use a wide interface bus (such as 1024-bit width) and parallel channel design, the bandwidth of the high-bandwidth memory 306 and high-bandwidth memory 307 is higher than that of traditional memory. Furthermore, based on vertical stacking technology, the high-bandwidth memory 306 and high-bandwidth memory 307 form a vertical conductive channel inside, which shortens the data transmission path and reduces latency and power consumption.
[0100] This application does not limit the number or capacity of high-bandwidth memory in chip 300. For example, chip 300 may include any number of high-bandwidth memory modules, such as 3, 4, or 10. Furthermore, the capacity of each high-bandwidth memory module may be any capacity, such as 20 gigabytes (GB), 55GB, 200GB, or 2000GB.
[0101] The first die 301 and the second die 302 are used to access data from the first partition 401, the on-chip high-bandwidth memory 306, and the on-chip high-bandwidth memory 307 based on the visible unified physical address.
[0102] Thus, chip 300 incorporates a first partition 401, used as a notepad storage, and on-chip high-bandwidth memory between the traditional LLC and DRAM, forming a multi-level memory. It supports data access to the multi-level memory via a unified physical address visible to the first die 301 and the second die 302. The first partition 401 supports high-bandwidth, low-latency hardware communication queues between dies of different architectures, enabling high-speed communication between them. The on-chip high-bandwidth memory supports memory-bound applications such as LLM inference, recommendation inference, and vector databases. This solves the problems of insufficient floating-point vector / matrix computing power and insufficient memory access bandwidth in general-purpose processors, improving the applicability of chip 300.
[0103] On the other hand, the traditional collaborative computing process between chips with different architectures, taking CPU and GPU (or NPU) as an example, involves the CPU serializing the operator / kernel, transmitting it to the GPU / NPU, and then having the general-purpose computing core on the GPU / NPU deserialize it before calling the local operator / kernel. In the embodiments of this application, the general-purpose computing die can directly access the on-chip shared high-bandwidth memory or distribute operators / kernels through a hardware-accelerated queue, avoiding the additional computing resource costs and latency associated with serialization and deserialization. Because the serialization and deserialization operations are saved, the general-purpose computing core does not need to be integrated on the DSA die, reducing chip area overhead.
[0104] In some possible application scenarios of chip 300, considering the traditional general computing workload requirements of chip 300, embodiments of this application may also add off-chip memory to chip 300. Figure 6 (not shown in the image), together with SRAM303, high-bandwidth memory 306, and high-bandwidth memory 307, constitute a multi-level memory.
[0105] In one possible implementation, chip 300 includes IO die 308 and IO die 309. IO dies 308 and IO die 309 are encapsulated on substrate 305, and the first die 301 and the second die 302 of chip 300 are communicatively connected to off-chip memory through IO dies 308 and IO die 309.
[0106] Off-chip memory refers to memory that is not directly integrated on the processor chip (such as chip 300), but is connected to chip 300 through an external bus (such as PCIe bus, double data rate (DDR) bus, universal serial bus (USB) etc.).
[0107] The embodiments of this application do not limit the type of off-chip memory, for example, the off-chip memory is DRAM.
[0108] like Figure 7 As shown, on-chip high-bandwidth memory is used to store memory pages whose access frequency is greater than or equal to a preset threshold, while off-chip memory is used to store memory pages whose access frequency is less than the preset threshold. The preset threshold can be flexibly adjusted according to the specific usage scenario and application type of chip 300. Memory pages whose access frequency is greater than or equal to the preset threshold can be called latency-sensitive application memory pages, and memory pages whose access frequency is less than the preset threshold can be called bandwidth-sensitive application memory pages.
[0109] In traditional CPUs, memory page swapping between hierarchical memory layers with varying bandwidth, capacity, and latency is primarily controlled by the application through an application-aware mechanism. This places significant development pressure on porting traditional applications. This application's embodiment, by sampling page access counts in on-chip and off-chip memory, automatically moves pages requiring high-bandwidth access to on-chip memory. This allows applications to seamlessly enjoy high-bandwidth memory, greatly reducing the workload of application porting. The on-chip memory can be high-bandwidth memory 306, high-bandwidth memory 307, or the first partition 401 of SRAM 303. The off-chip memory can be off-chip storage.
[0110] Thus, chip 300 incorporates a first partition 401 for note storage, on-chip high-bandwidth memory, and off-chip memory between the traditional LLC and DRAM, forming a structure as follows: Figure 7 The multi-level memory shown supports data access to the multi-level memory via a unified physical address visible on the first die 301 and the second die 302. The first partition 401 can support high-bandwidth, low-latency hardware communication queues between dies of different architectures to achieve high-speed communication between different dies. The on-chip high-bandwidth memory can support memory-bound applications, such as LLM inference, recommendation inference, and vector databases. The off-chip memory mainly serves traditional general-purpose computing workloads. This further improves the applicability of chip 300.
[0111] On the other hand, the three levels of memory of chip 300—namely, the first partition 401, the on-chip high-bandwidth memory, and the off-chip memory—are uniformly addressed. The DSA die can directly access host memory with memory semantics and share the LLC on the SRAM 303 of the underlying die 304. General-purpose computing dies and DSA dies, such as the first die 301 and the second die 302, can achieve consistent data access based on unified addressing and shared memory space. This avoids the problems of inter-chip bus bandwidth being lower than memory access bandwidth and the high design complexity of maintaining consistency check (CC) on the inter-chip bus, which are caused by traditional chips of different architectures accessing CPU-side host memory.
[0112] As one possible implementation, the various parts of the chip 300 in the above embodiments of this application, such as the first die 301, the second die 302, the SRAM 303, the bottom die 304, the high-bandwidth memory 306, the high-bandwidth memory 307, the IO die 308, and the IO die 309, can be packaged using 3DIC technology.
[0113] 3D-DIC is a chip technology that achieves high-density integration by vertically stacking multiple chips or wafers and utilizing advanced interconnect technologies. 3D-DIC forms a multi-layer device structure by vertically stacking multiple chips (such as logic chips, memory chips, and RF chips) in three-dimensional space. This stacking can be at the die or wafer level, with inter-layer electrical connections achieved through through-silicon vias (TSVs) or hybrid bonding technologies.
[0114] In conjunction with the chip provided in the above embodiments of this application, such as Figure 3 or Figure 6 The chip 300 shown in this application embodiment also provides a chip communication method.
[0115] Please refer to Figure 8 , Figure 8 This is a schematic flowchart illustrating a chip communication method provided in an embodiment of this application. Figure 3 or Figure 6 Based on the chip 300 shown, the chip communication method may include the following steps S801-S805.
[0116] S801, the first bare die 301 inserts the first task into the first asynchronous task queue.
[0117] If the first die 301 is a general-purpose computing die such as a CPU die, it is suitable for general-purpose computing tasks and can handle various types of tasks. The first task may include operating system instruction scheduling, multi-threaded concurrency, logical operations, data transfer, etc.
[0118] If the first chip 301 is a DSA chip, it is suitable for specific domain tasks. The first task may include AI inference, graphics rendering, network packet processing, tensor computation, etc.
[0119] For details regarding the first asynchronous task queue, please refer to [link / reference needed]. Figure 4 The relevant descriptions of the queue will not be repeated here.
[0120] Taking the first asynchronous task queue as a hardware atomic circular queue as an example, the first die 301 inserts the first task into the first asynchronous task queue as an enqueue operation. This enqueue operation includes: atomically reading the value of the tail pointer of the current first asynchronous task queue; calculating the next available position based on the value of the tail pointer of the current first asynchronous task queue; checking whether the queue is full; if not full, updating the tail pointer using an atomic CAS operation and writing the first task data.
[0121] S802, the second die 302 obtains the first task from the first asynchronous task queue.
[0122] Taking the first asynchronous task queue as a hardware atomic circular queue as an example, the second die 302 obtains the first task from the first asynchronous task queue as a dequeue operation of the hardware atomic circular queue. This dequeue operation includes: atomically reading the value of the head pointer of the current first asynchronous task queue; determining whether the queue is empty based on the value of the head pointer of the current first asynchronous task queue; if it is empty, returning failure; if it is not empty, updating the head pointer using an atomic CAS operation and reading the first task data.
[0123] S803, the second bare die 302 executes the first task and writes the result of the first task to the first partition.
[0124] The result of the first task is stored in the memory space corresponding to the first memory address in the first partition.
[0125] As one possible implementation, the second die 302 can also write the result of the first task of the first task into the memory space of the first die 301 that can access data based on a unified physical address, such as high-bandwidth memory 306, high-bandwidth memory 307, or off-chip memory.
[0126] S804, the second die 302 sends a notification message to the first die 301.
[0127] The notification message includes the first memory address.
[0128] S805, the first bare die 301 reads the first task result from the first partition based on the first memory address.
[0129] In a possible embodiment of this application, after S802, the second die 302 can use the enqueue operation of S801 to write the first task result to the first partition, and the first die 301 can use the dequeue operation of S802 to read the first task result from the first partition.
[0130] As one possible implementation, the operations of the first die 301 and the second die 302 in S801-S805 above can be implemented by the application programs corresponding to the first die 301 and the second die 302.
[0131] Thus, the first partition of the SRAM in the chip provided in this application exposes visible physical addresses to different dies, allowing programs on different dies to directly read and write the memory of the first partition to implement operations such as locks and queues. This supports communication between different dies based on hardware queues in lock-free scenarios, improving the communication efficiency between dies of different types of computing units.
[0132] In the case of a multi-level memory architecture in chip 300, in addition to implementing the aforementioned chip communication methods, chip 300 can also dynamically allocate memory pages non-explicitly between on-chip and off-chip memory. The following will combine... Figure 9 The steps of memory page allocation are explained in detail.
[0133] Please refer to Figure 9 , Figure 9 This is a flowchart illustrating a memory page allocation step provided in an embodiment of this application. Figure 3 or Figure 6 Based on the chip 300 shown, the memory page allocation step may include the following steps S901-S903.
[0134] S901. Determine the memory page access frequency of on-chip high-bandwidth memory and off-chip memory.
[0135] As one possible implementation, the chip 300 dynamically detects the application's load characteristics during runtime through a performance monitor unit (PMU). These load characteristics include the number of accesses (or access frequency) of memory pages (such as 4K pages) per unit time, based on statistical profiling extension (SPE).
[0136] S902: Store memory pages whose memory page access frequency is greater than or equal to a preset threshold to on-chip high-bandwidth memory.
[0137] Chip 300 uses a swap operation to move memory pages with access frequencies greater than or equal to a preset threshold from off-chip memory to on-chip high-bandwidth memory.
[0138] S903. Store memory pages whose memory page access frequency is less than a preset threshold to off-chip memory.
[0139] Chip 300 uses a swapping operation to move memory pages with access frequencies below a preset threshold from on-chip high-bandwidth memory to off-chip memory.
[0140] The preset threshold can be flexibly adjusted according to the frequency of data access by the application.
[0141] To complement the aforementioned chip communication method, this application also provides a chip communication device. Please refer to... Figure 10 , Figure 10 This is a schematic diagram of a chip communication device provided in an embodiment of this application.
[0142] The chip communication device 1000 includes a write module 1010, a read module 1020, and a notification module 1030. The write module 1010 instructs a first die to insert a first task into the first asynchronous task queue. The read module 1020 instructs a second die to retrieve the first task from the first asynchronous task queue. The write module 1010 instructs the second die to execute the first task and write the first task result to the first partition; the first task result is stored in the memory space corresponding to a first memory address in the first partition. The notification module 1030 instructs the second die to send a notification message to the first die. The read module 1020 instructs the first die to read the first task result from the first partition based on the first memory address.
[0143] As one possible implementation, the chip communication device 1000 further includes a detection module for determining the memory page access frequencies of the on-chip high-bandwidth memory and the off-chip memory. The write module 1010 and the read module 1020 are also configured to collaboratively store memory pages with access frequencies greater than or equal to a preset threshold into the on-chip high-bandwidth memory, and to store memory pages with access frequencies less than the preset threshold into the off-chip memory.
[0144] All of the above modules can be implemented in software, hardware, or a combination of both.
[0145] This application also provides an electronic device, such as... Figure 11 As shown, the electronic device 1100 includes a circuit board (not shown in the figure) and a chip system 1110. The chip system 1110 is disposed on the circuit board. Figure 11 As shown, the chip system 1110 includes Figure 3 or Figure 6 The chip shown is 300.
[0146] In some possible implementations, the electronic device 1100 also includes a display. The display is used to show the processing results of the tasks performed by the chip 300.
[0147] For example, electronic device 1100 can be a mobile device or an edge device. For example, electronic device 1100 can be a mobile phone, tablet computer, smartwatch, television, vehicle control device, and smart home device, etc. For example, electronic device 1100 can be a server, server rack, etc.
[0148] In some possible implementations, the chip system 1110 includes Figure 3 or Figure 6 The chip shown is 300.
[0149] In some possible implementations, the chip system 1110 may be provided with general or dedicated storage circuitry to store task results.
[0150] In some examples, when the application is running, the chip system 1110 can automatically execute the contents described in the above embodiments. In some examples, when the application is running, in response to user input, the chip system 1110 begins to execute the contents described in the above embodiments.
[0151] The chip 300 and its related hardware and circuits in the above embodiments of this application can be considered as a chip system. The first die 301 and its related hardware and circuits in the chip 300 can be considered as a chip system. The second die 302 and its related hardware and circuits in the chip 300 can be considered as a chip system. This application does not limit the specific scope of the chip system.
[0152] Next, combine Figure 12 The possible hardware structure of electronic device 1100 is described.
[0153] Figure 12 This embodiment provides a schematic diagram of the structure of an electronic device. Figure 2 .like Figure 12 As shown, the electronic device 1200 and Figure 11 The electronic device 1100 shown may be the same device. Electronic device 1200 includes a chip 1210, a bus 1220, a memory 1230, a communication interface 1240, and a memory unit 1250 (also referred to as main memory). The chip 1210, memory 1230, memory unit 1250, and communication interface 1240 are connected via the bus 1220.
[0154] It should be understood that, in this embodiment, chip 1210 is Figure 3 or Figure 6 The chip 300 is shown.
[0155] In a possible embodiment, electronic device 1200 may refer to chip 1210.
[0156] The communication interface 1240 is used to enable communication between the electronic device 1200 and external devices or components. In this embodiment, the electronic device 1200 is used to implement... Figure 1 When any network device or server is in use, the communication interface 1240 is used as a physical port for sending and receiving data packets.
[0157] Bus 1220 may include a pathway for transmitting information between the aforementioned components (such as chip 1210, memory unit 1250, and memory 1230). In addition to a data bus, bus 1220 may also include a power bus, a control bus, and a status signal bus, etc. However, for clarity, ... Figure 12 In this context, various buses are labeled as Bus 1220. Bus 1220 can be a PCIe bus, or an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL) memory interconnect, a Cache Coherent Interconnect for Accelerators (CCIX), etc. Bus 1220 can be categorized into address bus, data bus, and control bus.
[0158] As an example, electronic device 1200 may include multiple chips. A chip may be single-core or multi-core. Here, a chip may refer to one or more devices, circuits, and / or computing units used to process data (e.g., computer program instructions).
[0159] It is worth noting that, Figure 12 Taking the electronic device 1200 as an example, which includes one chip 1210 and one memory 1230, the chip 1210 and the memory 1230 are used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to business needs.
[0160] Memory cell 1250 may be volatile memory or non-volatile memory, or may include both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0161] The memory 1230 can correspond to the storage medium used to store computer instructions and other information in the above method embodiments, such as a disk, like a mechanical hard disk or a solid-state hard disk.
[0162] The aforementioned electronic device 1200 can be a general-purpose device or a special-purpose device. For example, electronic device 1200 can be an edge device (e.g., a box carrying a chip with processing capabilities). Alternatively, electronic device 1200 can also be a chip, network device, server, or other device with computing capabilities.
[0163] It should be understood that the electronic device 1200 according to this embodiment may correspond to the chip 300 in this embodiment or any device or apparatus containing the chip 300, and may correspond to the execution of the device according to the embodiment. Figure 8 or Figure 9 The corresponding subject in the Chinese method.
[0164] This application also provides a computer-readable storage medium including instructions; when the instructions are executed on the electronic device described in the above embodiments, the electronic device causes the electronic device to perform the chip communication method described in the above embodiments.
[0165] The processor involved in the embodiments of this application can be a chip. For example, it can be a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a system-on-chip (SoC), a CPU, a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), or other integrated chips.
[0166] The memory involved in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0167] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0168] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0169] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0170] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or modules may be electrical, mechanical, or other forms.
[0171] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located on one device or distributed across multiple devices. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0172] In addition, the functional modules in the various embodiments of this application can be integrated into one device, or each module can exist physically separately, or two or more modules can be integrated into one device.
[0173] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0174] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A chip, characterized in that, The chip includes a first die, a second die, and a static random access memory (SRAM). The first die is a general-purpose computing die, and the second die is a domain-specific architecture (DSA) die. The first die and the second die are respectively connected to the SRAM. The SRAM includes a first partition and a second partition. The first partition includes a data structure of a first asynchronous task queue and a second asynchronous task queue. The first asynchronous task queue is used to manage the tasks of the first die, and the second asynchronous task queue is used to manage the tasks of the second die. The second partition is a last-level cache LLC shared by the first die and the second die. The first and second dies are used to access data in the first partition to enable inter-dies communication based on the first and second asynchronous task queues.
2. The chip according to claim 1, characterized in that, The inter-die communication includes the following steps: The first die inserts the first task into the first asynchronous task queue; The second die retrieves the first task from the first asynchronous task queue; The second die executes the first task and writes the first task result of the first task into the first partition; The result of the first task is stored in the memory space corresponding to the first memory address in the first partition; The second die sends a notification message to the first die; the notification message includes the first memory address; The first die reads the first task result from the first partition based on the first memory address.
3. The chip according to claim 1 or 2, characterized in that, The first partition is a note storage area.
4. The chip according to any one of claims 1-3, characterized in that, The first asynchronous task queue and the second asynchronous task queue are hardware atomic circular queues.
5. The chip according to claim 4, characterized in that, The hardware atomic circular queue uses the storage space of the first partition as the queue lock address and the data blocks of the first partition to store queue data.
6. The chip according to any one of claims 1-5, characterized in that, The first partition is set to a non-cached state.
7. The chip according to any one of claims 1-6, characterized in that, The first and second dies are used to access data in the first partition using atomic instructions of the same architecture, based on a visible unified physical address.
8. The chip according to any one of claims 1-7, characterized in that, The chip also includes an underlying die, wherein the first die, the second die, and the SRAM are encapsulated together on the underlying die.
9. The chip according to claim 8, characterized in that, The chip also includes on-chip high-bandwidth memory, which is encapsulated in the underlying die and communicates with the first die and the second die through the underlying die.
10. The chip according to claim 9, characterized in that, The first and second dies are used to access data from the first partition and the on-chip high-bandwidth memory based on a visible unified physical address.
11. The chip according to claim 8 or 9, characterized in that, The chip also includes off-chip memory and input / output (IO) dies, with the first die and the second die connected to the off-chip memory via the IO dies; The first die and the second die are used to access data from the first partition, the on-chip high-bandwidth memory, and the off-chip memory based on a visible unified physical address.
12. The chip according to claim 11, characterized in that, The on-chip high-bandwidth memory is used to store memory pages whose memory page access frequency is greater than or equal to a preset threshold, and the off-chip memory is used to store memory pages whose memory page access frequency is less than the preset threshold.
13. A chip communication method, characterized in that, The chip is applied to a chip, which includes a first die, a second die, and a static random access memory (SRAM). The first die is a general-purpose computing die, and the second die is a domain-specific architecture (DSA) die. The first die and the second die are respectively connected to the SRAM. The SRAM stores a data structure for a first asynchronous task queue, which is used to manage tasks on the first die; the method includes: The first die inserts the first task into the first asynchronous task queue; The second die retrieves the first task from the first asynchronous task queue; The second die executes the first task and writes the first task result of the first task into the memory space corresponding to the first memory address of the SRAM; The second die sends the first memory address to the first die; The first die reads the first task result from the first partition based on the first memory address.
14. The method according to claim 13, characterized in that, The first die, the second die, and the SRAM are encapsulated in a bottom die. The chip also includes on-chip high-bandwidth memory and input / output (I / O) dies encapsulated in the bottom die. The first die and the second die are connected to the off-chip memory via the I / O dies. The method further includes: Determine the memory page access frequency of the on-chip high-bandwidth memory and the off-chip memory; Memory pages whose access frequency is greater than or equal to a preset threshold are stored in the on-chip high-bandwidth memory; Memory pages whose access frequency is less than the preset threshold are stored in the off-chip memory.
15. A chip communication device, characterized in that, The chip communication device includes the chip as described in any one of claims 1-12.
16. An electronic device, characterized in that, Includes the chip as described in any one of claims 1-12.