A data processing method and system based on 3D core particle integration and distributed access calculation, a terminal and a storage medium

CN122195927BActive Publication Date: 2026-09-22SHENZHEN MAITEXIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610614301.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-09-22
Estimated Expiration
2046-05-07

AI Technical Summary

Technical Problem

[0004]本发明的主要目的在于提供一种基于3D芯粒集成与分布式访存计算的数据处理方法、系统、终端及计算机可读存储介质,旨在解决现有技术中在处理大规模神经网络时,由于权重参数庞大、片外访存延迟高以及功耗大,导致大模型进行数据处理时的效率受到严重影响的问题

Benefits of technology

[0015]本发明中,获取多个芯片存储单元和混合键合,根据所述混合键合对多个所述芯片存储单元进行三维堆叠互连处理,得到三维互联结构;获取多个近存计算片,并根据多个所述近存计算片和所述三维互联结构构建多个存算一体芯粒;确定UCIE接口,通过所述UCIE接口对多个所述存算一体芯粒进行异构集成处理,得到高速互连网络;当接收到用户的访存指令时,根据所述访存指令确定所述高速互连网络中的目标计算单元和目标存储单元,并通过所述目标计算单元对所述目标存储单元进行访存读取处理和数据计算处理,得到目标计算结果。本发明通过将计算单元和存储单元进行垂直集成从而构建存算一体芯粒,并通过高速互联模块对存算一体芯粒进行异构集成处理,得到高速互连网络,实现了计算单元和存储单元之间以及存算一体芯粒之间的高效互联,形成了多芯粒协同的计算存储一体化架构,有效提高了大模型的带宽和计算效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195927B_ABST
    Figure CN122195927B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data processing, and discloses a data processing method, system, terminal and storage medium based on 3D core particle integration and distributed memory access calculation, which comprises the following steps: acquiring a plurality of chip storage units and hybrid bonding, performing three-dimensional stacking interconnection processing on the plurality of chip storage units according to the hybrid bonding, and obtaining a three-dimensional interconnection structure; acquiring a plurality of near-memory computing chips, and constructing a plurality of memory-compute integrated core particles together with the three-dimensional interconnection structure; determining a UCIE interface, performing heterogeneous integration processing on the plurality of memory-compute integrated core particles, and obtaining a high-speed interconnection network; when a memory access instruction is received, determining a target computing unit and a target storage unit in the high-speed interconnection network, and performing memory access reading processing and data calculation processing to obtain a target calculation result. Through the construction of the memory-compute integrated core particle and the high-speed interconnection network, a computing and storage integrated architecture of multiple core particles is formed, and the data processing efficiency of a large model is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data processing method, system, terminal, and computer-readable storage medium based on 3D chip integration and distributed memory access computing. Background Technology

[0002] In existing technologies, large model inference typically relies on general-purpose GPUs (Graphics Processing Units) or TPUs (Tensor Processing Units) for computation. These architectures are computation-centric, requiring frequent data transfer between storage and computation units, making memory access bandwidth a performance bottleneck. Especially when processing large-scale neural networks, the large number of weight parameters, high off-chip memory access latency, and high power consumption severely impact the efficiency of large model data processing.

[0003] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0004] The main objective of this invention is to provide a data processing method, system, terminal, and computer-readable storage medium based on 3D chip integration and distributed memory access computing. This invention aims to solve the problem that in the prior art, when processing large-scale neural networks, the efficiency of large models is severely affected by the large weight parameters, high off-chip memory access latency, and high power consumption.

[0005] To achieve the above objectives, the present invention provides a data processing method based on 3D chip integration and distributed memory access computing, the data processing method based on 3D chip integration and distributed memory access computing comprising the following steps: Multiple chip memory cells and hybrid bonding are obtained, and the multiple chip memory cells are subjected to three-dimensional stacking interconnection processing based on the hybrid bonding to obtain a three-dimensional interconnection structure. Multiple near-memory computing chips are acquired, and multiple in-memory computing chips are constructed based on the multiple near-memory computing chips and the three-dimensional interconnect structure; The UCIE interface is determined, and multiple in-memory computing chips are heterogeneously integrated through the UCIE interface to obtain a high-speed interconnect network. When a user's memory access instruction is received, the target computing unit and target storage unit in the high-speed interconnection network are determined according to the memory access instruction, and the target computing unit performs memory access reading and data calculation processing on the target storage unit to obtain the target calculation result.

[0006] Optionally, the data processing method based on 3D chip integration and distributed memory access computing, wherein obtaining multiple chip memory cells and hybrid bonding, and performing three-dimensional stacking interconnection processing on the multiple chip memory cells according to the hybrid bonding to obtain a three-dimensional interconnection structure, specifically includes: Multiple chip memory cells are obtained, and through-silicon vias are set in the multiple chip memory cells to obtain the vertical channels of the multiple chip memory cells; Copper pads in multiple chip memory cells are identified, and multiple hybrid bonds are obtained; Based on the multiple hybrid bonding and the copper pads, the multiple chip memory cells are vertically stacked and interconnected along the vertical channel to obtain a three-dimensional interconnect structure.

[0007] Optionally, the data processing method based on 3D chip integration and distributed memory access computing, wherein acquiring multiple near-memory computing chips and constructing multiple in-memory computing chips based on the multiple near-memory computing chips and the three-dimensional interconnect structure specifically includes: Multiple near-memory computing chips are obtained, each near-memory computing chip is placed below the three-dimensional interconnect structure, and the near-memory computing chips are vertically interconnected with the three-dimensional interconnect structure through the hybrid bonding to obtain multiple three-dimensional stacked structures; A packaging substrate is determined, and solder balls are used to solder the packaging substrate to the bottom of multiple three-dimensional stacked structures to obtain multiple in-memory computing chips.

[0008] Optionally, in the data processing method based on 3D chip integration and distributed memory access computing, each of the near-memory computing chips is a multi-core array architecture, including a dedicated cache unit and a direct memory access data scheduling unit; Access between each of the near-memory computing chips and each of the chip memory cells in the three-dimensional interconnect structure is independent of each other; The computation results are transmitted between different near-memory computing chips via pulse transmission.

[0009] Optionally, in the data processing method based on 3D chip integration and distributed memory access computing, the step of determining the UCIE interface and performing heterogeneous integration processing on multiple in-memory computing chips through the UCIE interface to obtain a high-speed interconnect network specifically involves: The UCIE interface is determined, and the UCIE interconnection protocol is used to interconnect multiple in-memory computing chips according to the UCIE interface to obtain a high-speed interconnection network; The UCIE interface employs a standardized communication protocol and high-density physical interconnect technology.

[0010] Optionally, the data processing method based on 3D chip integration and distributed memory access computing, wherein when a user's memory access instruction is received, the target computing unit and target storage unit in the high-speed interconnect network are determined according to the memory access instruction, and the target computing unit performs memory access reading and data computing processing on the target storage unit to obtain the target computing result, specifically includes: When a user's memory access instruction is received, the target in-memory computing chip and the target computing unit in the high-speed interconnect network are determined according to the memory access instruction. The target storage unit is determined by performing a proximity search on the target computing unit within the target in-memory computing chip. The target computing unit reads and performs data calculations on the target storage unit according to the vertical channel to obtain the target calculation result; The target calculation result is cascaded and transmitted between the near-memory computing chips using a data stream mode, and then the target calculation result is output.

[0011] Optionally, in the data processing method based on 3D chip integration and distributed memory access computing, the high-speed interconnect network integrates a hardware-aware task scheduler and a network status monitoring unit. The network status monitoring unit is used to perform data request pattern analysis on the near-memory computing chip to obtain the target data request pattern, and to perform access load analysis on the chip storage unit to obtain the access load analysis results. The hardware-aware task scheduler is used to perform optimal interconnection link analysis, link width adjustment, and transmission protocol adjustment based on the target data request pattern and the access load analysis results, so as to obtain the target network resource configuration results.

[0012] Furthermore, to achieve the above objectives, the present invention also provides a data processing system based on 3D chip integration and distributed memory access computing, wherein the data processing system based on 3D chip integration and distributed memory access computing includes: A three-dimensional stacked interconnect module is used to acquire multiple chip memory cells and hybrid bonding, and to perform three-dimensional stacked interconnect processing on the multiple chip memory cells according to the hybrid bonding to obtain a three-dimensional interconnect structure. The in-memory computing chip construction module is used to acquire multiple near-memory computing chips and construct multiple in-memory computing chips based on the multiple near-memory computing chips and the three-dimensional interconnect structure. A high-speed interconnect network construction module is used to determine the UCIE interface and perform heterogeneous integration processing on multiple in-memory computing chips through the UCIE interface to obtain a high-speed interconnect network. The target calculation result output module is used to determine the target computing unit and target storage unit in the high-speed interconnection network according to the user's memory access instruction when the user's memory access instruction is received, and to perform memory access reading and data calculation processing on the target storage unit through the target computing unit to obtain the target calculation result.

[0013] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a data processing program based on 3D chip integration and distributed memory access computing stored in the memory and executable on the processor, wherein when the data processing program based on 3D chip integration and distributed memory access computing is executed by the processor, it implements the steps of the data processing method based on 3D chip integration and distributed memory access computing as described above.

[0014] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a data processing program based on 3D chip integration and distributed memory access computing, and when the data processing program based on 3D chip integration and distributed memory access computing is executed by a processor, it implements the steps of the data processing method based on 3D chip integration and distributed memory access computing as described above.

[0015] In this invention, multiple chip storage units and hybrid bonding are obtained. Based on the hybrid bonding, the multiple chip storage units are subjected to three-dimensional stacking interconnection processing to obtain a three-dimensional interconnect structure. Multiple near-memory computing chips are obtained, and multiple in-memory computing granules are constructed based on the multiple near-memory computing chips and the three-dimensional interconnect structure. A UCIE interface is determined, and the multiple in-memory computing granules are heterogeneously integrated through the UCIE interface to obtain a high-speed interconnect network. When a user's memory access command is received, the target computing unit and target storage unit in the high-speed interconnect network are determined according to the memory access command. The target computing unit performs memory access and data computation processing on the target storage unit to obtain the target computation result. This invention constructs in-memory computing granules by vertically integrating computing units and storage units, and performs heterogeneous integration processing on the in-memory computing granules through a high-speed interconnect module to obtain a high-speed interconnect network. This achieves efficient interconnection between computing units and storage units, as well as between in-memory computing granules, forming a multi-granule collaborative computing and storage integrated architecture, effectively improving the bandwidth and computational efficiency of large models. Attached Figure Description

[0016] Figure 1 This is a flowchart of a preferred embodiment of the data processing method based on 3D chip integration and distributed memory access computing of the present invention; Figure 2This is a schematic diagram of the computing unit and storage unit vertically stacked in 3D according to a preferred embodiment of the data processing method based on 3D chip integration and distributed memory access computing of the present invention; Figure 3 This is a schematic diagram of the distributed architecture and vertical storage unit directly connected in the computing unit of the preferred embodiment of the data processing method based on 3D chip integration and distributed memory access computing of the present invention. Figure 4 This is a schematic diagram of the distributed mesh architecture and NoC on-chip network structure in the computing unit of a preferred embodiment of the data processing method based on 3D chip integration and distributed memory access computing of the present invention. Figure 5 This is a schematic diagram of multiple 3D integrated in-memory computing chips interconnected via UCIE, which is a preferred embodiment of the data processing method based on 3D chip integration and distributed memory access computing of the present invention. Figure 6 This is a structural diagram of a preferred embodiment of the data processing system based on 3D chip integration and distributed memory access computing of the present invention; Figure 7 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0018] In existing technologies, large model inference typically relies on general-purpose GPUs (Graphics Processing Units) or TPUs (Tensor Processing Units) for computation. These computation-centric architectures require frequent data transfer between storage and computation units, making memory access bandwidth a performance bottleneck. This is particularly problematic when processing large-scale neural networks, where the sheer size of weights, high off-chip memory latency, and high power consumption severely impact inference efficiency. Furthermore, existing accelerators often employ discrete memory architectures, which struggle to meet the demands for high-bandwidth, low-latency data delivery, hindering overall system performance improvement. To alleviate memory access bottlenecks, some solutions introduce high-bandwidth memory (HBM, often referred to as stacked memory) or near-memory computing architectures. However, these remain limited by the interconnect density and heat dissipation efficiency of two-dimensional planar integration, making efficient collaboration between storage and computational resources difficult. Simultaneously, as the scale of model parameters continues to grow, the traditional PCB (Printed Circuit Board) level interconnect bandwidth cannot match the data exchange requirements between chips, leading to a significant increase in communication overhead during multi-GPU expansion. In addition, existing 3D stacking technologies are mostly focused on the vertical integration of single functional units, lacking full-stack collaborative optimization of computing, storage and interconnection, which limits the further improvement of system energy efficiency ratio.

[0019] To address the aforementioned issues, this invention proposes a distributed interconnect memory access computing large-model inference card based on 3D chiplet integration, applicable to large-model inference computing. This invention achieves distributed memory access computing inference for large models by vertically integrating computing and storage units and heterogeneously integrating them at the board level in chiplet form with the industry-standard high-speed interconnect module (UCIE, Universal Chiplet Interconnect Express, an open high-speed interconnect standard between chips, generally referred to as Universal Chiplet Interconnect). This process achieves efficient interconnection, forming a multi-chiplet collaborative computing and storage integrated architecture. This architecture utilizes 3D stacking technology to shorten data paths, significantly reducing memory access latency and power consumption while increasing bandwidth density. Furthermore, high-speed interconnection between chips is achieved through the UCIE standard interface, supporting flexible expansion and heterogeneous integration, effectively alleviating the bandwidth bottleneck problem of traditional PCB interconnects. This invention also introduces a distributed memory access control mechanism to dynamically schedule the access of each computing unit to local and adjacent storage resources, avoid global memory contention, and improve parallel inference efficiency. At the same time, the computing unit and the vertically stacked storage unit adopt direct data access, avoiding complex on-chip network (NoC) routing and improving effective bandwidth and computing efficiency.

[0020] This invention achieves deep collaborative optimization of computing, storage, and communication, significantly improving the throughput and energy efficiency of large model inference, making it suitable for low-latency deployment scenarios with models boasting hundreds of billions of parameters. Simultaneously, this inference card supports multi-card cascading and cluster deployment, leveraging a low-latency interconnect network between chips to achieve efficient cross-device data synchronization, further enhancing system scalability and fault tolerance. Combined with dynamic voltage and frequency regulation technology, it intelligently adjusts power consumption modes according to load changes, optimizing energy efficiency while ensuring performance.

[0021] In practical applications, the architecture of this invention can be adapted to a variety of mainstream large model structures, including Transformer and MoE, demonstrating good versatility and compatibility, and providing solid support for the efficient deployment of large models at the edge and in the cloud.

[0022] In addition, the inference card set up in this invention incorporates a reserved expansion interface based on silicon photonics interconnect technology, which can support optical signal transmission between chips in the future, further breaking through the limits of bandwidth and energy efficiency, and providing a sustainable hardware foundation for the next generation of ultra-large-scale model inference.

[0023] The data processing method based on 3D chip integration and distributed memory access computing described in the preferred embodiment of the present invention, such as... Figure 1 As shown, the data processing method based on 3D chip integration and distributed memory access computing includes the following steps: Step S10: Obtain multiple chip memory cells and hybrid bonding, and perform three-dimensional stacking interconnection processing on the multiple chip memory cells according to the hybrid bonding to obtain a three-dimensional interconnection structure.

[0024] The construction of each three-dimensional interconnect structure requires four layers of chip storage units (i.e. Figure 2 The process involves vertically integrating four layers of memory cells (DRAM chips) and hybrid bonding. First, TSV vias are created within these four layers of memory cells. These TSV vias are then interconnected using hybrid bonding to construct a three-dimensional interconnect structure.

[0025] Specifically, multiple chip memory cells are obtained, and through-silicon vias are set in the multiple chip memory cells to obtain vertical channels of the multiple chip memory cells; copper pads in the multiple chip memory cells are determined, and multiple hybrid bonds are obtained; according to the multiple hybrid bonds and the copper pads, the multiple chip memory cells are subjected to three-dimensional vertical stacking and interconnection processing along the vertical channels to obtain a three-dimensional interconnect structure.

[0026] like Figure 2 As shown, the computational unit (PIM Die, i.e.) is described in detail. Figure 2 The near-memory computing chip and the storage unit (i.e., the chip storage unit in this invention, 4-layer CacheRAMDie, that is...) Figure 2 This is a three-dimensional stacked structure formed by integrating DRAM chips within the chip. This integration process involves first creating through-silicon vias (TSVs) within the chip. Figure 2 The TSV vias form vertical channels, and hybrid bonding is used to directly bond the copper pads of multiple chips face to face, thereby achieving high-density, short-distance three-dimensional stacked interconnects. This interconnect structure allows the near-in-memory (PIM) die to directly access adjacent memory layers, realizing an in-memory computing dataflow design. This design significantly reduces data transmission latency and improves bandwidth utilization.

[0027] Step S20: Obtain multiple near-memory computing chips, and construct multiple in-memory computing chips based on the multiple near-memory computing chips and the three-dimensional interconnect structure.

[0028] Each in-memory computing chip requires a three-dimensional interconnect structure and a near-memory computing chip to be constructed, such as Figure 2 As shown, the three-dimensional interconnect structure and the near-memory computing chip are also vertically distributed, with the near-memory computing chip positioned below the three-dimensional interconnect structure. The three-dimensional interconnect structure and the near-memory computing chip are connected via hybrid bonding, and the near-memory computing chip is connected to the packaging substrate below via solder balls, thereby constructing an in-memory computing chip. This invention enables the interconnection between the near-memory computing chip and the memory cells in the three-dimensional interconnect structure, allowing the near-memory computing chip to directly read the stored data from the memory cells in the three-dimensional interconnect structure.

[0029] Specifically, multiple near-memory computing chips are obtained, each near-memory computing chip is placed below the three-dimensional interconnect structure, and the near-memory computing chips are vertically interconnected with the three-dimensional interconnect structure through hybrid bonding to obtain multiple three-dimensional stacked structures; a packaging substrate is determined, and solder balls are used to solder the packaging substrate to the bottom of the multiple three-dimensional stacked structures to obtain multiple in-memory computing chips.

[0030] Furthermore, each of the near-memory computing chips is a multi-core array architecture, including a dedicated cache unit and a direct memory access data scheduling unit; the access between each of the near-memory computing chips and each of the chip memory units in the three-dimensional interconnect structure is independent of each other; the computing results are transmitted between different near-memory computing chips through pulse transmission.

[0031] like Figure 3 As shown, the multi-core array architecture inside the PIM Die is illustrated in detail. Each computing unit (PE, i.e., PIM Die) is equipped with a dedicated cache (i.e., ... Figure 3 SRAM in the memory and direct memory access data scheduling unit (i.e., SRAM) and direct memory access data scheduling unit (i.e. Figure 3The DMA in the stack is used to support fine-grained parallel computing and DMA memory access operations for group vector pulsation (i.e., the present invention can directly access memory through DMA without CPU involvement). Each PE is directly interconnected with the stacked DRMA through the nearest TSV to achieve high-density vertical interconnection (interconnection with the memory cell), ensuring that data flows efficiently between the computing unit and the stacked memory layer.

[0032] like Figure 4 The diagram illustrates the difference between NoC interconnects and distributed meshes. Figure 4 As shown in 'a', this is the NoC on-chip network architecture. The NoC connects each PE and 3D DRAM (such as...) through a routing network. Figure 4 In the 3D IO (of the PE), each PE can access the 3D IO of different areas through routing. However, NoC needs to handle extremely complex data paths, which can easily lead to data congestion when multiple PEs are working simultaneously, which is not conducive to improving bandwidth utilization.

[0033] like Figure 4 As shown in b, this is the distributed mesh in this invention. In contrast, the distributed architecture set up in this invention simplifies the data path by directly connecting the PE and 3D DRAM. Each PE accesses 3D IO independently, and the PEs transmit computation results in a pulsed manner. This distributed architecture effectively avoids the problems of complex data paths and helps improve bandwidth utilization. At the same time, it simplifies the data transmission path, reduces data congestion, and thus improves the overall computing efficiency.

[0034] Step S30: Determine the UCIE interface, and perform heterogeneous integration processing on multiple in-memory computing chips through the UCIE interface to obtain a high-speed interconnect network.

[0035] Once a single in-memory computing chip is constructed, multiple in-memory computing chips can be packaged together on a substrate. Each in-memory computing chip has a corresponding UCIE interface, which enables high-speed interconnection between chips.

[0036] Specifically, a UCIE interface is determined, and multiple in-memory computing chips are interconnected and communicated using the UCIE interconnection protocol according to the UCIE interface to obtain a high-speed interconnection network; wherein, the UCIE interface adopts a standardized communication protocol and high-density physical interconnection technology.

[0037] like Figure 5 As shown, the in-memory computing chip based on 3D integration technology forms a high-speed interconnect network through the UCIE (Universal Chiplet Interconnect Express) interconnect protocol, realizing the elastic expansion and collaborative scheduling of storage and computing resources.

[0038] The UCIE interface employs standardized communication protocols and high-density physical interconnect technology to ensure nanosecond-level latency and TB / s-level bandwidth for cross-chip data transmission, providing underlying architectural support for building large-scale heterogeneous computing systems. This architecture supports dynamic topology reconfiguration at runtime, adaptively adjusting to the bandwidth and latency requirements of computing tasks.

[0039] Step S40: When a user's memory access instruction is received, the target computing unit and the target storage unit in the high-speed interconnection network are determined according to the memory access instruction, and the target computing unit performs memory access reading and data calculation processing on the target storage unit to obtain the target calculation result.

[0040] Specifically, when a user's memory access command is received, the target in-memory computing chip and target computing unit in the high-speed interconnect network are determined according to the memory access command; the target computing unit in the target in-memory computing chip performs a proximity lookup to determine the target storage unit; the target computing unit reads data and performs data calculation on the target storage unit according to the vertical channel to obtain the target calculation result; the target calculation result is cascaded and transmitted between the proximity computing chips using a data stream mode, and the target calculation result is output.

[0041] Understandably, the timing coordination mechanism for pulsed transmission between PEs in a distributed mesh architecture relies on a globally synchronized clock cycle. Each PE operates strictly in a locked-step manner within each preset cycle, following the "receive-compute-send" steps, achieving overlapping scheduling between computation and communication. The specific process of its preset pipeline cycle is as follows: at the beginning of the cycle, the PE receives data from upstream, performs local computation, and at the end of the same cycle, sends the result calculated in the previous cycle to its downstream neighbor. In this way, multiple data blocks are processed continuously in the pulsed chain like a flowing stream, effectively hiding data transmission latency.

[0042] Furthermore, the high-speed interconnect network integrates a hardware-aware task scheduler and a network status monitoring unit; the network status monitoring unit is used to perform data request pattern analysis on the near-memory computing chip to obtain the target data request pattern, and to perform access load analysis on the chip storage unit to obtain the access load analysis result; the hardware-aware task scheduler is used to perform optimal interconnect link analysis, link width adjustment, and transmission protocol adjustment based on the target data request pattern and the access load analysis result to obtain the target network resource configuration result.

[0043] The core mechanism of this invention lies in the integration of a hardware-aware task scheduler and a network status monitoring unit, which can analyze the data request patterns of each computing chip and the access load of the storage chip in real time. Based on this information, the system dynamically enables or bypasses specific interconnect links and adjusts the link width and transmission protocol (such as switching to low-latency mode or high-bandwidth mode), thereby achieving dynamic optimization of network resource configuration within the task execution cycle to accurately match the differentiated needs of different scenarios, from high-throughput model inference to low-latency real-time decision-making.

[0044] Furthermore, in high-parallelism computing scenarios such as matrix multiplication, each chip achieves collaborative operation through a distributed computing array architecture. Data is cascaded and transmitted according to the data flow pattern of the pulsating array, making full use of computational locality and data reuse characteristics, and significantly reducing the pressure on global storage access.

[0045] Experimental data demonstrate that the sub-nanosecond latency of UCIE interconnects enables cross-chip scheduling efficiency approaching that of on-chip communication, resulting in a significant improvement in system-level energy efficiency. Particularly in deep learning inference tasks, this architecture effectively supports the efficient deployment of large-scale models through fine-grained resource allocation and low-latency communication. Combined with dynamic voltage and frequency regulation technology, the system can maintain optimal energy efficiency under different loads, meeting the needs of diverse application scenarios.

[0046] Furthermore, this architecture further enhances instruction-level parallelism efficiency through fine-grained hardware pipeline optimization and compiler-co-scheduled execution. Leveraging the high-bandwidth memory stacking advantage of 3D integration, each core can access multi-layer storage resources locally, forming a hierarchical caching system and reducing the overhead of remote data migration. In typical model inference such as MoE, the sparse activation characteristics of each expert subnetwork are highly compatible with the locality advantage of the distributed architecture. Computational units corresponding to inactive subnetworks can dynamically enter a low-power state, while active subnetworks achieve rapid parameter loading and result aggregation through UCIE high-speed links. The 3D stacked storage layer provides high-bandwidth weighted read capabilities for gated routing logic, ensuring millisecond-level dynamic path selection accuracy. Actual deployment shows that this architecture can still maintain over 80% computational resource utilization under models with hundreds of billions of parameters, reducing end-to-end inference latency by 60% compared to traditional architectures, while simultaneously reducing power consumption by 45%.

[0047] Furthermore, by combining the reserved interface for optoelectronic hybrid interconnects, micro-optical channels can be deployed between adjacent computing chips in the future to achieve ultra-low power consumption and ultra-high bandwidth chip-level signal transmission, providing a sustainable upgrade path for large-scale model inference.

[0048] Furthermore, the key difference between this invention and existing technologies lies in the following: This invention achieves sub-nanosecond-level communication latency between chips based on the UCIE standard, and supports operation by combining the collaborative design of 3D stacked storage and distributed computing arrays. In other words, this invention achieves sub-nanosecond-level communication between chips through the UCIE interconnect protocol, combining the high bandwidth characteristics of 3D stacked storage with the tightly coupled design of distributed computing units. During runtime, it dynamically reconstructs the computing topology based on model sparsity, improving the matching degree between data flow and computing resources. Through the collaboration of hardware-level fine-grained power management and compiler-driven load prediction mechanisms, it achieves joint scheduling of computing, storage, and communication resources, maximizing energy efficiency while ensuring QoS. The architecture of this invention supports multi-instance parallel inference and dynamic expert routing, and significantly reduces gating decision overhead by combining the low latency characteristics of UCIE links. In summary, the key protected technical solutions of this invention include: 1. Sub-nanosecond-level communication between chips based on UCIE and collaborative 3D high-bandwidth storage access; 2. Dynamic sparsity-aware computing topology reconstruction method; 3. Optoelectronic hybrid interconnect reserved architecture for future expansion of ultra-low-power optical links.

[0049] Possible design changes or modifications to this invention are as follows: For example, similar latency performance can be achieved using high-speed serial interconnects with non-UCIE protocols, or in-memory computing architectures can replace 3D stacked storage to reduce data transfer overhead. Furthermore, dynamic topology reconfiguration logic can be migrated to the software stack, and path scheduling can be implemented through firmware updates, reducing hardware complexity. Heterogeneous chip combinations can also be introduced, such as integrating analog computing units to execute specific operators, improving energy efficiency. The optoelectronic hybrid interface can also be extended to an all-optical interconnect, further increasing bandwidth density; these variations still fall within the core technology boundaries of this invention.

[0050] Furthermore, such as Figure 6 As shown, based on the above-described data processing method based on 3D chip integration and distributed memory access computing, the present invention also provides a data processing system based on 3D chip integration and distributed memory access computing, wherein the data processing system based on 3D chip integration and distributed memory access computing includes: The three-dimensional stacked interconnect module 51 is used to acquire multiple chip memory cells and hybrid bonding, and to perform three-dimensional stacked interconnect processing on the multiple chip memory cells according to the hybrid bonding to obtain a three-dimensional interconnect structure. The in-memory computing chip construction module 52 is used to acquire multiple near-memory computing chips and construct multiple in-memory computing chips based on the multiple near-memory computing chips and the three-dimensional interconnect structure. High-speed interconnect network construction module 53 is used to determine the UCIE interface and perform heterogeneous integration processing on multiple in-memory computing chips through the UCIE interface to obtain a high-speed interconnect network. The target calculation result output module 54 is used to determine the target computing unit and the target storage unit in the high-speed interconnection network according to the user's memory access instruction when the user's memory access instruction is received, and to perform memory access reading and data calculation processing on the target storage unit through the target computing unit to obtain the target calculation result.

[0051] Furthermore, such as Figure 7 As shown, based on the above-mentioned data processing method and system based on 3D chip integration and distributed memory access computing, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 7 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0052] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a data processing program 40 based on 3D chip integration and distributed memory access computing. This data processing program 40 can be executed by the processor 10 to implement the data processing method based on 3D chip integration and distributed memory access computing in this application.

[0053] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the data processing method based on 3D chip integration and distributed memory access computing.

[0054] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface.

[0055] In one embodiment, when the processor 10 executes the data processing program 40 based on 3D chip integration and distributed memory access computing in the memory 20, the following steps are performed: Multiple chip memory cells and hybrid bonding are obtained, and the multiple chip memory cells are subjected to three-dimensional stacking interconnection processing based on the hybrid bonding to obtain a three-dimensional interconnection structure. Multiple near-memory computing chips are acquired, and multiple in-memory computing chips are constructed based on the multiple near-memory computing chips and the three-dimensional interconnect structure; The UCIE interface is determined, and multiple in-memory computing chips are heterogeneously integrated through the UCIE interface to obtain a high-speed interconnect network. When a user's memory access instruction is received, the target computing unit and target storage unit in the high-speed interconnection network are determined according to the memory access instruction, and the target computing unit performs memory access reading and data calculation processing on the target storage unit to obtain the target calculation result.

[0056] The step of acquiring multiple chip memory cells and hybrid bonding, and performing three-dimensional stacking interconnection processing on the multiple chip memory cells based on the hybrid bonding to obtain a three-dimensional interconnection structure, specifically includes: Multiple chip memory cells are obtained, and through-silicon vias are set in the multiple chip memory cells to obtain the vertical channels of the multiple chip memory cells; Copper pads in multiple chip memory cells are identified, and multiple hybrid bonds are obtained; Based on the multiple hybrid bonding and the copper pads, the multiple chip memory cells are vertically stacked and interconnected along the vertical channel to obtain a three-dimensional interconnect structure.

[0057] The step of acquiring multiple near-memory computing chips and constructing multiple in-memory computing chips based on the multiple near-memory computing chips and the three-dimensional interconnect structure specifically includes: Multiple near-memory computing chips are obtained, each near-memory computing chip is placed below the three-dimensional interconnect structure, and the near-memory computing chips are vertically interconnected with the three-dimensional interconnect structure through the hybrid bonding to obtain multiple three-dimensional stacked structures; A packaging substrate is determined, and solder balls are used to solder the packaging substrate to the bottom of multiple three-dimensional stacked structures to obtain multiple in-memory computing chips.

[0058] Each of the near-memory computing chips is a multi-core array architecture, including a dedicated cache unit and a direct memory access data scheduling unit; Access between each of the near-memory computing chips and each of the chip memory cells in the three-dimensional interconnect structure is independent of each other; The computation results are transmitted between different near-memory computing chips via pulse transmission.

[0059] Specifically, determining the UCIE interface and performing heterogeneous integration of multiple in-memory computing chips through the UCIE interface to obtain a high-speed interconnect network involves: The UCIE interface is determined, and the UCIE interconnection protocol is used to interconnect multiple in-memory computing chips according to the UCIE interface to obtain a high-speed interconnection network; The UCIE interface employs a standardized communication protocol and high-density physical interconnect technology.

[0060] Specifically, when a user's memory access instruction is received, the target computing unit and target storage unit in the high-speed interconnect network are determined according to the memory access instruction, and the target computing unit performs memory access and data calculation processing on the target storage unit to obtain the target calculation result. When a user's memory access instruction is received, the target in-memory computing chip and the target computing unit in the high-speed interconnect network are determined according to the memory access instruction. The target storage unit is determined by performing a proximity search on the target computing unit within the target in-memory computing chip. The target computing unit reads and performs data calculations on the target storage unit according to the vertical channel to obtain the target calculation result; The target calculation result is cascaded and transmitted between the near-memory computing chips using a data stream mode, and then the target calculation result is output.

[0061] The high-speed interconnection network integrates a hardware-aware task scheduler and a network status monitoring unit. The network status monitoring unit is used to perform data request pattern analysis on the near-memory computing chip to obtain the target data request pattern, and to perform access load analysis on the chip storage unit to obtain the access load analysis results. The hardware-aware task scheduler is used to perform optimal interconnection link analysis, link width adjustment, and transmission protocol adjustment based on the target data request pattern and the access load analysis results, so as to obtain the target network resource configuration results.

[0062] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a data processing program based on 3D chip integration and distributed memory access computing, and the data processing program based on 3D chip integration and distributed memory access computing, when executed by a processor, implements the steps of the data processing method based on 3D chip integration and distributed memory access computing as described above.

[0063] In summary, this invention provides a data processing method, system, terminal, and storage medium based on 3D chip integration and distributed memory access computing. The method includes: acquiring multiple chip memory cells and hybrid bonding; performing three-dimensional stacking interconnection processing on the multiple chip memory cells according to the hybrid bonding to obtain a three-dimensional interconnect structure; acquiring multiple near-memory computing chips and constructing multiple in-memory computing chips based on the multiple near-memory computing chips and the three-dimensional interconnect structure; determining a UCIE interface and performing heterogeneous integration processing on the multiple in-memory computing chips through the UCIE interface to obtain a high-speed interconnect network; when receiving a user's memory access command, determining the target computing unit and target memory unit in the high-speed interconnect network according to the memory access command, and performing memory access reading processing and data computing processing on the target memory unit through the target computing unit to obtain the target computing result. This invention constructs an in-memory computing chip by vertically integrating computing units and storage units, and performs heterogeneous integration of the in-memory computing chip through a high-speed interconnect module to obtain a high-speed interconnect network. This achieves efficient interconnection between computing units and storage units, as well as between in-memory computing chips, forming a multi-chip collaborative computing and storage integrated architecture, which effectively improves the bandwidth and computing efficiency of large models.

[0064] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0065] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0066] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A data processing method based on 3D chip integration and distributed memory access computing, characterized in that, The data processing method based on 3D chip integration and distributed memory access computing includes: Multiple chip memory cells and hybrid bonding are obtained, and the multiple chip memory cells are subjected to three-dimensional stacking interconnection processing based on the hybrid bonding to obtain a three-dimensional interconnection structure. Multiple near-memory computing chips are acquired, and multiple in-memory computing chips are constructed based on the multiple near-memory computing chips and the three-dimensional interconnect structure; Each of the near-memory compute chips is a multi-core array architecture, including a dedicated cache unit and a direct memory access data scheduling unit; Access between each of the near-memory computing chips and each of the chip memory cells in the three-dimensional interconnect structure is independent of each other; The computation results are transmitted between different near-memory computing chips via pulse transmission; Configure a distributed architecture to directly connect PE and 3D DRAM; The UCIE interface is determined, and multiple in-memory computing chips are heterogeneously integrated through the UCIE interface to obtain a high-speed interconnect network. The process of determining the UCIE interface and then heterogeneously integrating multiple in-memory computing chips through the UCIE interface to obtain a high-speed interconnect network is as follows: The UCIE interface is determined, and the UCIE interconnection protocol is used to interconnect multiple in-memory computing chips according to the UCIE interface to obtain a high-speed interconnection network; The UCIE interface uses a standardized communication protocol and high-density physical interconnect technology. The high-speed interconnection network integrates a hardware-aware task scheduler and a network status monitoring unit. The network status monitoring unit is used to perform data request pattern analysis on the near-memory computing chip to obtain the target data request pattern, and to perform access load analysis on the chip storage unit to obtain the access load analysis results. The hardware-aware task scheduler is used to perform optimal interconnection link analysis, link width adjustment, and transmission protocol adjustment based on the target data request pattern and the access load analysis results, so as to obtain the target network resource configuration results. Real-time analysis of data request patterns and access load of storage cores for each computing core; dynamic activation or bypass of specific interconnect links; and adjustment of link width and transmission protocol; dynamic optimization of network resource configuration within the task execution cycle to accurately match differentiated needs from high-throughput model inference to low-latency real-time decision-making. When a user's memory access instruction is received, the target computing unit and target storage unit in the high-speed interconnection network are determined according to the memory access instruction, and the target computing unit performs memory access reading and data calculation processing on the target storage unit to obtain the target calculation result.

2. The data processing method based on 3D chip integration and distributed memory access computing according to claim 1, characterized in that, The step of acquiring multiple chip memory cells and hybrid bonding, and performing three-dimensional stacking interconnection processing on the multiple chip memory cells based on the hybrid bonding to obtain a three-dimensional interconnection structure, specifically includes: Multiple chip memory cells are obtained, and through-silicon vias are set in the multiple chip memory cells to obtain the vertical channels of the multiple chip memory cells; Copper pads in multiple chip memory cells are identified, and multiple hybrid bonds are obtained; Based on the multiple hybrid bonding and the copper pads, the multiple chip memory cells are vertically stacked and interconnected along the vertical channel to obtain a three-dimensional interconnect structure.

3. The data processing method based on 3D chip integration and distributed memory access computing according to claim 1, characterized in that, The process of acquiring multiple near-memory computing chips and constructing multiple in-memory computing chips based on the multiple near-memory computing chips and the three-dimensional interconnect structure specifically includes: Multiple near-memory computing chips are obtained, each near-memory computing chip is placed below the three-dimensional interconnect structure, and the near-memory computing chips are vertically interconnected with the three-dimensional interconnect structure through the hybrid bonding to obtain multiple three-dimensional stacked structures; A packaging substrate is determined, and solder balls are used to solder the packaging substrate to the bottom of multiple three-dimensional stacked structures to obtain multiple in-memory computing chips.

4. The data processing method based on 3D chip integration and distributed memory access computing according to claim 2, characterized in that, When a user's memory access instruction is received, the target computing unit and target storage unit in the high-speed interconnection network are determined according to the memory access instruction. Then, the target computing unit performs memory access and data computation processing on the target storage unit to obtain the target computation result. Specifically, this includes: When a user's memory access instruction is received, the target in-memory computing chip and the target computing unit in the high-speed interconnect network are determined according to the memory access instruction. The target storage unit is determined by performing a proximity search on the target computing unit within the target in-memory computing chip. The target computing unit reads and performs data calculations on the target storage unit according to the vertical channel to obtain the target calculation result; The target calculation result is cascaded and transmitted between the near-memory computing chips using a data stream mode, and then the target calculation result is output.

5. A data processing system based on 3D chip integration and distributed memory access computing, characterized in that, The data processing system based on 3D chip integration and distributed memory access computing is used to implement the data processing method based on 3D chip integration and distributed memory access computing as described in any one of claims 1-4, wherein the data processing system based on 3D chip integration and distributed memory access computing includes: A three-dimensional stacked interconnect module is used to acquire multiple chip memory cells and hybrid bonding, and to perform three-dimensional stacked interconnect processing on the multiple chip memory cells according to the hybrid bonding to obtain a three-dimensional interconnect structure. The in-memory computing chip construction module is used to acquire multiple near-memory computing chips and construct multiple in-memory computing chips based on the multiple near-memory computing chips and the three-dimensional interconnect structure. A high-speed interconnect network construction module is used to determine the UCIE interface and perform heterogeneous integration processing on multiple in-memory computing chips through the UCIE interface to obtain a high-speed interconnect network. The target calculation result output module is used to determine the target computing unit and target storage unit in the high-speed interconnection network according to the user's memory access instruction when the user's memory access instruction is received, and to perform memory access reading and data calculation processing on the target storage unit through the target computing unit to obtain the target calculation result.

6. A terminal, characterized in that, The terminal includes: a memory, a processor, and a data processing program based on 3D chip integration and distributed memory access computing stored in the memory and executable on the processor. When the data processing program based on 3D chip integration and distributed memory access computing is executed by the processor, it implements the steps of the data processing method based on 3D chip integration and distributed memory access computing as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a data processing program based on 3D chip integration and distributed memory access computing. When the data processing program based on 3D chip integration and distributed memory access computing is executed by a processor, it implements the steps of the data processing method based on 3D chip integration and distributed memory access computing as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Three-dimensional stacked near-memory computing architecture, chip and data access method

    CN121743246A