Near memory computing system, near memory computing intelligent accelerator card and intelligent equipment
By adopting 3D core-particle integration scheme and distributed storage design in near-memory computing systems, the problems of limited bandwidth, slow inference speed and high data access power consumption are solved for the deployment of the end-side of large language models, and more efficient data processing and lower power consumption are achieved.
Patent Information
- Application Number
- CN202510087276.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-13
AI Technical Summary
The deployment of large language models end-side deployments have problems such as bandwidth limitation, slow inference speed and high data access power consumption. The existing near-storage computing architecture adopts 2D or 2.5D integration, resulting in bandwidth and processing speed problems.
The 3D core-particle integration solution is adopted, and a near-memory computing system is formed through external interfaces, in-memory processing computing arrays and memory modules. The memory modules and in-memory processing computing arrays adopt a distributed design to realize distributed storage of weighted data and processing of static input data.
Improves data bandwidth, reduces latency, improves the overall performance of the hardware architecture, achieves higher throughput performance and faster data processing speeds, while reducing power consumption.
Smart Images

Figure CN120144527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of near-memory computing, and particularly to a near-memory computing system, a near-memory computing embodied intelligence acceleration card, and an intelligent device. Background Art
[0002] Currently, we are in an era of explosive growth of large language models (LLMs). These models play an important role in integrating into human life and improving efficiency, but at the same time, they also face major challenges in computing performance and resources. The training and inference of large language models usually require a large amount of data, especially frequent access to input data and model parameters, and their demand for memory capacity and storage bandwidth (such as Dynamic Random Access Memory, DRAM) exceeds the capabilities of the current largest-scale processors and memory systems. Therefore, there are problems of limited bandwidth, slow inference speed, and high power consumption for data access in the end-side deployment of large models.
[0003] To address the bottleneck brought about by the separation of storage and computing in traditional computing models, various in-memory computing architectures have been proposed and developed. However, existing near-memory computing (Process-In-Memory, PIM) architectures still adopt 2D or 2.5D integration methods. Although the power consumption problem is solved, the problem of limited bandwidth still exists. Moreover, as the integration density increases, more and more components are placed in a limited space, which leads to a decline in signal integrity and an increase in latency, making the processing speed of near-memory computing slower. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide a near-memory computing system, a near-memory computing embodied intelligence acceleration card, and an intelligent device to solve the problems of limited bandwidth, slow inference speed, and high power consumption for data access in the end-side deployment of large models, and at the same time adopt a 3D die integration solution to further solve the bandwidth and processing speed problems of 2D or 2.5D integrated PIM.
[0005] The technical solution of the present invention is as follows:
[0006] A near-memory computing system, comprising:
[0007] An external interface for connecting to an external device;
[0008] An in-memory processing computing array connected to the external interface, the in-memory processing computing array being configured to process large language model data output by an external device and output a processing result;
[0009] A memory module, connected to the in-memory processing computing array, is also connected to the external interface. The memory module stores the computing weights of the in-memory processing computing array, and is used to store the processing results output by the in-memory processing computing array, as well as the large language model data output by external devices. The memory module and the in-memory processing computing array adopt a distributed design.
[0010] Optionally, the in-memory processing computing array includes:
[0011] Multiple in-memory processing computing units that parallelly process the large language model data output by external devices and output corresponding processing results. The multiple in-memory processing computing units and the in-memory processing computing array adopt a distributed design.
[0012] Optionally, the in-memory processing computing unit includes:
[0013] A data interface, connected to the external interface, is also connected to the memory module;
[0014] A processor memory, connected to the data interface, is used to process the large language model data and output the processing results;
[0015] An accumulator, connected to the processor memory, is used to perform an accumulative calculation on the processing results output by the processor memory and then output them;
[0016] A first-in-first-out queue, connected to the accumulator, is also connected to the data interface. The first-in-first-out queue is used to synchronize the processing results output by multiple in-memory processing computing units and then output them to the memory module.
[0017] Optionally, the memory module includes:
[0018] A double data rate synchronous dynamic random access memory, connected to the in-memory processing computing array, is used to store the processing results output by the in-memory processing computing array;
[0019] A high bandwidth memory, connected to the in-memory processing computing array, is used to store the computing weights of the in-memory processing computing array.
[0020] Optionally, the near-memory computing system includes a packaging structure. The in-memory processing computing array and the memory module are integrated in the packaging structure. The packaging structure includes:
[0021] A substrate;
[0022] The interposer layer is disposed on the substrate. A plurality of through-silicon vias are formed in the interposer layer. The in-memory processing computing array and the memory module are stacked on the interposer layer, and the in-memory processing computing array and the memory module are connected through the plurality of through-silicon vias.
[0023] Optionally, at least one of a non-volatile memory, a communication network module, and a sensing module is further integrated in the packaging structure.
[0024] Optionally, the in-memory processing computing array is used for processing matrix multiplication and high-speed vector processing in large language model data. The near-memory computing system further includes:
[0025] An auxiliary computing unit is connected to the external interface, and the auxiliary computing unit is used for computing other operators in large language model data.
[0026] Optionally, the near-memory computing system further includes:
[0027] A direct memory access controller is connected to the in-memory processing computing array, and the direct memory access controller is further connected to the external interface. The direct memory access controller is used for directly transmitting large language model data output by an external device to the in-memory processing computing array.
[0028] The present invention also provides a near-memory computing embodied intelligence acceleration card, including the near-memory computing system as described above.
[0029] The present invention also provides an intelligent device, including the near-memory computing embodied intelligence acceleration card as described above.
[0030] The technical solution of the present invention constitutes a near-memory computing system through an external interface, an in-memory processing computing array, and a memory module. Among them, the external interface can be used to connect to external devices, enabling the external devices to output large language model data to the near-memory computing system through the external interface; the in-memory processing computing array is connected to the external interface, and the in-memory processing computing array can receive and process the large language model data output by the external devices and output the processing results; the memory module is connected to the in-memory processing computing array and the external interface. The memory module can store the computing weights of the in-memory processing computing array, and the memory module can also store the processing results output by the in-memory processing computing array, as well as store the large language model data output by the external devices; the memory module and the in-memory processing computing array adopt a distributed design. This solution adopts distributed storage of weight data and static input data. The distributed weight storage greatly improves the data bandwidth, and the static input data based on in-memory processing reduces the data transmission volume. Therefore, the overall performance of the hardware architecture is improved, the latency is reduced, and at the same time, the distributed in-memory processing architecture can achieve various independent parallel computations of different in-memory processing, achieving higher throughput performance, thereby realizing faster data processing speed and lower power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0032] Figure 1 It is a schematic diagram of the functional modules of an embodiment of the near-memory computing system of the present invention.
[0033] Figure 2 It is a schematic diagram of the functional modules of an embodiment of the in-memory processing computing unit in the near-memory computing system of the present invention.
[0034] Figure 3 It is a schematic diagram of the overall architecture of an embodiment of the near-memory computing system of the present invention.
[0035] Figure 4 It is a schematic diagram of the packaging structure in the near-memory computing system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] To make the purpose, technical solutions and effects of the present invention clearer and more definite, the following further elaborates on the present invention with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0037] In the embodiments and the claims, unless the article is specifically defined in the context, the articles "a", "an", "the" and "said" may also include the plural forms. If the description of "first", "second", etc. is involved in the embodiments of the present invention, the descriptions of "first", "second", etc. are for descriptive purposes only, and should not be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second" may explicitly or implicitly include at least one of such features.
[0038] It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements may also be present. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.
[0039] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0040] In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0041] Currently, we are in an era of explosive growth of large language models (LLMs). These models play an important role in integrating into human life and improving efficiency, etc., but at the same time, they also face major challenges in computing performance and resources. The training and inference of large language models usually require a large amount of data, especially frequent access to input data and model parameters, and their demand for memory capacity and storage bandwidth (such as Dynamic Random Access Memory, DRAM for short) exceeds the capabilities of the current largest-scale processors and memory systems.
[0042] To address the bottleneck caused by the separation of storage and computing in traditional computing models, various in-memory computing architectures have been proposed and developed. However, in existing Process-In-Memory (PIM) architectures, as the integration density increases, more and more components are placed in a limited space, which leads to a decline in signal integrity and an increase in latency, resulting in slower processing speed and higher power consumption of in-memory computing. Meeting high-bandwidth requirements requires more complex packaging technologies and high-quality circuit board designs. Especially in a data center environment, the adoption of High-Bandwidth-Memory (HBM) will further drive up the overall cost. Although storage technologies are constantly evolving, the iteration speed of large-capacity storage devices such as Solid-State-Drives (SSDs) and Dynamic-Random-Access-Memories (DRAMs) is relatively slow, and high-bandwidth and low-latency storage solutions are not yet widely applied.
[0043] To solve the above problems, the present invention proposes an in-memory computing system.
[0044] Referring to Figure 1 , in one embodiment, the in-memory computing system includes:
[0045] An external interface 10 for connecting to external devices;
[0046] An in-memory processing computing array 20 connected to the external interface 10, where the in-memory processing computing array 20 is used to process large language model data output by external devices and output processing results;
[0047] A memory module 30 connected to the in-memory processing computing array 20, where the memory module 30 is also connected to the external interface 10. The memory module 30 stores the computing weights of the in-memory processing computing array 20, and the memory module 30 is used to store the processing results output by the in-memory processing computing array 20, as well as store large language model data output by external devices; the memory module 30 and the in-memory processing computing array 20 adopt a distributed design.
[0048] In this embodiment, the external interface 10 may be a Peripheral Component Interconnect Express (PCIe) interface. The near-memory computing system can interact with external devices for data and instruction control through the PCIe interface, and uses a shared bus for communication to control the registers in each module of the control system. The data output by the external device can also perform read-write interactions with the memory module 30 through the bus. The in-memory processing computing array 20 is an architecture that combines computing and data storage to improve the efficiency and speed of data processing. The in-memory processing computing array 20 can directly perform calculations in the memory, reducing the need for data transmission, thereby improving the overall performance. The in-memory processing computing array 20 can be composed of multiple groups of in-memory processing computing units, and multiple groups of in-memory processing computing units can perform parallel processing on the large language model data output by the external device, which can also significantly improve the computing efficiency and performance. The memory module 30 can be composed of multiple memories, such as on-chip memories and off-chip memories and other storage devices. The on-chip memory can be integrated inside the processor chip in the system and is usually used to store temporary data and instructions for quick access, such as caches and registers. The off-chip memory refers to the memory located outside the processor chip and is usually used for long-term storage of data and programs, such as dynamic random access memory, static random access memory, and flash memory. Through the memory module 30, the processing results output by the in-memory processing computing array 20 can be stored, as well as the large language model data output by the external device, and the computing weights of the in-memory processing computing array 20 can also be stored. The distributed design is a design method for system architecture that distributes each component of the system to multiple physical or virtual locations, and the specific distribution locations can be set according to the actual situation and user requirements. And adopting a distributed design for the memory module 30 and the in-memory processing computing array 20 can improve the scalability, reliability, and flexibility of the system.
[0049] It can be understood that in this solution, the large model acceleration technology based on near-memory computing adopts distributed storage of weight data and static input data. The distributed weight storage greatly improves the data bandwidth, and the static input data based on near-memory computing reduces the data transmission volume. Therefore, the overall performance of the hardware architecture is improved, the latency is reduced, and at the same time, the distributed near-memory computing architecture can achieve various independent parallel computations for different near-memory computations, achieving higher throughput performance. The integrated solution of near-memory computing not only improves the performance of data processing but also can optimize the power consumption and space utilization of the system. The goal is to achieve faster data processing speed, lower power consumption, and smaller chip area, thereby solving the problems of limited bandwidth, slow inference speed, and high data access power consumption in the end-side deployment of large models.
[0050] The technical solution of the present invention constitutes a near-memory computing system through an external interface 10, an in-memory processing computing array 20, and a memory module 30. Among them, the external interface 10 can be used to connect to external devices, enabling the external devices to output large language model data into the near-memory computing system through the external interface 10; the in-memory processing computing array 20 is connected to the external interface 10, and the in-memory processing computing array 20 can receive and process the large language model data output by the external devices and output the processing results; the memory module 30 is connected to the in-memory processing computing array 20 and the external interface 10. The memory module 30 can store the computing weights of the in-memory processing computing array 20. The memory module 30 can also store the processing results output by the in-memory processing computing array 20, as well as store the large language model data output by the external devices; the memory module 30 and the in-memory processing computing array 20 adopt a distributed design. This solution adopts distributed storage of weight data and static input data. The distributed weight storage greatly improves the data bandwidth, and the static input data based on in-memory processing reduces the data transmission. Therefore, the overall performance of the hardware architecture is improved, the latency is reduced, and at the same time, the distributed in-memory processing architecture can implement various independent parallel computations of different in-memory processing, achieving higher throughput performance, thereby achieving faster data processing speed and lower power consumption.
[0051] In one embodiment, the in-memory processing computing array 20 includes:
[0052] A plurality of in-memory processing computing units, and the plurality of in-memory processing computing units process the large language model data output by the external devices in parallel and output the corresponding processing results. The plurality of in-memory processing computing units and the in-memory processing computing array 20 adopt a distributed design.
[0053] In this embodiment, the in-memory processing computing array 20 can be composed of multiple groups of in-memory processing computing units, such as 32 groups, and the specific quantity can be set according to the actual situation and user requirements, which is not limited in this embodiment. Each in-memory processing computing unit can support operations of 2x128 groups of INT4xFP16 and 2x32 groups of FP16xFP16, and adopts parallel processing. In this way, the in-memory processing computing array 20 composed of a plurality of in-memory processing computing units can have a faster data processing speed.
[0054] Refer to Figure 2 In one embodiment, the in-memory processing computing unit includes:
[0055] A data interface, connected to the external interface 10, and the data interface is also connected to the memory module 30;
[0056] A processor memory, connected to the data interface, and the processor memory is used to process the large language model data and output the processing results;
[0057] An accumulator, connected to the processor memory, for accumulating and calculating the processing results output by the processor memory and then outputting them;
[0058] A first-in-first-out queue, connected to the accumulator, and also connected to the data interface, for synchronizing the processing results output by multiple in-memory processing calculation units and then outputting them to the memory module 30.
[0059] In this embodiment, the data interface (referred to as Interface in Figure 2 ) can communicate with other parts of the system through a custom interface protocol. The data interface can include a read data tag (PIM_RD_ID), read data (PIM_RD_DATA[RW:0]), read valid (PIM_RD_VLD), write data tag (PIM_WR_DATA_ID), write data (PIM_WR_DATA[WW:0]), address (PIM_ADDR[AW:0]), write enable (PIM_WR), etc. The data interface can process complex commands from the processor or other control units and initiate data read / write requests or read the processed activation data to the in-memory processing calculation unit (referred to as PIM PE in Figure 2 ). The first-in-first-out queue (First-in First-out, abbreviated as FIFO) is a buffer for temporarily storing output data to ensure that the data can be processed orderly and consistently before being sent to the next part of the system. The first-in-first-out queue can smooth the speed difference between the calculation and external communication and maintain the integrity of the data stream. Figure 2 The buffer (Buffer) in Figure 2 can also temporarily store the data received by the data interface; the accumulator (referred to as Accumulator in Figure 2In the middle is the PIM, which includes 128x2 groups of INT4xFP16 and 2x32 groups of FP16xFP16 parallel computing units for performing data-intensive computing tasks such as matrix multiplication and high-speed vector processing. The weight data required by the in-memory processing computing unit can be stored in the processor memory, which supports high-speed read and write operations, improving the efficiency of data processing. The weight data can be split by multiplication tasks, matrix-blocked, and packed into a globally unified data parallel format outside the chip, and then sent to the processor memory for storage through the AXI protocol interface of the host computer. The encoder can parse the instructions sent by the host AXI protocol and send them to the control register in the processor memory. At the same time, each processor memory independently completes part of the matrix calculation, and finally completes the synchronization of the calculation through the first-in-first-out queue and returns the calculation result to the host computer. In this embodiment, the design of in-memory computing not only enhances the flexibility of the data stream but also improves the execution efficiency of computing tasks. The design of the in-memory processing computing unit not only focuses on the integration of storage and computing but also considers scalability and customization. Custom signals and interfaces allow designers to adjust and optimize the in-memory processing computing unit according to specific application requirements. Therefore, the in-memory processing computing unit can not only perform standard storage operations but also process complex computing tasks, meeting the strict requirements for speed and energy efficiency in compute-intensive applications.
[0060] In one embodiment, the memory module 30 includes:
[0061] A double data rate synchronous dynamic random access memory, connected to the in-memory processing computing array 20, and the double data rate synchronous dynamic random access memory is used to store the processing results output by the in-memory processing computing array 20;
[0062] A high-bandwidth memory, connected to the in-memory processing computing array 20, and the high-bandwidth memory is used to store the computing weights of the in-memory processing computing array 20.
[0063] In this embodiment, the memory module 30 can be composed of a double data rate synchronous dynamic random access memory and a high-bandwidth memory. The double data rate synchronous dynamic random access memory realizes a higher data bandwidth by transmitting data simultaneously on the rising and falling edges of the clock; the high-bandwidth memory adopts a stacked chip method and connects multiple memory chips through a high-bandwidth silicon interposer to achieve an extremely high data transfer rate. Therefore, in this embodiment, the double data rate synchronous dynamic random access memory can be used to store the processing results output by the in-memory processing computing array 20, and the high-bandwidth memory can be used to store the computing weights of the in-memory processing computing array 20. In addition, the memory module 30 can also include on-chip memory integrated inside the processor chip in the system for storing temporary data and instructions for fast access, such as caches and registers.
[0064] Reference Figure 4 In one embodiment, the near-memory computing system includes a package structure, and the in-memory processing computing array 20 and the memory module 30 are integrated into the package structure, and the package structure includes:
[0065] A substrate;
[0066] An interposer disposed on the substrate, and a plurality of through-silicon vias are formed on the interposer. The in-memory processing computing array 20 and the memory module 30 are stacked on the interposer, and the in-memory processing computing array 20 and the memory module 30 are connected through the plurality of through-silicon vias.
[0067] In this embodiment, the substrate ( Figure 4 in this case, a PCB) can be a circuit board. The substrate can provide functions such as electrical connection, protection, support, heat dissipation, and assembly for the chip, so as to achieve the purpose of multi-pin, reducing the volume of the packaged product, improving electrical performance and heat dissipation, ultra-high density or multi-chip modularization. The 3D die stacking integration scheme is as Figure 4 shown. By using 3D-DRAM hybrid bonding technology, dynamic random access memory (DRAM) chips are stacked in the same package. It integrates multiple chips (or dies) in one package through an interposer. This interposer is usually made of silicon wafer, and dense through-silicon vias ( Figure 4 in this case, TSV) and metal wirings are provided on the interposer for connecting dies and DRAM modules, allowing high-density micro-interconnection, thereby achieving extremely high storage density and bandwidth to meet the requirements of high-performance computing and large-scale data processing. 3D-DRAM uses a tight stacking method, which can make the connection distance between storage units shorter, the data transmission speed faster, and the energy consumption lower. This solution further integrates the in-memory processing computing unit into the interposer ( Figure 4 in this case, PIM Die), which can interact with DRAM stably and efficiently, further improving the data transmission bandwidth and reducing the latency, and can achieve more efficient deployment and execution of large model inference tasks. The 3D die stacking integration scheme of this embodiment can integrate the in-memory processing computing unit through the interposer and design a standardized interface to dock with the DRAM layer. In this way, PIM-DRAM interoperability between different manufacturers can be achieved, with better compatibility. In subsequent designs, by adapting to standard communication protocols and hardware interfaces, the compatibility between the in-memory processing computing unit and traditional computing systems as well as deployment-side hardware can also be improved. Figure 4Micro Bump and Bump in it are two different connection technologies used to achieve electrical connections between chips, especially for connections between stacked chips (such as DRAM stacks) and processors; Bump is a relatively large solder ball, usually between dozens of microns and hundreds of microns, used for chip packaging and connection; Micro Bump is a smaller solder ball, usually between a few microns and dozens of microns, used for higher-density connections.
[0068] Furthermore, in one embodiment, at least one of a non-volatile memory, a communication network module, and a sensing module is also integrated in the packaging structure. In this embodiment, based on different group vector pulsation sizes, functional modules such as a non-volatile memory, a communication network module, and a sensing module can be integrated in the packaging structure to support different-scale model deployment requirements in different scenarios.
[0069] In one embodiment, the in-memory processing computing array 20 is used to process matrix multiplication and high-speed vector processing in large language model data, and the near-memory computing system further includes:
[0070] An auxiliary computing unit, connected to the external interface 10, and the auxiliary computing unit is used to calculate other operators in the large language model data.
[0071] In this embodiment, in addition to matrix multiplication and high-speed vectors in large language model data that need to be processed, there are many other operators, such as activation functions (such as ReLU, Sigmoid, etc.), normalization operations (such as Batch Normalization, Layer Normalization, etc.), convolution operations, and pooling operations, etc. By calculating other operators through the auxiliary computing unit, a more flexible and efficient computing architecture can be achieved.
[0072] In one embodiment, the near-memory computing system further includes:
[0073] A direct memory access controller, connected to the in-memory processing computing array 20, and the direct memory access controller is also connected to the external interface 10, and the direct memory access controller is used to directly transfer the large language model data output by the external device to the in-memory processing computing array 20.
[0074] In this embodiment, the direct memory access controller adopts the Direct Memory Access (DMA) technology, which allows peripheral devices (such as hard disks, network adapters, graphics cards, etc.) to directly interact with the memory without the intervention of the CPU. In this way, using the direct memory access controller to directly transfer the large language model data output by the external device to the in-memory processing computing array 20 for processing and calculation can significantly improve the data transfer efficiency and system performance. The overall architecture of the near-memory computing system can refer toFigure 3 。
[0075] The present invention also provides a near-memory computing embodied intelligence acceleration card.
[0076] In one embodiment, the near-memory computing embodied intelligence acceleration card includes the near-memory computing system described above. It can be understood that since the above near-memory computing system is used in the near-memory computing embodied intelligence acceleration card of the present invention, the embodiments of the near-memory computing embodied intelligence acceleration card of the present invention include all the technical solutions of all the embodiments of the above near-memory computing system, and the achieved technical effects are also exactly the same, so they will not be elaborated here. In the near-memory computing embodied intelligence acceleration card of this embodiment, the near-memory computing processor, memory and other functional modules are integrated in a 3D stacking manner, significantly improving the performance and efficiency of near-memory computing. This integration method shortens the data transmission path, reduces latency, and thus speeds up data processing. Compared with traditional planar integration, 3D die integration technology can achieve higher bandwidth and lower power consumption, which is crucial for processing big data and performing complex computing tasks. In addition, 3D die integration also improves the integration density of the chip and reduces the volume of the device, enabling more functions to be integrated in a limited space. On the other hand, the near-memory computing embodied intelligence acceleration card based on 3D die integration can achieve dialogue and instruction sending through display terminals such as personal computers and mobile phones, so as to realize that the embodied intelligence card concurrently drives multiple intelligent hardware to complete corresponding functions. And the near-memory computing embodied intelligence acceleration card proposed in this embodiment can not only be applied to large language models, but also be extended to other types of neural network architectures.
[0077] The present invention also provides an intelligent device.
[0078] In one embodiment, the intelligent device includes the near-memory computing embodied intelligence acceleration card described above. It can be understood that since the above-mentioned near-memory computing embodied intelligence acceleration card is used in the intelligent device of the present invention, therefore, the embodiments of the intelligent device of the present invention include all the technical solutions of all the embodiments of the above-mentioned near-memory computing embodied intelligence acceleration card, and the achieved technical effects are also exactly the same, which will not be elaborated here. The intelligent device in this embodiment has significant advantages in multi-task parallel processing and high-throughput data transmission. For example, it shows extremely high efficiency in a complex intelligent device ecosystem. Through the near-memory computing embodied intelligence acceleration card in this solution, users can simultaneously interact with multiple intelligent hardware devices in real time without waiting for a single task to complete, truly achieving efficient parallel drive. For example, users can send instructions to multiple devices almost seamlessly through voice or other input methods. Users can, while controlling a humanoid robot to perform complex actions through voice commands, simultaneously let a robotic arm execute precise operations such as assembly or welding. During this process, a floor cleaning robot can also receive instructions and independently perform room cleaning tasks. In addition, the motion control of a robotic dog can also respond to commands immediately and coordinate with other devices. The embodied intelligence acceleration card in this solution supports this ability of multiple hardware to respond simultaneously, thanks to its powerful data throughput and parallel computing capabilities. It can simultaneously process data inputs from multiple sensors, quickly perform data analysis and decision-making, so as to ensure that the actions of all devices can be coordinated and interfere with each other. This high-throughput parallel processing greatly improves the response speed and flexibility of the entire system, enabling multiple devices to cooperate together to complete complex task scenarios, expanding the application scenarios and efficiency of intelligent hardware. Therefore, the characteristic of the embodied intelligence acceleration card in this solution lies in its high-throughput parallel processing ability, which not only speeds up the response speed of the intelligent hardware system, but also ensures the efficient cooperation and real-time control of multiple devices in a complex environment.
[0079] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A near storage computing system, characterized in that: include: External interface, used to connect external devices; An in-memory processing and computing array connected to the external interface, the in-memory processing and computing array is used to process the large language model data output by the external device and output the processing results; A memory module is connected to the in-memory processing and computing array. The memory module is also connected to the external interface. The memory module stores the computing weights of the in-memory processing and computing array. The memory module is used to store the processing results output by the in-memory processing and computing array, and to store large language model data output by an external device. The memory module and the in-memory processing and computing array adopt a distributed design.
2. The near storage computing system according to claim 1, characterized in that: The in-memory processing computing array includes: Multiple in-memory processing and computing units, the multiple in-memory processing and computing units process large language model data output by external devices in parallel and output corresponding processing results, and the multiple in-memory processing and computing units and the in-memory processing and computing array adopt a distributed design.
3. The near storage computing system according to claim 2, characterized in that: The in-memory processing and computing unit comprises: A data interface connected to the external interface, the data interface also connected to the memory module; A processor memory connected to the data interface, the processor memory being used to process large language model data and output processing results; an accumulator connected to the processor memory, the accumulator being used to perform cumulative calculation on the processing results output by the processor memory and then output the result; A first-in-first-out queue is connected to the accumulator and is also connected to the data interface. The first-in-first-out queue is used to synchronize the processing results output by multiple in-memory processing and calculation units and then output them to the memory module.
4. The near storage computing system according to claim 1, wherein: The memory module comprises: A double data rate synchronous dynamic random access memory connected to the in-memory processing and computing array, the double data rate synchronous dynamic random access memory is used to store the processing results output by the in-memory processing and computing array; A high bandwidth memory is connected to the in-memory processing and computing array, and the high bandwidth memory is used to store the computing weights of the in-memory processing and computing array.
5. The near storage computing system according to claim 1, wherein: The near-memory computing system comprises a packaging structure, in which the in-memory processing computing array and the memory module are integrated, and the packaging structure comprises: substrate; An interposer is arranged on the substrate, a plurality of through silicon vias are opened on the interposer, the in-memory processing and computing array and the memory module are stacked and arranged on the interposer, and the in-memory processing and computing array and the memory module are connected through the plurality of through silicon vias.
6. The near storage computing system according to claim 5, characterized in that: The packaging structure also integrates at least one of a non-volatile memory, a communication network module and a sensor module.
7. The near storage computing system according to claim 1, wherein: The in-memory processing computing array is used to process matrix multiplication and high-speed vector processing in large language model data, and the near-memory computing system also includes: An auxiliary computing unit is connected to the external interface, and is used to calculate other operators in the large language model data.
8. The near storage computing system according to claim 1, wherein: Also includes: A direct memory access controller is connected to the in-memory processing and computing array. The direct memory access controller is also connected to an external interface. The direct memory access controller is used to directly transmit large language model data output by an external device to the in-memory processing and computing array.
9. A near-memory computing embodied intelligent acceleration card, characterized in that: It comprises a near storage computing system as described in any one of claims 1 to 8.
10. A smart device, characterized in that: It includes the near-memory computing embodied intelligent acceleration card as described in claim 9.