Near memory computing coprocessor system
By integrating a near-memory computing coprocessor around the storage module, the problems of slow data transfer speed and high energy consumption in the von Neumann architecture are solved, enabling efficient processing of memory-intensive computing tasks and improving system performance and energy efficiency.
Patent Information
- Application Number
- CN202511099656.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
In the von Neumann architecture, a large amount of data is repeatedly moved between the computing core and memory during the computation process. The bus data bandwidth limits the data transfer speed, causing the arithmetic unit to wait frequently, reducing the utilization of the arithmetic unit, and resulting in a decrease in system performance.
Design a near-memory computing coprocessor system, including a main processor, a near-memory computing coprocessor and a storage module connected by communication. By integrating the near-memory computing coprocessor around the storage module, a direct connection between the high-bandwidth bus and the storage module is realized, and memory-intensive computing tasks are offloaded to the near-memory computing coprocessor to complete, thereby reducing memory access latency and data transfer overhead.
It improves the overall computing efficiency of the system, reduces memory access latency and data transfer overhead, and enhances the performance and energy efficiency of the computing system.
Smart Images

Figure CN120994379A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of coprocessor, and in particular to a near-memory computing coprocessor system. BACKGROUND
[0002] The Von Neumann architecture is the basic design paradigm of modern computers, which is usually composed of a main processor and a storage module. The storage unit (memory) and the computing unit (CPU) are physically separated, and the instruction and data transmission process during the computation is completed through the bus.
[0003] The mainstream computer storage unit is usually composed of DDR, LPDDR, GDDR, HBM storage modules based on DRAM technology, and other volatile or non-volatile storage modules. During the computation, the instructions and data required by the processor are transported from the storage module to the main processor through a long-distance (10-20 cm) signal transmission path via the bus. After the computation is completed in the main processor, the computation results are returned to the storage module through the same signal transmission path. In the past use scenarios, the computation task has a high computation access ratio, and the bus bandwidth is sufficient to support the computation demand. The computation delay of the arithmetic unit in the main processor becomes the main factor limiting the performance of the main processor. The Von Neumann architecture realizes universality and flexibility through modular design, and becomes the cornerstone of computer science.
[0004] However, the performance of the current main processor and the development of the memory bandwidth have a serious mismatch, resulting in that the access speed of the current memory seriously lags behind the computation speed of the main processor. The memory bottleneck causes the high-performance processor to be difficult to play its due role. A large amount of data needs to be repeatedly transported between the computation core and the memory during the computation, and the bus data bandwidth limits the data transport speed, resulting in frequent waiting of the arithmetic unit for obtaining computation data, reducing the utilization rate of the arithmetic unit, and causing the system performance to decline. In addition, the frequent data transport process will bring a large amount of power consumption overhead. The energy consumption of the arithmetic unit for reading 1 bit of data from the external memory is more than 200 times the energy consumption of the arithmetic unit for completing one computation. SUMMARY
[0005] The present application provides a near-memory computing coprocessor system to solve the technical problem that in the existing Von Neumann architecture, a large amount of data needs to be repeatedly transported between the computation core and the memory during the computation, the bus data bandwidth limits the data transport speed, resulting in frequent waiting of the arithmetic unit for obtaining computation data, reducing the utilization rate of the arithmetic unit, and causing the system performance to decline.
[0006] The present application provides a near-memory computing coprocessor system, comprising:
[0007] A main processor, a near-memory computing coprocessor, and a storage module are connected in communication. The storage module is provided with a plurality of storage particles. The storage particles are volatile storage particles or non-volatile storage particles.
[0008] The near-memory computing coprocessor includes
[0009] A communication protocol decoding unit is configured to receive a memory control instruction sent by the main processor and parse the memory control instruction to obtain executable decoding. The memory control instruction includes a near-memory computing start instruction and a configuration acquisition instruction.
[0010] A processing unit is configured to control the behavior of the near-memory computing coprocessor according to the executable decoding, so that the near-memory computing coprocessor runs in a set mode. The set mode includes a computing mode, a pass-through mode, and a compatible mode.
[0011] A near-memory computing unit is configured to perform near-memory computing.
[0012] An on-chip buffer is configured to cache data to be near-memory computed.
[0013] A plurality of first memory controllers are connected to the storage particles through a high-speed selector. The number of the first memory controllers is equal to the number of the storage particles. The first memory controllers are configured to independently access the data stored in the storage particles. The high-speed selector is configured to determine the access control right of the storage particles for the main processor or the near-memory computing coprocessor.
[0014] In some embodiments, the main processor is provided with a second memory controller. The second memory controller is configured to send a memory control instruction to the near-memory computing coprocessor.
[0015] The first memory controller is provided with a first memory physical layer interface.
[0016] The storage particles are provided with a storage particle interface. The first memory controller is connected to the storage particles through the storage particle interface and the first memory physical layer interface. The specification of the storage particle interface is the same as that of the first memory physical layer interface.
[0017] In some embodiments, when the amount of data processed by the main processor is greater than a preset data amount, the main processor generates the near-memory computing start instruction and near-memory computing configuration information and sends them to the communication protocol decoding unit and the on-chip buffer, respectively.
[0018] The communication protocol decoding unit is configured to decode the near-memory-computing start instruction into a start decoding sent to the processing unit, and start the processing unit;
[0019] The processing unit is configured to:
[0020] According to the start decoding, release the control right of the memory granule to the main processor;
[0021] According to the near-memory-computing configuration information, activate the first memory physical layer interface corresponding to the memory granule interface, and access the data stored in the memory granule;
[0022] After the access is completed, return the control right to the main processor.
[0023] In some embodiments, the step of returning the control right to the main processor after the access is completed is further configured to:
[0024] send a control right return instruction to the main processor;
[0025] The main processor is further configured to:
[0026] receive the control right return instruction, and send a control instruction to the memory granule; if a reply signal of the memory granule can be received, the control right has been returned to the main processor.
[0027] In some embodiments, the main processor is configured with a memory control word lookup table; the memory control word lookup table includes: address intervals of each memory granule, memory control words, and calibration parameters;
[0028] The second memory controller stores the memory control words and calibration parameters;
[0029] The main processor is configured to:
[0030] obtain an access address; the access address includes: an address interval of the memory granule to be accessed;
[0031] based on the access address, determine the memory control words and calibration parameters corresponding to the memory granule by using the memory control word lookup table;
[0032] The second memory controller is further configured to:
[0033] generate the memory control instruction and the near-memory-computing configuration information based on the memory control words and calibration parameters.
[0034] In some embodiments, when the number of storage particles to be accessed is equal to the total number of storage particles in the storage module, the near-memory computing coprocessor operates in a compute mode; when the number of storage particles to be accessed is less than the total number of storage particles in the storage module, the near-memory computing coprocessor operates in a compatible mode.
[0035] When the near-memory computing coprocessor operates in the compute mode, the host processor does not have control over the storage particles; when the near-memory computing coprocessor operates in the compatible mode, the host processor has partial control over the storage particles.
[0036] In some embodiments, before the host processor receives the control return instruction or the near-memory computing coprocessor receives the memory control instruction, the computing data cached in the host processor or the near-memory computing coprocessor is updated to the storage particles; after the computing data is updated, the control of the storage particles is returned to the host processor or the near-memory computing coprocessor is allowed to have control over the storage particles.
[0037] In some embodiments, when the amount of data processed by the host processor is less than or equal to a preset data amount, the near-memory computing coprocessor operates in the pass-through mode; the host processor is configured to send the configuration acquisition instruction to the near-memory computing coprocessor.
[0038] The near-memory computing coprocessor is further configured to:
[0039] receive the configuration acquisition instruction, so that the near-memory computing unit, the on-chip buffer, the first memory controller, and the first memory physical layer interface enter a sleep state; and control, by the processing unit, the high-speed selector to select the configuration acquisition instruction sent by the second memory controller for reception;
[0040] According to the configuration acquisition instruction, access the data stored in the storage particles.
[0041] In some embodiments, the host processor is further configured to:
[0042] send an address access request to the communication protocol decoding unit, which is parsed to determine the storage particles to be accessed;
[0043] If the host processor does not have control over the storage particles, interrupt the access between the first memory controller and the storage particles; and establish the access between the host processor and the storage particles.
[0044] In some embodiments, the near-memory computing coprocessor is further provided with a second memory physical layer interface; the second memory controller is provided with a third memory physical layer interface;
[0045] The main processor is connected to the near-memory computing coprocessor through the third memory physical layer interface connected to the second memory physical layer interface;
[0046] The near-memory computing coprocessors are cascaded through the second memory physical layer interface connected to the first memory physical layer interface.
[0047] The application provides a near-memory computing coprocessor system, comprising a main processor, a near-memory computing coprocessor, and a storage module connected in communication; the storage module is provided with a plurality of storage particles; the storage particles are volatile storage particles or non-volatile storage particles; the near-memory computing coprocessor comprises a communication protocol decoding unit configured to receive a memory control instruction sent by the main processor and analyze the memory control instruction to obtain an executable code; the memory control instruction comprises a near-memory computing starting instruction and a configuration acquisition instruction; a processing unit configured to control the behavior of the near-memory computing coprocessor according to the executable code to make the near-memory computing coprocessor run in a set mode; the set mode comprises a computing mode, a transparent mode, and a compatible mode; a near-memory computing unit configured to perform near-memory computing; an on-chip buffer configured to cache data to be near-memory computed; a plurality of first memory controllers connected to the storage particles through a high-speed selector; the number of the first memory controllers is equal to the number of the storage particles; the first memory controllers are configured to independently access the data stored in the storage particles; and the high-speed selector is configured to determine the memory access control right of the main processor or the near-memory computing coprocessor to the storage particles to unload the processing of the memory-intensive task to the near-memory computing coprocessor system under the control of the main processor, thereby greatly improving the overall computing efficiency of the system. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0049] Figure 1 It is a structural schematic diagram of the near-memory computing coprocessor system (four channels) in the present application.
[0050] Figure 2 It is a structural schematic diagram of the von Neumann architecture.
[0051] Figure 3 The structure diagram of the near-memory computing coprocessor applied to the Von Neumann architecture in the present application;
[0052] Figure 4 The structure diagram of the connection of various components of the near-memory computing coprocessor system in the present application;
[0053] Figure 5 The structure diagram of the near-memory computing coprocessor system in the present application in the transparent mode;
[0054] Figure 6 The diagram of the memory control word lookup table in the present application;
[0055] Figure 7 The structure diagram of the near-memory computing coprocessor system in the present application in the compatible mode;
[0056] Figure 8 The structure diagram of the near-memory computing coprocessor system in the present application in the cascade mode.
[0057] Explanation of reference signs:
[0058] 1-main processor; 11-second memory controller; 111-third memory physical layer interface; 2-near-memory computing coprocessor; 21-communication protocol decoding unit; 22-processing unit; 23-near-memory computing unit; 24-on-chip buffer; 25-first memory controller; 251-first memory physical layer interface; 26-high-speed selector; 27-second memory physical layer interface; 3-storage module; 31-storage grain; 311-storage grain interface. DETAILED DESCRIPTION
[0059] In order to enable persons skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should fall within the scope of protection of the present application.
[0060] For example, the Von Neumann architecture is the basic design paradigm of modern computers, which is usually composed of a main processor and a storage module. The storage unit (memory) and the computing unit (CPU) are physically separated, and the instruction and data transmission process in the computing process is completed through the bus (Bus), such as Figure 2The main stream computer storage unit is usually built by DDR, LPDDR, GDDR, HBM storage module based on DRAM technology, and other volatile or non-volatile storage modules. During the calculation, the instructions and data required by the processor are carried to the main processor from the storage module through the long-distance (10-20 cm) signal transmission path through the bus. After the calculation is completed in the main processor, the calculation result is returned to the storage module through the same signal transmission path. In the current use scenario, the calculation task has a high calculation storage ratio, and the bus bandwidth is sufficient to support the calculation demand, and the calculation delay of the arithmetic unit becomes the main factor limiting the performance of the processor. The von Neumann architecture realizes the generality and flexibility through modular design, and becomes the cornerstone of computer science.
[0061] However, in the development process of computers, the performance of the processor and the development of the memory bandwidth have a serious mismatch, resulting in a serious lag in the access speed of the current memory behind the calculation speed of the processor. The memory bottleneck makes it difficult for high-performance processors to perform their due functions. With the continuous rise of artificial intelligence, big data and other technologies, the main memory access operation in the von Neumann computing architecture has become a key performance and energy bottleneck of the hardware platform, resulting in the "memory wall" problem of high-performance computing. A large amount of data needs to be repeatedly carried between the computing core and the memory during the calculation, and the bus data bandwidth limits the data carrying speed, resulting in frequent waiting of the arithmetic unit for calculation data, reducing the utilization rate of the arithmetic unit, and causing the system performance to decline. In addition, the frequent data carrying process will bring a lot of power consumption overhead. In the advanced process node below 7nm, the energy consumption of the arithmetic unit to read 1bit data from the external memory is more than 200 times the energy consumption of the arithmetic unit to complete a calculation once.
[0062] With the continuous improvement of integrated circuit process, the performance of the processor is improved exponentially, but the bandwidth and access power consumption of the memory module develop slowly, and the imbalance between the development speed of the two causes the "memory wall" to become a major obstacle to the performance improvement of the von Neumann architecture data computing system. With the development of cloud computing and artificial intelligence technology in recent years, the calculation task has greater memory demand, and the delay and power consumption caused by slow data carrying and large carrying energy have become the key bottleneck of the computing system.
[0063] A direct method to solve the memory access bottleneck of existing computing systems is to improve the memory access bandwidth of the storage module. High Bandwidth Memory (HBM) realizes the stacking of multiple layers of DRAM storage chips through Through-Silicon Vias (TSV) technology, and realizes high data bandwidth by integrating a wider bus and more parallel channels. HBM achieves higher data bandwidth with smaller volume and less power, thereby reducing the additional load delay and platform power consumption caused by the data access process. However, this will cause the cost of storage particles to rise sharply, and will affect the design of the storage bus interface of the main processor of the computing system, so that it can only be applied to a few customized solutions of high-performance computing systems, and cannot be adapted to widely used low-performance computing systems.
[0064] Another method to solve the memory access bottleneck of existing computing systems is in-memory computing. Storage manufacturers transfer computing tasks to the inside of the storage unit by integrating arithmetic units at various levels of the storage particle, and use the multi-level internal bus inside the storage unit to achieve a memory access bandwidth far exceeding the external interface, while avoiding multi-level memory transfer. However, this method destroys the structure of the storage array, and requires chip design under advanced processes to meet timing requirements, which is difficult to design and has high production costs. In addition, the arithmetic unit integrated inside the storage particle needs to complete the computing task based on a special instruction set, which requires modification of existing content control instructions and communication protocols, and cannot be compatible with existing computing systems.
[0065] Due to the fact that in some technologies, the Von Neumann architecture requires a large amount of data to be repeatedly transferred between the computing core and the memory during the computing process, the bus data bandwidth limits the data transfer speed, causing the arithmetic unit to frequently wait for computing data, reducing the utilization rate of the arithmetic unit, and causing system performance to decline, in order to solve the technical problem, the present application provides a near-memory computing coprocessor system, the structure of each part of the near-memory computing coprocessor system is described as follows:
[0066] The present application provides a near-memory computing coprocessor system, comprising:
[0067] The main processor 1, the near-memory computing coprocessor 2, and the storage module 3 are communicatively connected; the storage module 3 is provided with a plurality of storage particles 31; the storage particles 31 are volatile storage particles or non-volatile storage particles.
[0068] The near-memory computing coprocessor 2 comprises:
[0069] The communication protocol decoding unit 21 is configured to receive the memory control instruction sent by the main processor 1 and parse the memory control instruction to obtain an executable decoding; the memory control instruction includes: a near-memory computing start instruction and a configuration acquisition instruction; the processing unit 22 is configured to control the behavior of the near-memory computing coprocessor 2 according to the executable decoding, so that the near-memory computing coprocessor 2 runs in a set mode; the set mode includes: a computing mode, a transparent transmission mode and a compatible mode.
[0070] In this embodiment, the transparent transmission mode means that the near-memory computing coprocessor 2 does not participate in computing, and only completes the transparent transmission of the storage bus signal of the main processor 1. The near-memory computing coprocessor 2 proposed in the application suspends the computing function, and all control signals, address signals and data signals on the storage bus connected with the second memory controller 11 of the main processor 1 are transmitted through the near-memory computing coprocessor 2 to complete the interaction with the storage module 3. At this time, the working mode and the memory access mode of the main processor 1 are the same as those of the current von Neumann architecture computing system. The computing mode and the compatible mode mean that the near-memory computing coprocessor 2 assists the main processor 1 to complete the corresponding computing task. The near-memory computing coprocessor 2 proposed in the application has the functions of control, computing and memory access. The main processor 1 offloads the computing task to the near-memory computing coprocessor 2, the near-memory computing coprocessor 2 takes over the control right of the storage module 3, realizes efficient memory access by using a high-bandwidth bus, and completes various memory-intensive tasks in the field of efficient processing AI computing.
[0071] The near-memory computing unit 23 is configured to perform near-memory computing; the on-chip buffer 24 is configured to cache data to be near-memory computed; a plurality of first memory controllers 25 are connected with the storage particles 31 through a high-speed selector 26; the number of the first memory controllers 25 is equal to the number of the storage particles 31; the first memory controllers 25 are configured to independently access the data stored in the storage particles 31; and the high-speed selector 26 is configured to determine the memory access control right of the storage particles 31 for the main processor 1 or the near-memory computing coprocessor 2.
[0072] The application provides a near-memory computing coprocessor system, which does not change the memory access mode of the arithmetic unit-memory unit of the existing von Neumann architecture. The near-memory computing coprocessor 2 and one or more storage particles 31 are additionally added around the storage module 3 to realize 2D / 3D integration (distance < 5 cm), so that the memory access bandwidth of the near-memory computing coprocessor 2 can break through the bandwidth limitation of the traditional storage module 3, reduce the data carrying distance while reducing the memory access delay, realize the low-delay and low-power memory access of the near-memory computing coprocessor 2 to the storage module 3. For memory-intensive computing tasks, the related processing is directly completed in the near-memory computing coprocessor 2 without being moved to the outside of the storage module 3. The addition of the near-memory computing coprocessor 2 reduces the memory access delay and data carrying cost of the computing system during the execution of the memory-intensive computing task, and the computing efficiency of the traditional von Neumann architecture computing system is greatly improved.
[0073] The near-memory computing coprocessor 2 architecture proposed in the application is connected with the main processor 1 through a storage bus at one end, and is connected with the storage module 3 through a high-bandwidth bus at the other end, and can also be realized as in-package or on-chip integration with the storage module 3, thereby forming a new storage module 3 with near-memory computing function, replacing the current storage module 3 connected with the main processor 1 and improving the overall performance of the computing system.
[0074] The application does not change the computing architecture of the existing main processor 1 and storage module 3, and integrates the near-memory computing coprocessor 2 near the storage module 3, as shown in the drawing, to assist the main processor 1 to complete the memory-intensive computing task. Figure 3 The near-memory computing coprocessor 2 and the storage module 3 have higher storage bandwidth, thereby reducing the memory access cost and being able to complete the computing task with higher efficiency. During the running process, the memory-intensive task is unloaded to the near-memory computing coprocessor 2 under the control of the main processor 1 to complete the processing, thereby greatly improving the overall computing efficiency of the system.
[0075] For example, the computing core currently using the common von Neumann architecture computing architecture and the storage module 3 are usually separate chips, and the interconnection is completed in the form of PCB wiring or chip packaging. Limited by the area, power consumption and memory channel in the design process of the main processor 1 and the chip, the distance between the storage module 3 and the main processor 1 is large, and limited by the signal integrity problem, the number and interface speed of the interconnection interface cannot be effectively improved, thereby reducing the interconnection bandwidth. The near-memory computing coprocessor 2 reserves more memory controller interfaces in the design process, one end receives the memory controller signal controlled by the main processor 1, and the other end is connected with the storage grain 31 through the led-out multiple first memory controllers 25, thereby realizing the storage expansion and bandwidth improvement. The near-memory computing coprocessor 2 has a memory interface, and the interface specification of the near-memory computing coprocessor 2 is the same as that of the storage grain 31, which is used to receive the content control signal sent by the second memory controller 11 of the main processor 1, and has a single-channel data bandwidth. The near-memory computing coprocessor 2 contains multiple first memory controllers 25, which realize the interconnection with multiple storage grains 31 by using multiple memory channels, and has a multi-channel data bandwidth, thereby realizing high-bandwidth multi-channel parallel memory access. Figure 4 The near-memory computing coprocessor 2 with one memory interface and four first memory controllers 25 is shown in the figure, and compared with the main processor 1, it has four times the memory bandwidth for the four storage grains 31 integrated in the system. The type of the storage grain 31 can be a volatile storage grain such as DRAM, SRAM, etc., or a non-volatile storage grain such as Flash, ReRAM, MRAM, etc.
[0076] In this embodiment, the main processor 1 is provided with a second memory controller 11; the second memory controller 11 is configured to send a memory control instruction to the near-memory computing coprocessor 2; and the first memory controller 25 is provided with a first memory physical layer interface 251.
[0077] The storage grain 31 is provided with a storage grain interface 311; the first memory controller 25 is connected to the storage grain 31 through the storage grain interface 311 and the first memory physical layer interface 251; the specification of the storage grain interface 311 is the same as that of the first memory physical layer interface 251; the external interface of the near-memory computing coprocessor 2 proposed in the present application is completely the same as the external interface and protocol of the current storage module 3, without the need to add additional control lines or modify the existing interaction instructions in the interface, and can be directly inserted between the main processor 1 and the storage module 3 without changing other parts of the computing system. The interaction between the main processor 1 and the near-memory computing coprocessor 2 is completed based on a low-bandwidth memory bus, and is realized by reading and writing a control word at a specific address in the storage space.
[0078] In this embodiment, in addition to the 2D integration shown in 4, the near-memory computing coprocessor 2 can also implement 3D stacked integration with the storage grain 31, connecting the near-memory computing coprocessor 2 with one or more layers of storage grains 31 through hybrid bonding, which can greatly improve the number and density of interconnections, and using 3D stacking technology can achieve greater bandwidth improvement between the near-memory computing coprocessor 2 and the storage grain 31. Compared with the main processor 1, the interconnection bandwidth between the near-memory computing coprocessor 2 and the single storage grain 31 is higher, which can further improve the memory bandwidth. Among them, the connection mode of the storage grain interface 311 and the first memory physical layer interface 25 depends on the chip process used. If the near-memory computing coprocessor 2 and the storage grain 31 are in the form of separate packaging, the interconnection between the corresponding pins of the discrete packages is completed using short-distance PCB wiring. If they are integrated in the same package, the connection is achieved using the packaging form.
[0079] In this embodiment, when the amount of data processed by the main processor 1 is greater than the preset data amount, the main processor 1 generates the near-memory computing start instruction and the near-memory computing configuration information and sends them to the communication protocol decoding unit 21 and the on-chip buffer 24, respectively; the communication protocol decoding unit 21 is configured to decode the near-memory computing start instruction into a start code and send it to the processing unit 22 to start the processing unit 22.
[0080] The processing unit 22 is configured to:
[0081] According to the start code, the control right of the main processor 1 over the storage grain 31 is released; according to the near-memory computing configuration information, the corresponding first memory physical layer interface 251 and the storage grain interface 311 are activated, and the data stored in the corresponding storage grain 31 is accessed; after the access is completed, the control right is returned to the main processor 1.
[0082] In this embodiment, as Figure 1The diagram shows the internal architecture of a near-memory computing coprocessor 2 with a single memory interface and four first memory controllers 25. The processing unit 22 (CPU) receives memory control instructions from the main processor 1 via a standard physical layer (PHY) interface, completing the interconnection with the main processor 1. The near-memory computing coprocessor 2 has an internal communication protocol decoding unit 21 that parses memory instructions. The parsed host commands are directly passed to the internal processing unit 22 of the near-memory computing coprocessor 2, enabling behavioral control of the near-memory computing coprocessor 2. The near-memory computing unit 23 within the near-memory computing coprocessor 2 performs high-speed parallel near-memory computing functions to match the high memory access bandwidth of the near-memory computing coprocessor 2. The near-memory computing coprocessor 2 uses four first memory controllers 25 to achieve independent access to four memory chips 31. Four high-speed selectors 26 are used to determine whether the main processor 1 or the near-memory computing coprocessor 2 has access control over the memory chips 31. After configuration by the main processor 1, the near-memory computing coprocessor 2 can operate in computing mode, pass-through mode, and compatibility mode.
[0083] Specifically, such as Figure 1 The diagram shows the structure of the near-memory computing coprocessor 2 in computing mode. The near-memory computing coprocessor system defines a dedicated address range for the main processor 1 to configure the near-memory computing coprocessor 2. When the host needs to enable the computing function of the near-memory computing coprocessor 2, it needs to write the corresponding data information into this address range. The overall process consists of the following 5 steps:
[0084] 1. The main processor 1 writes a near-memory computing start instruction into the specified reserved address range. The communication protocol decoding unit 21 of the near-memory computing coprocessor 2 completes the instruction decoding and sends the start instruction to the processing unit 22.
[0085] 2. The main processor 1 generates near-memory computing configuration information, writes it into the near-memory computing reserved address range, and stores it in the on-chip buffer 24 of the near-memory computing coprocessor 2 after decoding. Then, the main processor 1 performs a cache refresh to synchronize the cache contents with the memory. After that, it writes a memory release control word to the specified address or performs a write operation to the reserved register as a switching command to release the bus. The JEDCE communication protocol decoding unit 21 of the near-memory computing coprocessor 2 generates a memory release signal after decoding.
[0086] 3. The near-memory computing coprocessor 2 activates the near-memory computing unit 23 and the first memory controllers 25 and the memory interface based on the memory release signal, and starts the near-memory computing based on the near-memory computing configuration information written by the host processor 1. In this process, the multiple first memory controllers 25 of the near-memory computing coprocessor 2 work in parallel, and high-bandwidth access to the storage grains 31 is achieved by using multiple memory channels. The DMA controller achieves efficient data interaction with the storage module 3, and the near-memory computing unit 23 and the on-chip buffer 24 cooperate to complete the efficient computing process of the memory-intensive task.
[0087] 4. After the near-memory computing task is executed, the near-memory computing coprocessor 2 performs Cache refresh to complete the synchronization between the Cache content and the memory, and then returns the control right of the near-memory computing coprocessor 2 to the bus, and generates a response signal based on the communication protocol decoding unit 21 and informs the host processor 1.
[0088] 5. The host processor 1 releases the memory control and waits for the response signal of the near-memory computing coprocessor 2, and considers that the bus control right has been returned after detecting the response signal of the near-memory computing coprocessor 2. Then, a corresponding control command is sent to the storage grain 31 to verify whether the bus is connected, and if the reply of the storage grain 31 is received again, it means that the bus control is re-received, and the second memory controller 11 of the host processor 1 continues to complete the memory operation of the computing program running process.
[0089] In this embodiment, the step of returning the control right to the host processor 1 after the memory access is completed is further configured as:
[0090] sending a control right return instruction to the host processor 1.
[0091] The host processor 1 is further configured to:
[0092] receive the control right return instruction, and send a control instruction to the storage grain 31; if the reply signal of the storage grain 31 can be received, the control right has been returned to the host processor 1.
[0093] In this embodiment, the host processor 1 is configured with a memory control word lookup table; the memory control word lookup table includes: address intervals, memory control words, and calibration parameters of each storage grain 31; the second memory controller 11 stores the memory control words and the calibration parameters.
[0094] The host processor 1 is configured to:
[0095] The access address is obtained; the access address includes an address range of the storage particle 31 to be accessed; based on the access address, the memory control word lookup table is used to determine the memory control word and calibration parameters corresponding to the storage particle 31; and the second memory controller 11 is further configured to generate the memory control instruction and near-computing configuration information based on the memory control word and calibration parameters.
[0096] For example, the main processor 1 and the near-computing coprocessor 2 have a unified physical address size, which is determined by the hardware connection of the near-computing coprocessor system based on the near-computing coprocessor 2 and does not change with the switching of the control right of the storage particle 31 between the main processor 1 and the near-computing coprocessor 2 during the working process. During the working process, the main processor 1 completes the management of the memory space in a virtual address, and the operating system constructs a virtual address mapping table to realize the address space allocation for each process in a paging mechanism. When the main processor 1 delivers the control right of the storage particle 31 to the near-computing coprocessor 2, the memory page corresponding to the storage particle 31 on the virtual address mapping table is frozen. During the near-computing task configuration, the second memory controller 11 of the main processor 1 outputs the virtual address to the actual address conversion, and the near-computing coprocessor 2 directly performs the reading operation in a physical address during the near-computing working process, which is similar to the bare machine architecture of the embedded system.
[0097] In the transparent mode, the plurality of storage particles 31 connected to the near-computing coprocessor 2 can complete the overall control process through the second memory controller 11 of the main processor 1, and the storage addresses of the storage particles 31 are implemented to be superimposed in the high bits to complete the storage space expansion. In the transparent mode, the near-computing coprocessor 2 functions similarly to a storage expander, realizes the expansion of the memory interface of the main processor 1, and increases the number of storage particles 31 accessed thereby. In the transparent mode, the main processor 1 can control the multiplexer (MUX) to selectively activate the storage particles 31 by writing control instructions into the internal processing unit 22 of the near-computing coprocessor 2, so that any number of storage particles 31 participate in the system operation of the main processor 1. The second memory physical layer interface 27 of the near-computing coprocessor 2 has a bidirectional driving buffer to ensure the driving capability of the transparent memory signal and improve the signal integrity. The main processor 1 calculates the access target based on the current access address, configures the control word of the memory controller, establishes the connection with each storage particle 31, and realizes the access operation on each storage particle 31.
[0098] Specifically, the plurality of storage particles 31 connected to the near-memory computing coprocessor 2 can have different specifications, but need to meet the interface standard to ensure that the host processor 1 can complete the corresponding access through the second memory controller 11. The second memory controller 11 of the host processor 1 needs to establish a connection with each memory particle 31 based on the address range of the access when switching between the memory particles 31, and in this process, the master-slave second memory physical layer interface 27 needs to perform an adaptive dynamic calibration process to meet the timing requirements of the memory protocol, and finally generate a set of memory controller control words to adjust the phase, amplitude and other attributes of the interface signal to ensure the integrity of the signal transmission, and finally establish a stable connection between the master-slave interfaces. When the near-memory computing coprocessor 2 works in a transparent transmission mode, the second memory controller 11 of the host processor 1 needs to switch between the plurality of storage particles 31 extended by the near-memory computing coprocessor 2 based on the memory address space. To realize fast link establishment between the master-slave interfaces, the host processor 1 divides the storage area to build a second memory controller 11 control word lookup table to realize fast configuration of the second memory controller 11. The memory controller built-in storage unit at the host processor 1 end can cache the first memory physical layer interface 251 and the memory controller end training setting and calibration parameters, and quickly configure in the link establishment process to speed up the storage link training speed and improve the training time. Figure 6 The control logic of the host processor 1 in the near-memory computing coprocessor system is shown in FIG. 14. Four storage particles 31 are extended by the near-memory computing coprocessor 2, and their capacities are 4GB, 2GB, 2GB and 1GB respectively, and the physical address allocation and cascading are completed in order. In the transparent transmission mode, the host processor 1 completes the memory access process of instructions and data based on the running program. To realize the access of a single memory controller to different specifications of storage chips, the host processor 1 internally builds a memory control word lookup table, finds the corresponding memory control word based on the accessed address interval to complete the configuration of the memory controller, and realizes the access of the corresponding storage particle 31. The remaining storage particles 31 not accessed are in self-refresh state to maintain the data content in the storage particles 31. When the system is powered on, the host processor 1 establishes a connection with each storage particle 31 by running the power-on process, establishes the port connection through adaptive dynamic calibration, and stores the training parameters obtained at the first memory physical layer interface 251 end in the table. In the memory access process, the host processor 1 sends read and write signals through the second memory controller 11 to control the memory interface, and the near-memory computing coprocessor 2 completes the reception, and then performs internal forwarding transmission to realize the memory access operation of the extended storage particle 31.
[0099] In this embodiment, when the number of the storage grains 31 to be accessed is equal to the total number of the storage grains 31 in the storage module 3, the near-memory computing coprocessor 2 runs in the computing mode; when the number of the storage grains 31 to be accessed is less than the total number of the storage grains 31 in the storage module 3, the near-memory computing coprocessor 2 runs in the compatible mode.
[0100] When the near-memory computing coprocessor 2 runs in the computing mode, the host processor 1 does not have the control right of the storage grains 31; when the near-memory computing coprocessor 2 runs in the compatible mode, the host processor 1 has the partial control right of the storage grains 31.
[0101] In this embodiment, in order to further improve the link building speed, the near-memory computing coprocessor 2 described in the present application internally integrates a dedicated cache unit, builds a calibration parameter lookup table of the near-memory computing coprocessor 2, completes the configuration of the control word of the storage interface integrated by the near-memory computing coprocessor 2 and the training result, realizes the rapid establishment of the storage interface link in the switching process, and further reduces the switching delay. In addition to the dedicated cache unit, the near-memory computing coprocessor internally integrates a dedicated configuration circuit, which can perform parallel configuration operations on each memory controller and storage interface under the control signal of the internal processing unit 22. The configuration process is completed through the write register mode of the interface, which further shortens the interface training and calibration delay and reduces the waiting time generated in the memory chip control right switching process.
[0102] In the computing mode, the host processor 1 can realize the enable operation of the corresponding selector by writing the corresponding control word to the processing unit 22 of the near-memory computing coprocessor 2, so as to give the control right of the storage grains 31 to the first memory controller 25 of the near-memory computing unit 23. In the computing process, the near-memory computing coprocessor 2 completes the corresponding access and computing behavior under the control of the processing unit 22 based on the written near-memory computing configuration. Since the near-memory computing coprocessor 2 internally integrates multiple first memory controllers 25, the expansion of the storage channel is realized, that is, the parallel access to each storage grain 31 can be completed in the computing process. The near-memory computing unit 23 in the near-memory computing coprocessor 2 is composed of a parallel computing array, including a matrix computing unit, a vector computing unit and a floating point computing unit, which realizes parallel high-throughput computing in cooperation with the on-chip buffer 24 and high access bandwidth. After the near-memory computing is completed, the internal processing unit 22 of the near-memory computing coprocessor 2 controls the internal high-speed selector 26 to return the control right of the storage grains 31. The handover process of the control right of the storage grains 31 between the near-memory computing coprocessor 2 and the host processor 1 can be completed through the polling access of the host processor 1 to the storage grains 31, or can be realized through the interrupt operation caused by the response signal (such as the Alert signal in the JEDEC protocol of the DDR grain) in the storage protocol.
[0103] Specifically, by configuring near-memory computing information, the main processor 1 can selectively release the storage granules 31. Figure 7 Taking the scenario shown as an example, the data required for near-memory computation is already stored in storage particles (2-4). After completing the near-memory computation configuration, the main processor 1 releases control of storage particles (2-4) to the near-memory computation coprocessor 2. The near-memory computation coprocessor 2 uses the configuration information to access the storage particles (2-4) with a three-channel memory bandwidth, completing the near-memory computation. The main processor 1 still retains control of storage particle (1) and executes the program based on the physical address corresponding to storage particle (1). At this time, the main processor 1 and the near-memory computation coprocessor 2 achieve collaborative execution. During the near-memory computation process, the main processor 1 periodically checks the status of the near-memory computation coprocessor 2 through program polling. When the near-memory computation coprocessor 2 finishes execution, the processing unit 22 controls the high-speed selector 26 to return control of storage particle 31. After the main processor 1 detects that the near-memory computation coprocessor 2 has finished execution through polling, it reads the near-memory computation result from the specified address space according to the near-memory computation configuration.
[0104] In this embodiment, before the main processor 1 receives the control return instruction or the near-memory computing coprocessor 2 receives the memory control instruction, the computing data cached in the main processor 1 or the near-memory computing coprocessor 2 is updated to the storage particle 31; after the computing data is updated, the control of the storage particle 31 is returned to the main processor 1 or the near-memory computing coprocessor 2 is given control of the storage particle 31.
[0105] Specifically, before the main processor 1 receives the control relinquishment instruction, it updates the computation data cached in the near-memory computing coprocessor 2 to the storage particle 31; after the computation data update is complete, it returns control of the storage particle 31 to the main processor 1. Before the near-memory computing coprocessor 2 receives the memory control instruction, it updates the computation data cached in the main processor 1 to the storage particle 31; after the computation data update is complete, the near-memory computing coprocessor 2 acquires control of the storage particle 31.
[0106] Specifically, through the above data updating process, the memory page table data taken out from the storage grain 31 is processed in the main processor 1 or the near-memory computing coprocessor 2, and the data in the memory page table data may be changed during the processing. The memory page table data is stored in the storage grain 31 again during the memory page table replacement, so as to overwrite the previous page table and ensure the consistency of the memory page table data. Before the main processor 1 or the near-memory computing coprocessor 2 loses the control right of the storage grain 31, all data in the cache of the main processor 1 or the near-memory computing coprocessor 2 is stored back into the storage grain 31. It is ensured that the main processor 1 or the near-memory computing coprocessor 2 can load the latest data for calculation during the working process after the switching.
[0107] In this embodiment, when the amount of data processed by the main processor 1 is less than or equal to a preset data amount, the near-memory computing coprocessor 2 operates in the transparent transmission mode; and the main processor 1 is configured to send the configuration obtaining instruction to the near-memory computing coprocessor 2.
[0108] The near-memory computing coprocessor 2 is further configured to:
[0109] receive the configuration obtaining instruction, so that the near-memory computing unit 23, the on-chip buffer 24, the first memory controller 25, and the first memory physical layer interface 251 enter the sleep state; and control the high-speed selector 26 to select the configuration obtaining instruction sent by the second memory controller 11 through the processing unit 22; and according to the configuration obtaining instruction, access the data stored in the storage grain 31.
[0110] Specifically, the transparent transmission mode of the near-memory computing coprocessor 2 is as shown in Figure 5 The main processor 1 sends the configuration obtaining instruction to the near-memory computing coprocessor 2. The internal processing unit 22 of the near-memory computing coprocessor 2 sends a control signal to make the high-speed selector 26 select the high-speed signal sent by the second memory controller 11 of the main processor 1. The near-memory computing unit 23, the on-chip buffer 24, the first memory controller 25, and the first memory physical layer interface 251 of the near-memory computing coprocessor 2 are in the sleep state. The memory control instruction sent by the main processor 1 is transparently transmitted through the near-memory computing coprocessor 2, and directly acts on the storage grain 31 to complete the read-write control. In the transparent transmission mode, the main processor 1 realizes the access to the four storage grains 31 through a single memory channel, and the near-memory computing coprocessor 2 realizes the expansion of the storage space of the main processor 1.
[0111] Exemplarily, in the above computing mode, pass-through mode, and compatible mode, the main processor 1 can complete data interaction with the near-memory computing coprocessor 2 based on the second memory controller 11 and the storage interface, and follow the storage communication protocol. The interaction between the near-memory computing coprocessor 2 and the main processor 1 is completely completed through reading and writing of a specific storage address segment in the storage grain 31. The overall operation process of the near-memory computing is based on the near-memory computing configuration information written by the main processor 1 before the computing starts. The near-memory computing result is written into the corresponding address segment according to the configuration information, and the main processor 1 reads the computing result from the corresponding address segment after regaining the memory control right.
[0112] In this embodiment, the main processor 1 is further configured to:
[0113] send an address access request to the communication protocol decoding unit 21, analyze the storage grain 31 to be accessed through the communication protocol decoding unit 21, interrupt the access between the first memory controller 25 and the storage grain 31 if the main processor 1 does not enjoy the control right of the storage grain 31, and establish the access between the main processor 1 and the storage grain 31.
[0114] Specifically, during the execution of the near-memory computing task, the main processor 1 can complete the access to the released storage grain 31 through the near-memory computing coprocessor 2. During the execution of the main processor 1, the second memory controller 11 sends a corresponding memory signal through the third memory physical layer interface 111. Since the main processor 1 has released the control right of the storage grain 31, the high-speed selector 26 bypasses the memory signal of the main processor 1 and cannot directly control it. The access signal is acquired by the communication protocol decoding unit 21 in the near-memory computing coprocessor 2, and the detailed information of the access request is obtained after analysis. The communication protocol decoding unit 21 judges the memory information and generates an interrupt request. The processing unit 22 in the near-memory computing coprocessor 2 responds to the interrupt request, calls the interrupt service program, and completes the access to the storage grain 31 based on the information provided by the communication protocol decoding unit 21. After the access is completed, the communication protocol decoding unit 21 generates a corresponding memory control instruction to realize the interaction with the second memory controller 11 of the main processor 1 and realize the access operation of the main processor 1. The processing unit 22 of the near-memory computing coprocessor 2 sets the interrupt mask state of the communication protocol decoding unit 21 based on the near-memory computing configuration information, modifies the corresponding frequency of the processing unit 22 for the access interrupt of the storage grain 31, thereby controlling the indirect access bandwidth of the main processor 1 based on the near-memory computing coprocessor 2, and balancing the storage bandwidth occupation of the main processor 1 access and the near-memory computing behavior for each storage grain 31.
[0115] In this embodiment, the near-memory computing co-processor 2 is also provided with a second memory physical layer interface 27; the second memory controller 11 is provided with a third memory physical layer interface 111; the host processor 1 is connected to the near-memory computing co-processor 2 through the third memory physical layer interface 111 connected to the second memory physical layer interface 27; and the near-memory computing co-processors 2 are connected to each other through the second memory physical layer interface 27 connected to the first memory physical layer interface 251.
[0116] The near-memory computing co-processor 2 described in the present application can realize multi-level interconnected tree expansion, thereby further improving the memory space and near-memory computing memory bandwidth of the near-memory computing co-processor system. The second memory physical layer interface 27 of the near-memory computing co-processor 2 can be connected to the third memory physical layer interface 111 of the host processor 1, or can be connected to the selector output signal of other near-memory computing co-processors 2 in the expansion structure, realizing multi-level cascade of the near-memory computing co-processor 2. In the overall expansion architecture interconnection topology of the tree structure, the host processor 1 is the root node, and the storage grain 31 is the leaf node. When two near-memory computing co-processors 2 are cascaded, the near-memory computing co-processor 2 that completes the storage signal output through the selector is called the upper co-processor, which is at a higher level in the tree diagram; the near-memory computing co-processor 2 that completes the storage signal reception through the second memory physical layer interface 27 signal is called the lower co-processor, which is at a lower level in the tree diagram. Figure 8 A two-level interconnected expansion topology of the near-memory computing co-processor 2 is shown in FIG. 2. Each near-memory computing co-processor 2 has a set of second memory physical layer interfaces 27 for receiving memory signals, and integrates two sets of memory controllers and first memory physical layer interface 251 architectures to complete storage expansion. Through two-level expansion, support for four storage grains 31 is realized. Figure 8 The first-level near-memory computing co-processor 2 is numbered 1-1 in FIG. 2, and the second-level near-memory computing co-processors 2 are numbered 2-1 and 2-2 after tree expansion. It is worth noting that in actual architecture, the number of DDR controllers and first memory physical layer interfaces 251 integrated by each near-memory computing co-processor 2 is not necessarily the same, and near-memory computing co-processors 2 with different expansion capabilities can also be cascaded, and the number of child nodes of each node in the tree structure formed at this time is not necessarily the same.
[0117] In the embodiment shown in FIG. 3, the near-memory computing co-processor 2 is provided with a second memory physical layer interface 27, and the host processor 1 is provided with a third memory physical layer interface 111. The host processor 1 is connected to the near-memory computing co-processor 2 through the third memory physical layer interface 111 connected to the second memory physical layer interface 27. The near-memory computing co-processors 2 are connected to each other through the second memory physical layer interface 27 connected to the first memory physical layer interface 251. Figure 8The near-memory computing co-processor 2 shown in the middle integrates two sets of memory controllers and a first memory physical layer interface 251, and storage expansion is completed through two levels of near-memory computing co-processor 2 cascading. When both two levels of near-memory computing co-processor 2 work in the transparent transmission mode, the main processor 1 can complete the access to all storage grains 31, which is similar to the working principle of the single near-memory computing co-processor 2 in the transparent transmission mode. When the multi-level near-memory computing co-processor 2 works in the computing mode, the main processor 1 can configure the scheduling of the multi-level near-memory computing co-processor 2, including the access behavior and the control right allocation of the storage grain 31. For example, when the right side memory first memory physical layer interface 251 of the near-memory computing co-processor 1-1 and the second memory physical layer interface 27 above the near-memory computing co-processor 2-2 both work in the transparent transmission mode, the main processor 1 still retains the control right of the storage grain 31. During the working process, the main processor 1 and the near-memory computing co-processor 1-1 can realize indirect access to the storage grain 1-3 through the communication protocol decoding unit 21. The near-memory computing co-processor 2-1 realizes near-memory computing on the storage grain (1) and the storage grain (2) by using a double-channel access bandwidth. The near-memory computing co-processor 2-2 realizes near-memory computing on the storage grain (3) by using a single-channel access bandwidth. It is worth noting that in the actual architecture, the number of cascaded levels of the near-memory computing co-processor 2 can be higher, and the number of child nodes of a single near-memory computing co-processor 2 can be more. During the working process, the root node main processor 1 and the higher level near-memory computing co-processor 2 in the tree structure can realize access to the lower layer leaf node storage grain 31 through direct transparent transmission or indirect access.
[0118] It is worth noting that for the near-memory computing co-processor 2 storage expansion tree interconnection topology described in the present application, a depth-limited search method is used to arrange the address space of each storage grain 31. The behavior scheduling of the main processor 1 for each level of near-memory computing co-processor 2 and the access control of the storage grain 31 are all based on the control method of the near-memory computing co-processor 2 described in the present application.
[0119] The present application provides a near-memory computing co-processor system, which has the following advantages:
[0120] 1. Low cost: the present application can directly reuse commercial storage grains 31 and corresponding access protocols, without changing the storage chip architecture, avoiding the cost of redesign or tape-out. The overall architecture only needs to complete the design of the near-memory computing co-processor 2, and complete 2D / 3D integration with commercial storage grains based on application scenarios and requirements.
[0121] 2. Good compatibility: the application can be directly used in existing computing systems, one end is connected to the second memory controller 11 of the main processor 1 and the main interface physical layer through the integrated memory from the interface, and the other end is connected to the memory main interface physical layer of multiple groups of first memory controllers 25 and multiple storage particles 31. In the transparent mode, it can be used as a traditional storage module 3, while realizing the expansion of the system storage capacity. In the computing mode, the application integrates multiple groups of memory controllers and memory interfaces to further improve the number of storage particle 31 interconnections and the computing efficiency of storage-intensive tasks. There is no need to modify other parts of the system to adapt to the near-memory computing coprocessor 2.
[0122] 3. Good flexibility: the near-memory computing coprocessor 2 described in the application and the main processor 1 and the storage particle 31 in the existing computer form a near-memory computing coprocessor system. The overall system can work in multiple states. When the near-memory computing coprocessor 2 runs in the transparent mode, the main processor 1 can directly access multiple storage particles 31 connected to the near-memory computing coprocessor 2 through its memory controller, realizing the function of storage expansion. When the near-memory computing coprocessor 2 runs in the computing mode, the main processor 1 releases the control of the storage particles 31 to the near-memory computing coprocessor 2 to complete the near-memory computing function, and the released storage particles 31 are controllable. When the control of the storage particles 31 belongs to the near-memory computing coprocessor 2, the main processor 1 can still complete the indirect memory access operation by controlling the near-memory computing coprocessor 2, and the indirect memory access bandwidth is controllable. The interaction between the main processor 1 and the near-memory computing coprocessor 2 and the storage particles 31 is realized through address access operation, all signals are transmitted through the storage interface, and the actual system has high flexibility during operation.
[0123] The above detailed description of the embodiments of the application further describes the purposes, technical solutions and beneficial effects of the embodiments of the application. It should be understood that the above is only a specific embodiment of the application, and is not used to limit the protection scope of the embodiments of the application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the embodiments of the application should be included in the protection scope of the embodiments of the application.
Claims
1. A near-memory computing coprocessor system, comprising: The application relates to a main processor (1), a near-memory computing coprocessor (2) and a storage module (3) connected in communication. The near-memory computing coprocessor (2) comprises: a communication protocol decoding unit (21) configured to receive a memory control instruction sent by the main processor (1) and to parse the memory control instruction to obtain executable decoding; the memory control instruction comprises a near-memory computing starting instruction and a configuration acquisition instruction; a processing unit (22) configured to control the behavior of the near-memory computing coprocessor (2) according to the executable decoding, so that the near-memory computing coprocessor (2) operates in a set mode; the set mode comprises a computing mode, a transparent mode and a compatible mode; a near-memory computing unit (23) configured to perform near-memory computing; an on-chip buffer (24) configured to cache data to be subjected to near-memory computing; a plurality of first memory controllers (25) connected to the storage particles (31) through a high-speed selector (26); the number of the first memory controllers (25) is equal to the number of the storage particles (31); the first memory controllers (25) are configured to independently access the data stored in the storage particles (31); and the high-speed selector (25) is configured to determine the access control right of the main processor (1) or the near-memory computing coprocessor (2) to the storage particles (31). The main processor (1) is provided with a second memory controller (11); the second memory controller (11) is configured to send a memory control instruction to the near-memory computing coprocessor (2); 2. The near-memory computing coprocessor system of claim 1, wherein, The first memory controller (25) is provided with a first memory physical layer interface (251); The storage particle (31) is provided with a storage particle interface (311); the first memory controller (25) is connected to the storage particle (31) through the storage particle interface (311) and the first memory physical layer interface (251); the specification of the storage particle interface (311) is the same as that of the first memory physical layer interface (251). When the amount of data processed by the main processor (1) is greater than a preset data amount, the main processor (1) generates a near-memory computing starting instruction and near-memory computing configuration information and sends the near-memory computing starting instruction and the near-memory computing configuration information to the communication protocol decoding unit (21) and the on-chip buffer (24), respectively; 3. The near-memory computing coprocessor system of claim 2, wherein, The communication protocol decoding unit (21) is configured to decode the near-memory computing starting instruction into a starting decoding and send the starting decoding to the processing unit (22) to start the processing unit (22); The processing unit (22) is configured to: release the control right of the main processor (1) to the storage particles (31) according to the starting decoding. According to the near-computing configuration information, the first memory physical layer interface (251) corresponding to the memory particle interface (311) is activated, and data stored in the memory particle (31) corresponding to the memory particle interface (311) is accessed; After the to-be-accessed data is accessed, the control right is returned to the main processor (1).
4. The near-memory computing coprocessor system of claim 3, wherein, The step of returning the control right to the main processor (1) after the to-be-accessed data is accessed is further configured to: send a control right return instruction to the main processor (1); The main processor (1) is further configured to: receive the control right return instruction and send a control instruction to the memory particle (31); if a reply signal of the memory particle (31) can be received, the control right has been returned to the main processor (1).
5. The near-memory computing coprocessor system of claim 3, wherein, The main processor (1) is configured with a memory control word lookup table; The memory control word lookup table includes: address intervals, memory control words, and calibration parameters of each memory particle (31); The second memory controller (11) stores the memory control words and calibration parameters; The main processor (1) is configured to: obtain an access address; the access address includes an address interval of the memory particle (31) to be accessed; based on the access address, the memory control word lookup table is used to determine the memory control word and the calibration parameter corresponding to the memory particle (31); The second memory controller (11) is further configured to: based on the memory control word and the calibration parameter, the memory control instruction and the near-computing configuration information are generated.
6. The near-memory computing coprocessor system of claim 4, wherein, When the number of to-be-accessed memory particles (31) is equal to the total number of memory particles (31) in the storage module (3), the near-computing coprocessor (2) runs in a computing mode; when the number of to-be-accessed memory particles (31) is less than the total number of memory particles (31) in the storage module (3), the near-computing coprocessor (2) runs in a compatible mode; When the near-computing coprocessor (2) runs in the computing mode, the main processor (1) does not enjoy the control right of the memory particle (31); when the near-computing coprocessor (2) runs in the compatible mode, the main processor (1) enjoys part of the control right of the memory particle (31).
7. The near-memory computing coprocessor system of claim 6, wherein, Before the main processor (1) receives the control right return instruction or the near-computing coprocessor (2) receives the memory control instruction, the calculation data cached in the main processor (1) or the near-computing coprocessor (2) is updated to the memory particle (31); after the calculation data is updated, the control right of the memory particle (31) is returned to the main processor (1) or the near-computing coprocessor (2) enjoys the control right of the memory particle (31).
8. The near-memory computing coprocessor system of claim 3, wherein, When the amount of data processed by the main processor (1) is less than or equal to a preset data amount, the near-computing coprocessor (2) runs in the transparent mode; the main processor (1) is configured to send the configuration acquisition instruction to the near-computing coprocessor (2); The near-computing coprocessor (2) is further configured to: The configuration obtaining instruction is received, so that the near-memory computing unit (23), the on-chip buffer (24), the first memory controller (25), and the first memory physical layer interface (251) interface enter a sleep state; and the high-speed selector (25) is controlled by the processing unit (22) to select the configuration obtaining instruction sent by the second memory controller (11) for reception; According to the configuration obtaining instruction, data stored in the storage grain (31) is accessed.
9. The near-memory computing coprocessor system of claim 3, wherein, The main processor (1) is further configured to: Send an address access request to the communication protocol decoding unit (21), analyze it through the communication protocol decoding unit (21), and determine the storage grain (31) to be accessed; If the main processor (1) does not have control over the storage grain (31), interrupt the access between the first memory controller (25) and the storage grain (31); and establish access between the main processor (1) and the storage grain (31).
10. The near-memory computing coprocessor system of claim 2, wherein, The near-memory computing coprocessor (2) is further provided with a second memory physical layer interface (27); and the second memory controller (11) is provided with a third memory physical layer interface (111); The third memory physical layer interface (111) is connected to the second memory physical layer interface (27), so that the main processor (1) is connected to the near-memory computing coprocessor (2); The second memory physical layer interface (27) is connected to the first memory physical layer interface (251), so that the near-memory computing coprocessors (2) are cascaded with each other.