Processor memory sharing method based on Crossbar chip architecture
Through the processor shared memory method based on Crossbar chip architecture, the chip design architecture in the existing technology has solved the problems of high latency, large-scale data processing and memory access, high energy consumption and high cost, and has achieved efficient and low-latency data exchange and memory sharing, supporting the computing needs of artificial intelligence large models.
Patent Information
- Application Number
- CN202510087618.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2025-06-03
AI Technical Summary
When facing large-scale data processing and memory access, the existing chip design architecture has problems such as high latency, large energy consumption, and high design and manufacturing costs, which is difficult to meet the needs of artificial intelligence large models for memory space and data processing capabilities.
The processor shared memory method based on the Crossbar chip architecture is adopted, and the processor can share access to intermediate-level memory and remote memory through the Crossbar structure. The multi-port Crossbar structure provides extremely low-latency data exchange and supports the memory space required for large-scale computing.
It realizes a simple and easy-to-implement chip design architecture, and data exchange with extremely low latency reduces the memory access delay of heterogeneous processors during computing, supports memory space requirements for large-model computing, and reduces design and manufacturing costs.
Smart Images

Figure CN120086178A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semiconductor integrated circuits, and in particular to a method for a processor to share memory based on a Crossbar chip architecture. Background Art
[0002] With the development of technology, the application of artificial intelligence is becoming more and more extensive. The training of large artificial intelligence models requires a large amount of memory space and efficient data processing capabilities. Therefore, the importance of chip design architectures in the field of artificial intelligence is becoming increasingly prominent. That is, designing a chip architecture that can both meet the demand for a large amount of memory space and reduce memory access latency has become an important research direction.
[0003] In addition, microelectronics packaging technology and high-speed network communication devices are also important factors affecting chip performance. That is, the design and manufacturing levels of chip design architectures, microelectronics packaging technology, and high-speed network communication devices directly affect the performance and cost of products.
[0004] Existing chip design architectures mainly use a central processing unit (CPU) as the core and connect each functional module through a bus structure. This architecture can meet basic computing requirements to a certain extent. However, with the emergence of large artificial intelligence models, the demand for memory space and data processing capabilities is increasing, and existing chip design architectures are increasingly unable to meet these demands.
[0005] In addition, existing chip design architectures have problems of high latency, high energy consumption, and high design and manufacturing costs when facing large-scale data processing and memory access.
[0006] The disclosure of the above background art content is only used to assist in understanding the concept and technical solution of the present application. It does not necessarily belong to the prior art of the present application, nor does it necessarily provide technical guidance. Without clear evidence that the above content was publicly available before the filing date of the present application, the above background art should not be used to evaluate the novelty and inventiveness of the present application. Summary of the Invention
[0007] The object of the present invention is to provide an improved method for a processor to share memory based on a Crossbar chip architecture. In addition to exclusive proximal memory, the processor can share access to intermediate-level memory or remote memory to support the memory space required for large model operations.
[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0009] A method for a processor to share memory based on a Crossbar chip architecture. The processor realizes the sharing of intermediate-level memory and remote memory based on a chip with a Crossbar structure. Among them, the chip includes a single Crossbar and multiple processor ports, a memory processing unit, and a network data processing unit that perform data exchange through the Crossbar. The Crossbar has a switching matrix; multiple processors are connected to the corresponding processor ports, the intermediate-level memory is connected to the corresponding memory processing unit, and the remote memory is set on a network device and connected to the corresponding network data processing unit;
[0010] The sharing of the intermediate-level memory and the remote memory is realized in the following way:
[0011] When a processor requests to write data to a target remote memory, the processor first writes the data into the intermediate-level memory Z in an idle state through the Crossbar structure, and then controls the switch on the path between the intermediate-level memory Z and the target remote memory in the switching matrix of the Crossbar to close, so that the data written into the intermediate-level memory Z is transmitted to the target remote memory;
[0012] When a processor requests to read data from a target remote memory, the data is first transmitted to the intermediate-level memory X in an idle state through the Crossbar structure by means of RDMA, and then the switch on the path between the intermediate-level memory X and the processor in the switching matrix of the Crossbar is controlled to close, so that the processor can read the data temporarily stored in the intermediate-level memory X.
[0013] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the processor writes data into the intermediate-level memory Z in an idle state through the Crossbar structure in the following way:
[0014] Taking the processor port corresponding to the processor as the Crossbar input port and the memory processing unit corresponding to the intermediate-level memory Z as the Crossbar output port, control the switch on the path for conducting the input port and the output port in the switching matrix of the Crossbar to close.
[0015] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, when the memory space requirement of the processor is greater than the proximal memory of the processor, the processor sends a request to the Crossbar to share the intermediate-level memory. By looking up the table in the computer, if there is an idle intermediate-level memory Y, then taking the processor port corresponding to the processor as the Crossbar input port and the memory processing unit corresponding to the intermediate-level memory Y as the Crossbar output port, control the switch on the path for conducting the input port and the output port in the switching matrix of the Crossbar to close.
[0016] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the processor includes a CPU and a GPU, and the CPU and the GPU are connected to their respective corresponding processor ports;
[0017] Each CPU is directly connected to the corresponding proximal memory outside the Crossbar structure, and each GPU is directly connected to the corresponding proximal memory outside the Crossbar structure.
[0018] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the GPU does not access the intermediate-level memory and the remote memory through the CPU.
[0019] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the proximal memory includes a proximal DDR memory and a proximal HBM memory. Among them, the proximal DDR memory is connected to the CPU, and the proximal HBM memory is connected to the GPU in a one-to-one correspondence.
[0020] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the processor port includes at least one PCIe port and multiple communication protocol ports. Among them, the PCIe port is docked to the CPU, and the communication protocol ports are connected to the GPU in a one-to-one correspondence.
[0021] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, adjust the quantity allocation of the processor port, the memory processing unit, and the network data processing unit according to the multi-port structure of the Crossbar structure.
[0022] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the intermediate-level memory includes an intermediate-level DDR memory and an intermediate-level HBM memory. Some memory processing units are connected to the intermediate-level DDR memory, and the remaining memory processing units are connected to the intermediate-level HBM memory.
[0023] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the memory processing unit includes one or more DDR controllers and one or more HBM controllers. Among them, the DDR controller is connected to the intermediate-level DDR memory, and the HBM controller is connected to the intermediate-level HBM memory.
[0024] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the DDR controller is connected to the intermediate-level DDR memory through a PCIe bus, and the HBM controller is connected to the intermediate-level HBM memory through a communication protocol bus.
[0025] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the remote memory is configured on an Ethernet device or an Infiniband device according to the network communication mode of the network data processing unit.
[0026] The beneficial effects brought by the technical solutions provided by the present invention are as follows:
[0027] a. Provide a concise and easy-to-implement chip design architecture centered on data exchange functions. This architecture utilizes a multi-port Crossbar structure to provide data exchange with extremely low latency, ensuring the smoothness of data exchange and greatly reducing the latency of memory access during the operation of heterogeneous processors.
[0028] b. Integrate and expand the function of shared memory: In addition to the exclusive proximal memory in the chip architecture, the processor can share access to the intermediate-level memory or the remote memory to support the memory space required for large model operations.
[0029] c. Provide an improved chip architecture suitable for heterogeneous processors, changing the traditional CPU-centered architecture to a GPU processor mainly focused on computing power, adapting to the computing of large-scale data processing and ensuring the stability of the overall system.
[0030] d. With the concept of System on Chip, the chip architecture is miniaturized to a high-density system-on-chip, which has important practical value for power-sensitive system devices: Under a single package structure, the latency of data exchange can be controlled, reducing the operating cost formed by the high power consumption derived from a large number of operations in the artificial intelligence computing center. At the same time, it is also beneficial to improve the performance and stability of the system. Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0032] Figure 1 Schematic diagram of the structure of the Crossbar data exchange center provided for an exemplary embodiment of the present invention;
[0033] Figure 2 Schematic diagram of the chip architecture provided for an exemplary embodiment of the present invention;
[0034] Figure 3 Schematic diagram of the multi-level memory architecture provided for an exemplary embodiment of the present invention;
[0035] Figure 4 Flowchart of a chip architecture design method provided for an exemplary embodiment of the present invention. Detailed implementation manners
[0036] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0038] The present invention aims to provide an improved chip design architecture centered on data exchange function. This architecture utilizes a multi-port Crossbar structure and can provide extremely low-latency data exchange.
[0039] In an embodiment of the present invention, a chip architecture integrating heterogeneous processor shared memory is provided, including a CPU, a GPU, a memory unit, multiple processor ports arranged based on a Crossbar structure, one or more memory processing units, and a network data processing unit. Data exchange is performed between the processor ports, the memory processing units, and the network data processing units through the Crossbar structure; the Crossbar structure is as Figure 1 shown. The Crossbar structure is an efficient switching structure widely used in the computer field, that is, efficient data routing and switching are achieved between multiple ports: data enters from the input port and is routed to the corresponding output port through the switch matrix. Figure 1 The 4×3 switch matrix in is only for illustration. The present invention does not limit the number of switches in the switch matrix of the Crossbar, that is, it does not limit the number of ports / processing units connected thereto. Figure 1The port / processing unit in [the relevant context] is the processor port, memory processing unit, and network data processing unit in this embodiment.
[0040] In a specific embodiment, such as Figure 2 The Crossbar structure shown has 16 ports and can adopt a 4×4 switch matrix. Taking the chip architecture shown as an example for illustration: Figure 2 The chip architecture shown is used as an example for illustration:
[0041] In this architecture, the port / processing units of the 16 Crossbars are respectively 8 processor ports (1 PCIe port, 7 communication protocol ports), 4 memory processing units (2 DDR controllers, 2 HBM controllers), and 4 network data processing units; obviously, in different embodiments, the quantities of PCIe ports, communication protocol ports, DDR controllers, HBM controllers, and network data processing units can be adjusted and variably allocated.
[0042] Although Figure 2 it is shown in [the relevant context] that only one of the 8 processor ports is connected to the CPU processor, and the remaining seven are connected to the GPU processor, which is just a schematic illustration and does not limit such quantity or proportional relationship. However, in this embodiment, the number of GPUs being more than the number of CPUs is used to reflect a newly evolved chip architecture with the GPU processor, which focuses on computing power, as the core, replacing the traditional chip architecture with the CPU as the core. This enables the newly evolved chip architecture in this embodiment to better adapt to the computing power requirements of artificial intelligence large models and effectively alleviates the problem of high latency that occurs when the traditional CPU-core architecture processes large-scale data. In contrast, the data processing efficiency of the chip architecture in this embodiment is higher.
[0043] The connection relationship between the Crossbar structure of the chip and the external processors and memory units of the chip is as follows:
[0044] The CPU and GPU are connected to the corresponding processor ports: the CPU is connected to the PCIe port, and the GPU is connected to the communication protocol ports one by one; in this embodiment, the communication protocol can adopt a public protocol or a private protocol as the basis for jobs for memory access to ensure the reliability and efficiency of data transmission;
[0045] The memory unit includes proximal memory, intermediate-level memory, and distal memory. Among them, the proximal memory has the shortest latency and is called Level1 memory. As shown in Figure 3 The proximal memory includes proximal DDR memory and proximal HBM memory. Among them, the proximal DDR memory is connected to the CPU, and the proximal HBM memory is connected to the GPU one by one;
[0046] The latency of the intermediate-level memory is the second, and it is called the memory of Level 2. It is connected to the corresponding memory processing unit. For example, Figure 3 As shown, according to the memory processing unit, it is divided into a DDR controller and an HBM controller. Correspondingly, the intermediate-level memory includes an intermediate-level DDR memory and an intermediate-level HBM memory. The DDR controller is connected to the intermediate-level DDR memory through a PCIe bus, and the HBM controller is connected to the intermediate-level HBM memory through a communication protocol bus;
[0047] The remote memory is set on the network device, and the network device is connected to the corresponding network data processing unit; the latency of the remote memory is the longest, and it is called the memory of Level 3. The network data processing unit has the ability of Remote Direct Memory Access (RDMA). If the network data processing unit uses the RDMA on Ethernet communication method, the remote memory is configured on the Ethernet device; if the network data processing unit uses the RDMA on Infiniband communication method, the remote memory is configured on the Infiniband device.
[0048] Based on the connection between the above Crossbar structure and the memory unit, both the CPU and the GPU are configured to be able to access the intermediate-level memory and the remote memory through the Crossbar structure. The first scenario of the processor sharing memory based on the Crossbar chip architecture is as Figure 4 shown: When the memory space requirement is greater than the proximal memory, or when there is data that needs to be exchanged through the shared memory, the corresponding processor (CPU / GPU) sends a request to the Crossbar to share the intermediate-level memory. For example, according to the latency priority, it first requests to access the intermediate-level memory of Level 2. By looking up the table in the computer, it can be determined that there is an idle intermediate-level memory Y, then it can be successfully accessed by the processor that sends the sharing request to meet the large memory space requirement; in this case, if the processor is a CPU / GPU, the corresponding PCIe port / communication protocol port is the input port of the Crossbar, and the DDR controller / HBM controller corresponding to the idle intermediate-level memory is the output port of the Crossbar. Then, the switch on the path in the control switch matrix that conducts the two is closed to realize the bidirectional conduction of the path from the processor to the intermediate-level memory Y, and finally realize the sharing of the intermediate-level memory Y by the processor.
[0049] Continue to refer to Figure 4The second scenario of processor - shared memory based on the Crossbar chip architecture: In a communication scenario, when a processor requests to write data to a remote memory so that the data can be further sent to the outside through a network device. In this case, the processor first writes the data to an intermediate - level memory in an idle state, and then moves the data from the intermediate - level memory to the remote memory. That is, first, the computer can determine through table - lookup that there is an intermediate - level memory Z in an idle state, then control the switches on the path between the processor and the intermediate - level memory Z in the switch matrix of the Crossbar to close, and then the processor completes the data writing; then control the switches on the path between this intermediate - level memory Z and the remote memory in the switch matrix of the Crossbar to close, so that the written data is moved from the intermediate - level memory Z to the remote memory.
[0050] The above is the scenario where the processor requests to write data. There is also a scenario where, based on communication requirements, data in the remote memory needs to be read. The third scenario of processor - shared memory based on the Crossbar chip architecture is as Figure 4 shown. When the required data is stored in the remote shared memory (i.e., the remote memory of Level 3), data can be moved to the proximal shared memory (i.e., the intermediate - level memory of Level 2) through RDMA: First, the computer can determine through table - lookup that there is an intermediate - level memory X in an idle state, then control the switches on the path between the remote memory and the intermediate - level memory X in the switch matrix of the Crossbar to close, so that the data is moved from the remote memory to the intermediate - level memory X, and then control the switches on the path between the processor and the intermediate - level memory X in the switch matrix of the Crossbar to close, finally realizing the sharing of this remote memory by the processor and reducing the latency of data access.
[0051] The chip architecture of this embodiment can meet the requirements of a large amount of memory space when running an artificial - intelligence large model, and can expand the memory elastically according to the requirements; and can set the priority level of memory expansion according to the speed of latency, so as to shorten the latency of accessing the expanded memory and improve the efficiency.
[0052] This embodiment can adopt a low - power design concept, further integrate the chip architecture into a high - density SoC (System on Chip), which can appropriately reduce the operating cost formed by the high power consumption derived from a large number of operations in the artificial - intelligence operation center, and is also conducive to improving the performance and stability of the system. And because it is integrated in a single package structure, the latency of data exchange can be controlled, which cannot be achieved by the existing discrete architectures. During the manufacturing process, advanced micro - electronic packaging technologies, such as chiplet or 3D packaging, or in a single - chip manner, are adopted to reduce the manufacturing cost of the chip, and are also conducive to improving the performance and stability of the chip.
[0053] In an embodiment of the present invention, a chip integrating heterogeneous processor shared memory is provided, including a plurality of processor ports arranged based on a Crossbar structure, one or more memory processing units, and a network data processing unit. Data exchange is performed between the processor ports, the memory processing units, and the network data processing unit through the Crossbar structure;
[0054] Compared with the chip architecture provided in the previous embodiment, the chip of this embodiment does not include processors (CPU, GPU), memory units (proximal memory, intermediate-level memory, and distal memory), and network devices.
[0055] The chip integrating heterogeneous processor shared memory in this embodiment belongs to the same concept as the chip architecture provided in the above embodiment, that is, the processor realizes access to other external memories that are not its own exclusive through the ports of the Crossbar structure of the chip, that is, realizes memory sharing, effectively solving the demand for a large amount of memory space in the training of artificial intelligence large models or other scenarios.
[0056] The processor ports of the Crossbar structure of the chip in this embodiment are configured to be connected to the CPU and GPU respectively; the memory processing unit is configured to be connected to the intermediate-level memory; the network data processing unit is configured to be connected to the network device, and the network device is configured with a distal memory;
[0057] Both the CPU and GPU are configured to be able to access the intermediate-level memory and / or distal memory through the Crossbar structure.
[0058] In an embodiment of the present invention, a chip architecture design method for integrating processor shared memory is provided. The chip architecture design method includes the following steps:
[0059] Construct a Crossbar as a data exchange center so that data interaction can be performed between the processor ports and the memory processing units through the Crossbar;
[0060] Connect the processor ports to the processors and connect the memory processing units to the intermediate-level memory;
[0061] Configure the processor to be able to access the intermediate-level memory through the Crossbar, specifically as Figure 4As shown, when the memory space requirement is greater than the proximal memory, or when there is data that needs to be exchanged through the shared memory (requesting to read or write data to the remote memory), the corresponding processor sends a request to access the intermediate-level memory to the Crossbar. Similar to the chip architecture embodiment described above, if an intermediate-level memory in an idle state is traversed, the switch on the path between the processor port corresponding to the processor in the switch matrix of the Crossbar and the memory processing unit corresponding to the idle intermediate-level memory is closed, enabling the processor to share the intermediate-level memory.
[0062] A network data processing unit can be further designed to enable data interaction with the processor port and the memory processing unit through the Crossbar;
[0063] Connect the network data processing unit to a network device, where the network device is configured with a remote memory;
[0064] Configure the processor to be able to access the remote memory through the Crossbar. Specifically, when the dedicated proximal memory of the processor is less than the current memory space requirement, the corresponding processor sends a request to access the remote memory to the Crossbar. Similar to the chip architecture embodiment described above, if a remote memory in an idle state is traversed, the switch on the path between the processor port corresponding to the processor in the switch matrix of the Crossbar and the network data processing unit corresponding to the idle remote memory is closed, enabling the processor to share the terminal memory.
[0065] In the case where there are both an intermediate-level memory and a remote memory in an idle state, the memory with a shorter latency - the intermediate-level memory - can be preferentially selected for sharing.
[0066] The chip architecture design method provided above is applicable to heterogeneous processors including CPUs and GPUs; correspondingly, the number of the processor ports is configured to be plural to be connected to one or more CPUs and GPUs in a one-to-one correspondence; and proximal memories corresponding to the CPUs and GPUs in a one-to-one correspondence are configured;
[0067] When the proximal memory is less than the current memory space requirement, the corresponding processor sends a request to access the intermediate-level memory and / or the remote memory to the Crossbar.
[0068] The processor shared memory method based on the Crossbar chip architecture provided by the present invention is applicable to fields such as artificial intelligence applications, big data processing, high-performance computing, and network communication devices. First, the chip architecture of the present invention can effectively solve the demand for a large amount of memory space during the training of artificial intelligence large models and improve the efficiency of artificial intelligence applications; second, the chip architecture of the present invention can greatly reduce the latency of memory access, improve the efficiency of big data processing, and meet the requirements of high-performance computing; third, the chip architecture of the present invention can improve the data transmission efficiency of network communication devices through RDMA technology and meet the requirements of high-speed network communication; finally, the chip design architecture of the present invention can reduce the design and manufacturing costs and is suitable for being installed in a data center or the backbone network layer of cloud computing. Because of the Crossbar structure for fast data exchange, the computing center can adapt to different computing power requirements, flexibly increase or decrease the number of GPUs in the rack without affecting the computing performance; it can use external DDR / HBM to provide additional required memory space; if there are resources on the network, it can also connect to the regional network through the network data processing unit to obtain the required data.
[0069] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0070] The above are only specific embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A processor shared memory method based on Crossbar chip architecture, characterized in that: The processor realizes sharing of intermediate memory and remote memory based on a chip with a Crossbar structure, wherein the chip includes a single Crossbar and multiple processor ports, a memory processing unit, and a network data processing unit for data exchange through the Crossbar, and the Crossbar has a switch matrix; multiple processors are connected to corresponding processor ports, the intermediate memory is connected to corresponding memory processing units, and the remote memory is set on a network device and connected to a corresponding network data processing unit; The sharing of intermediate memory and remote memory is achieved through the following methods: When the processor requests to write data into the target remote memory, the processor first writes the data into the idle intermediate memory Z through the Crossbar structure, and then controls the switch on the path between the intermediate memory Z and the target remote memory in the switch matrix of the Crossbar to close, so that the data written into the intermediate memory Z is transmitted to the target remote memory; When the processor requests to read data in the target remote memory, the data is first transmitted to the idle intermediate memory X through the Crossbar structure by RDMA, and then the switch on the path between the intermediate memory X and the processor in the Crossbar switch matrix is controlled to be closed, so that the processor can read the data temporarily stored in the intermediate memory X.
2. The processor shared memory method according to claim 1, characterized in that: The processor writes data to the idle intermediate memory Z through the Crossbar structure in the following way: The processor port corresponding to the processor is used as the Crossbar input port, and the memory processing unit corresponding to the intermediate memory Z is used as the Crossbar output port to control the closure of switches on the path for conducting the input port and the output port in the switch matrix of the Crossbar.
3. The processor shared memory method according to claim 1, characterized in that: When the memory space requirement of the processor is greater than the proximal memory of the processor, the processor sends a request for sharing the intermediate-level memory to Crossbar, and determines through a computer table lookup that there is an idle intermediate-level memory Y. The processor port corresponding to the processor is used as the Crossbar input port, and the memory processing unit corresponding to the intermediate-level memory Y is used as the Crossbar output port, and the switch on the path for conducting the input port and the output port in the switch matrix of the Crossbar is controlled to be closed.
4. The processor shared memory method according to claim 1, characterized in that: The processor includes a CPU and a GPU, and the CPU and the GPU are connected to respective corresponding processor ports; Each CPU is directly connected to the corresponding proximal memory outside the Crossbar structure, and each GPU is directly connected to the corresponding proximal memory outside the Crossbar structure.
5. The processor shared memory method according to claim 4, characterized in that: The GPU accesses the intermediate level memory and the remote memory without going through the CPU.
6. The processor shared memory method according to claim 4, characterized in that: The proximal memory includes a proximal DDR memory and a proximal HBM memory, wherein the proximal DDR memory is connected to the CPU, and the proximal HBM memory is connected to the GPU in a one-to-one correspondence.
7. The processor shared memory method according to claim 4, characterized in that: The processor port includes at least one PCIe port and a plurality of communication protocol ports, wherein the PCIe port is connected to the CPU, and the communication protocol ports are connected to the GPU in a one-to-one correspondence.
8. The processor shared memory method according to claim 1, characterized in that: The number allocation of processor ports, memory processing units and network data processing units is adjusted according to the multi-port structure of the Crossbar structure.
9. The processor shared memory method according to claim 1, characterized in that: The intermediate-level memory includes an intermediate-level DDR memory and an intermediate-level HBM memory, some memory processing units are connected to the intermediate-level DDR memory, and the remaining memory processing units are connected to the intermediate-level HBM memory.
10. The processor shared memory method according to claim 9, characterized in that: The memory processing unit includes one or more DDR controllers and one or more HBM controllers, wherein the DDR controller is connected to the intermediate-level DDR memory, and the HBM controller is connected to the intermediate-level HBM memory.
11. The processor shared memory method according to claim 10, characterized in that: The DDR controller is connected to the intermediate-level DDR memory via a PCIe bus, and the HBM controller is connected to the intermediate-level HBM memory via a communication protocol bus.
12. The processor shared memory method according to any one of claims 1 to 11, characterized in that: The remote memory is configured on an Ethernet device or an Infiniband device according to the network communication mode of the network data processing unit.