Memory sharing control method and device, computer device, and system

The memory sharing control method and device address inefficiencies in multi-core processor memory utilization by dynamically allocating shared memory resources, improving computing efficiency and reducing costs.

JP7700420B2Active Publication Date: 2025-07-01HUAWEI TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2023555494
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-31
Filing Date
2022-03-14
Publication Date
2025-07-01
Estimated Expiration
2042-03-14

AI Technical Summary

Technical Problem

The increasing demand for computing resources in big data applications and the limitations of multi-core processors due to the slowdown in semiconductor technology development lead to inefficiencies in memory utilization, with memory costs accounting for a significant portion of server operating costs.

Method used

A memory sharing control method and device that allows multiple processing units to access shared memory resources through a memory sharing control device, which allocates memory dynamically and flexibly, utilizing technologies like FPGA chips and ASICs to manage memory access and optimize resource utilization.

Benefits of technology

Improves memory resource utilization by allowing different processing units to access memory during different periods, enhancing computing efficiency and reducing total cost of ownership (TCO) through optimized memory allocation and access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700420000004
    Figure 0007700420000004
  • Figure 0007700420000005
    Figure 0007700420000005
  • Figure 0007700420000006
    Figure 0007700420000006
Patent Text Reader

Abstract

A memory sharing control method and device, a computer device, and a system for improving memory resource utilization are provided. In the computer device, a memory sharing control device is disposed between a processor and a memory pool, and the processor accesses the memory pool through the memory sharing control device. Different processing units, such as processors or cores within a processor, access at least one memory in the memory pool at different time periods, so that the memory is shared by multiple processing units and memory resource utilization is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and more particularly, to a memory sharing control method and device, and a system.

Background Art

[0002] With the popularization of big data technology, applications in various fields have increasing requirements for computing resources. Large-scale computing represented by applications such as graph computing and deep learning represents the latest application development direction. Furthermore, with the slowdown of semiconductor technology development, the scalable performance of applications cannot be continuously improved during the upgrade of processors, and multi-core processors are gradually becoming mainstream.

[0003] Multi-core processor systems have increasingly high requirements for memory capacity. As an essential component in servers, the cost of memory accounts for 30% - 40% of the total operating cost of the server. Improving the utilization rate of memory is an important means to reduce the total cost of operations (TCO).

Summary of the Invention

[0004] This application provides a memory sharing control method and device, a computer device, and a system for improving the utilization rate of memory resources.

[0005] According to a first aspect, this application provides a computer device including at least two processing units, a memory sharing control device, and a memory pool, where the processing units are processors, cores within a processor, or a combination of cores within a processor, the memory pool includes one or more memories, the at least two processing units are coupled to the memory sharing control device, The memory sharing control device is configured to separately allocate the memory from the memory pool to at least two processing units, and at least one memory in the memory pool is accessible by different processing units during different periods. At least two processing units are configured to access the memory allocated via the memory sharing control device.

[0006] At least two processing units in the computer device can access at least one memory in the memory pool during different periods via the memory sharing control device to realize memory sharing by multiple processing units, thereby improving the utilization rate of memory resources.

[0007] Optionally, the fact that at least one memory in the memory pool is accessible by different processing units during different periods means that any two of the at least two processing units can separately access at least one memory in the memory pool during different periods. For example, the at least two processing units include a first processing unit and a second processing unit. In the first period, the first memory in the memory pool is accessed by the first processing unit, and the second processing unit cannot access the first memory. In the second period, the first memory in the memory pool is accessed by the second processing unit, and the first processing unit cannot access the first memory. Optionally, the processor may be a central processing unit (CPU), and one CPU may include two or more cores.

[0008] Optionally, at least one of the at least two processing units may be a processor, a core within a processor, a combination of multiple cores within a processor, or a combination of multiple cores in different processors. A combination of multiple cores within a processor is used as a processing unit, or a combination of multiple cores in different processors is used as a processing unit. Thus, in a parallel computing scenario, multiple different cores access the same memory when executing tasks in parallel, thereby improving the efficiency of executing parallel computing by multiple different cores.

[0009] Optionally, the memory sharing control device may separately allocate memory from the memory pool to at least two processing units based on received control instructions sent by the operating system within the computer device. Specifically, a driver within the operating system may send control instructions used to allocate memory within the memory pool to at least two processing units to the memory sharing control device on a dedicated channel. The operating system is implemented by the CPU within the computer device by executing relevant code. The CPU executing the operating system has a privilege mode, and in this mode, the driver within the operating system can send control instructions to the memory sharing control device on a dedicated or designated channel.

[0010] Optionally, the memory sharing control device may be implemented by using a field programmable gate array (FPGA) chip, an application-specific integrated circuit (ASIC), or other similar chips. The circuit function of the ASIC is defined at the beginning of the design, and the ASIC is characterized by high chip integration, ease of achieving a large number of tape-outs, low cost of a single tape-out, small size, etc.

[0011] In some possible implementation manners, at least two processing units are connected to a memory sharing control device via a serial bus, the first processing unit in at least two processing units is configured to send a first memory access request in the form of a serial signal to the memory sharing device via the serial bus, and the first memory access request is used to access a first memory allocated to the first processing unit.

[0012] The serial bus has characteristics of high bandwidth and low latency. At least two processing units are connected to the memory sharing control device via the serial bus, thereby ensuring the efficiency of data transmission between the processing unit and the memory sharing control device.

[0013] Optionally, the serial bus is a memory semantic bus. The memory semantic bus includes, but is not limited to, a bus based on a quick path interconnect (QPI), a peripheral component interconnect express (PCIe), a Huawei Cache Coherence System (HCCS), or a compute express link (CXL) interconnect.

[0014] Optionally, the memory access request generated by the first processing unit is a memory access request in the form of a parallel signal. The first processing unit can convert the memory access request in the form of a parallel signal into the first memory access request in the form of a serial signal through an interface capable of realizing the conversion between the parallel signal and the serial signal, for example, a Serdes interface, and send the first memory access request in the form of a serial signal to the memory sharing device via the serial bus.

[0015] In some possible implementation manners, the memory sharing control device includes a processor interface, and the processor interface receives a first memory access request, and is configured to convert the first memory access request into a second memory access request in a parallel signal format.

[0016] The processor interface converts the first memory access request into a second memory access request in a parallel signal format, so that the memory sharing control device can access the first memory and realize memory without changing the existing memory access architecture.

[0017] Optionally, the processor interface is an interface capable of realizing the conversion between a parallel signal and a serial signal, for example, it may be a Serdes interface.

[0018] In some possible implementation manners, the memory sharing control device includes a control unit, and the control unit is configured to establish a correspondence relationship between the memory address of the first memory in the memory pool and the first processing unit in at least two processing units in order to allocate the first memory from the memory pool to the first processing unit.

[0019] Optionally, the correspondence relationship between the memory address of the first memory and the first processing unit may be dynamically adjusted. For example, the correspondence relationship between the memory address of the first memory and the first processing unit may be dynamically adjusted as needed.

[0020] Optionally, the memory address of the first memory may be a segment of consecutive physical memory addresses in the memory pool. The segment of consecutive physical memory addresses in the memory pool can simplify the management of the first memory. Obviously, the memory address of the first memory may alternatively be several segments of discontinuous physical memory addresses in the memory pool.

[0021] Optionally, the memory address information of the first memory includes the start address of the first memory and the size of the first memory. The first processing unit has an identifier, and establishing the correspondence between the memory address of the first memory and the first processing unit may be to establish the correspondence between the unique identifier of the first processing unit and the memory address information of the first memory.

[0022] In some possible implementation manners, the memory sharing control device includes a control unit, and the control unit is configured to virtualize a plurality of virtual memory devices from a memory pool. The physical memory corresponding to the first virtual memory device in the plurality of virtual memory devices is the first memory, and is configured to allocate the first virtual memory device to the first processing unit. Optionally, the virtual memory device corresponds to a segment of consecutive physical memory addresses in the memory pool. The virtual memory device corresponds to a segment of consecutive physical memory addresses in the memory pool, thereby simplifying the management of the virtual memory device. Obviously, the virtual memory device may alternatively correspond to some segments of discontinuous physical memory addresses in the memory pool.

[0023] Optionally, the first virtual memory device may be allocated to the first processing unit by establishing an access control table. For example, the access control table may include information such as the identifier of the first processing unit, the identifier of the first virtual memory device, and the start address and size of the memory corresponding to the first virtual memory device. The access control table may further include permission information for the first processing unit to access the first virtual memory device, attribute information of the accessed memory (including but not limited to information regarding whether the memory is a persistent memory), etc.

[0024] In some possible implementation manners, the control unit When a preset condition is satisfied, cancel the correspondence between the first virtual memory device and the first processing unit, and is further configured to establish a correspondence between the first virtual memory device and a second processing unit within at least two processing units.

[0025] Optionally, the correspondence between the virtual memory device and the processing unit may be dynamically adjusted based on the memory resource requirements of at least two processing units.

[0026] The correspondence between the virtual memory device and the processing unit is dynamically adjusted, so that the memory resource requirements of different processing units in different service scenarios can be flexibly adapted, and the utilization rate of memory resources can be improved.

[0027] Optionally, the preset condition may be that the memory access requirement of the first processing unit decreases and the memory access requirement of the second processing unit increases.

[0028] Optionally, the control unit When a preset condition is satisfied, cancel the correspondence between the first memory and the first virtual memory device, establish a correspondence between the first memory and a second virtual memory device within a plurality of virtual memory devices, and further configure to allocate the second virtual memory device to a second processing unit within at least two processing units. In this case, there is no need to change the correspondence between the virtual memory device and the physical memory address in the memory pool, and only the correspondence between the virtual memory device and different processing units needs to be changed, so that different processing units can access the same physical memory during different periods.

[0029] In some possible implementation manners, the memory sharing control device further includes a cache unit, and the cache unit is configured to cache data read by any one of at least two processing units from a memory pool, or to cache data evicted by any one of at least two processing units.

[0030] The efficiency of accessing memory data by a processing unit can be further improved by using the cache unit.

[0031] Optionally, the cache unit may include a level 1 cache and a level 2 cache. The level 1 cache may be a small-capacity cache with a higher read / write speed than that of the level 2 cache. For example, the level 1 cache may be a 100-megabyte (MB) nanosecond-level cache. The level 2 cache may be a large-capacity cache with a lower read / write speed than that of the level 1 cache. For example, the level 2 cache may be a 1-gigabyte (GB) dynamic random access memory (DRAM). The level 1 cache and the level 2 cache are used, so that while the data access speed of the processor can be improved by using the cache, the cache space can be increased, the range in which the processor can quickly access the memory by using the cache is extended, and the memory access rate of the processor resource pool is generally further improved.

[0032] Optionally, the data in the memory may first be cached in the level 2 cache, and then, based on the requirements of the processing unit for the memory data, the data in the level 2 cache is cached in the level 1 cache. Alternatively, to ensure that the level 1 cache has sufficient space for other processing units to cache data for use, data that does not need to be evicted or temporarily processed by the processing unit may be cached in the level 1 cache, and some of the data evicted from the level 1 cache by the processing unit may be cached in the level 2 cache.

[0033] In some possible implementation manners, the memory sharing control device further includes a prefetch engine, and the prefetch engine is configured to prefetch data that needs to be read by any one of at least two processing units from the memory pool and cache the data in the cache unit.

[0034] Optionally, the prefetch engine may implement intelligent data prediction by using a specified algorithm or a related artificial intelligence (AI) algorithm to further improve the efficiency of the processing unit accessing the memory data.

[0035] In some possible implementation manners, the memory sharing control device further includes a quality of service (QoS) engine, and the QoS engine is configured to realize optimized storage of data that needs to be cached by any one of at least two processing units in the cache unit. By using the QoS engine, different capabilities of caching memory data accessed by different processing units can be realized in the cache unit 304. For example, a memory access request initiated by a processing unit with a high priority has exclusive cache space in the cache unit 304. In this way, it can be ensured that the data accessed by the processing unit can be cached in a timely manner, thereby ensuring the service processing quality of this type of processing unit.

[0036] In some possible implementation manners, the memory sharing control device further includes a compression / decompression engine, and the compression / decompression engine is configured to compress or decompress data related to memory access.

[0037] Optionally, the function of the compression / decompression engine may be disabled.

[0038] The compression / decompression engine may compress the data written into the memory by the processing unit at a granularity of 4 kilobits (KB) per page by using a compression ratio algorithm, and then write the compressed data into the memory, or when the processing unit reads the compressed data in the memory, decompress the data to be read, and then send the decompressed data to the processor. In this way, the data transmission rate can be improved, and the efficiency of accessing memory data by the processing unit can be further improved. Optionally, the compression / decompression engine may be disabled.

[0039] Optionally, the memory sharing control device further includes a storage unit, and the storage unit includes software code of at least one of a QoS engine, a prefetch engine, and a compression / decompression engine. The memory sharing control device may read the code in the storage unit to implement the corresponding function.

[0040] Optionally, at least one of the QoS engine, the prefetch engine, and the compression / decompression engine may be implemented by using the control logic of the memory sharing control device.

[0041] In some possible implementation manners, the first processing unit further has a local memory, and the local memory is used for the memory access of the first processing unit. Optionally, the first processing unit may preferentially access the local memory. The first processing unit has a higher speed of accessing the local memory, so that the speed of accessing the memory by the first processing unit can be further improved.

[0042] In some possible implementation manners, the multiple memories included in the memory pool are of different media types. For example, the memory pool may include at least one of the following memory media, namely, DRAM, phase change memory (PCM), storage class memory (SCM), static random access memory (SRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), NAND flash memory, spin-transfer torque random access memory (STT-RAM), or resistive random access memory (RRAM). The memory pool may further include a dual in-line memory module (DIMM) or a solid-state disk (SSD).

[0043] The different memory media can meet the memory resource requirements when different processing units process different services. For example, DRAM has characteristics of high read / write speed and volatility, and the memory of DRAM may be allocated to a processing unit that starts hot data access. PCM has the characteristic of non-volatility, and the memory of PCM may be allocated to a processing unit that accesses data that needs to be stored for a long time. In this way, while the memory resources are shared, the flexibility of memory access control can be improved.

[0044] For example, the memory pool includes a volatile DRAM storage medium and a non-volatile PCM storage medium. The DRAM and PCM within the memory pool may have a parallel architecture and may not have hierarchical levels. Alternatively, a non-parallel architecture where DRAM is used as a cache and PCM is used as main memory may be used. DRAM may be used as a first-level storage medium, and PCM may be used as a second-level storage medium. For an architecture where DRAM and PCM are parallel to each other, the control unit may store frequently accessed hot data in DRAM. In other words, the control unit may establish a correspondence relationship between the processing unit that starts the process of accessing frequently accessed hot data and the virtual memory device corresponding to the memory of DRAM. In this way, the read / write speed of memory data and the service life of the main memory system can be improved. The control unit may further establish a correspondence relationship between the processing unit that starts the process of accessing infrequently accessed cold data and the virtual memory device corresponding to the memory of PCM in order to store infrequently accessed cold data in PCM. In this way, the security of important data can be ensured based on the non-volatile feature of PCM. For an architecture where DRAM and PCM are not parallel to each other, based on the characteristics of high integration of PCM and low read / write latency of DRAM, the control unit may use PCM as main memory and DRAM as a cache to store various types of data. In this way, the memory access efficiency and performance can be further improved.

[0045] According to a second aspect, this application provides a system including at least two computer devices according to the first aspect, and the at least two computer devices according to the first aspect are connected to each other through a network.

[0046] The computer device of the system can access not only the memory pool on the computer device through the memory sharing control device to improve the memory utilization rate, but also the memory pool on other computer devices through the network. The range of the memory pool is expanded, so that the utilization rate of the memory resources can be further improved.

[0047] Optionally, the memory sharing control device in the computer device in the system may alternatively have the function of a network adapter and can send the access request of the processing unit in the system to other computer devices in the system through the network to access the memory of other computer devices.

[0048] Optionally, the computer device in the system may alternatively include a network adapter with a serial-parallel interface (for example, a Serdes interface). The memory sharing control device in the computer device may send the memory access request of the processing unit in the system to other computer devices in the system through the network by using the network adapter to access the memory of other computer devices.

[0049] Optionally, the computer device in the system may be connected through an Ethernet-based network or a unified bus (U-bus)-based network.

[0050] According to a third aspect, this application provides a memory sharing control device, and the memory sharing control device includes a control unit, a processor interface, and a memory interface.

[0051] The processor interface is configured to receive memory access requests sent by at least two processing units, and the processing units are processors, cores in the processor, or combinations of cores in the processor.

[0052] The control unit is configured to separately allocate the memory from the memory pool to at least two processing units, and at least one memory in the memory pool is accessible by different processing units during different periods.

[0053] The control unit is further configured to access the memory allocated to at least two processing units through a memory interface.

[0054] Through a memory sharing control device, different processing units can access at least one memory in the memory pool during different periods, so that the memory resource requirements of the processing units can be met and the utilization of memory resources can be improved.

[0055] Optionally, the fact that at least one memory in the memory pool is accessible by different processing units during different periods means that any two of the at least two processing units can separately access at least one memory in the memory pool during different periods. For example, the at least two processing units include a first processing unit and a second processing unit. In the first period, the first memory in the memory pool is accessed by the first processing unit, and the second processing unit cannot access the first memory. In the second period, the first memory in the memory pool is accessed by the second processing unit, and the first processing unit cannot access the first memory.

[0056] Optionally, the memory interface may be a double data rate (DDR) controller, or the memory interface may be a memory controller having a PCM control function.

[0057] Optionally, the memory sharing control device may separately allocate memory from the memory pool to at least two processing units based on the received control instructions transmitted by the operating system within the computer device. Specifically, a driver within the operating system may transmit, over a dedicated channel, control instructions used to allocate memory within the memory pool to at least two processing units to the memory sharing control device. The operating system is implemented by the CPU within the computer device by executing the relevant code. The CPU executing the operating system has a privilege mode, in which a driver within the operating system can transmit control instructions to the memory sharing control device over a dedicated or designated channel.

[0058] Optionally, the memory sharing control device may be implemented by an FPGA chip, an ASIC, or other similar chips.

[0059] In some possible implementation manners, the processor interface is further configured to receive, via a serial bus, a first memory access request transmitted in a serial signal format by a first processing unit within at least two processing units, and the first memory access request is used to access a first memory allocated to the first processing unit.

[0060] The serial bus has characteristics of high bandwidth and low latency. The first memory access request transmitted by the first processing unit within at least two processing units is received via the serial bus, thereby ensuring the efficiency of data transmission between the processing unit and the memory sharing control device.

[0061] Optionally, the serial bus is a memory semantic bus. The memory semantic bus includes, but is not limited to, buses based on QPI, PCIe, HCCS, or CXL protocol interconnections.

[0062] In some possible implementation manners, the processor interface is further configured to convert a first memory access request into a second memory access request in a parallel signal format and send the second memory access request to a control unit.

[0063] The control unit is further configured to access a first memory based on the second memory access request through a memory interface.

[0064] Optionally, the processor interface is an interface capable of realizing conversion between a parallel signal and a serial signal, for example, it may be a Serdes interface.

[0065] In some possible implementation manners, the control unit is further configured to establish a correspondence between a memory address of a first memory in a memory pool and a first processing unit in order to allocate the first memory from the memory pool to the first processing unit.

[0066] Optionally, the correspondence between the memory address of the first memory and the first processing unit is dynamically adjustable. For example, the correspondence between the memory address of the first memory and the first processing unit may be dynamically adjusted as needed.

[0067] Optionally, the memory address of the first memory may be a segment of consecutive physical memory addresses in the memory pool. The segment of consecutive physical memory addresses in the memory pool can simplify the management of the first memory. Obviously, as an alternative, the memory address of the first memory may be several segments of non - consecutive physical memory addresses in the memory pool.

[0068] Optionally, the memory address information of the first memory includes the start address of the first memory and the size of the first memory. The first processing unit has an identifier, and establishing the correspondence between the memory address of the first memory and the first processing unit may be to establish the correspondence between the unique identifier of the first processing unit and the memory address information of the first memory. In some possible implementation manners, the control unit is further configured to virtualize a plurality of virtual memory devices from a memory pool, and the physical memory corresponding to the first virtual memory device in the plurality of virtual memory devices is the first memory. It is further configured to allocate the first virtual memory device to the first processing unit.

[0069] Optionally, the virtual memory device corresponds to a segment of consecutive physical memory addresses in the memory pool. The virtual memory device corresponds to a segment of consecutive physical memory addresses in the memory pool, thereby simplifying the management of the virtual memory device. Obviously, the virtual memory device may alternatively correspond to some segments of discontinuous physical memory addresses in the memory pool.

[0070] Optionally, the first virtual memory device may be allocated to the first processing unit by establishing an access control table. For example, the access control table may include information such as the identifier of the first processing unit, the identifier of the first virtual memory device, and the start address and size of the memory corresponding to the first virtual memory device. The access control table may further include permission information for the first processing unit to access the first virtual memory device, attribute information of the accessed memory (including but not limited to information regarding whether the memory is a persistent memory), etc.

[0071] In some possible implementation manners, when a preset condition is satisfied, the control unit further cancels the correspondence relationship between the first virtual memory device and the first processing unit, and establishes a correspondence relationship between the first virtual memory device and a second processing unit among at least two processing units.

[0072] Optionally, the correspondence relationship between the virtual memory device and the processing unit may be dynamically adjusted based on the memory resource requirements of at least two processing units.

[0073] The correspondence relationship between the virtual memory device and the processing unit is dynamically adjusted, so that the memory resource requirements of different processing units in different service scenarios can be flexibly adapted, and the utilization rate of memory resources can be improved.

[0074] Optionally, the control unit When a preset condition is satisfied, cancels the correspondence relationship between the first memory and the first virtual memory device, establishes a correspondence relationship between the first memory and a second virtual memory device among a plurality of virtual memory devices, and further configures the second virtual memory device to be assigned to a second processing unit among at least two processing units. In this case, it is not necessary to change the correspondence relationship between the virtual memory device and the physical memory address in the memory pool, and only the correspondence relationship between the virtual memory device and different processing units needs to be changed, so that different processing units can access the same physical memory during different periods. In some possible implementation manners, the memory sharing control device further includes a cache unit.

[0075] The cache unit is configured to cache data read by any one of at least two processing units from the memory pool, or to cache data evicted by any one of at least two processing units.

[0076] The efficiency of accessing memory data by the processing unit can be further improved by using the cache unit.

[0077] Optionally, the cache unit may include a level 1 cache and a level 2 cache. The level 1 cache may be a small-capacity cache with a higher read / write speed than that of the level 2 cache. For example, the level 1 cache may be a 100MB nanosecond-level cache. The level 2 cache may be a large-capacity cache with a lower read / write speed than that of the level 1 cache. For example, the level 2 cache may be a 1GB DRAM. When the level 1 cache and the level 2 cache are used, thereby the data access speed of the processor can be improved by using the cache, while the cache space can be increased, the range in which the processor can quickly access the memory by using the cache is extended, and the memory access rate of the processor resource pool is generally further improved.

[0078] In some possible implementation manners, the memory sharing control device further includes a prefetch engine, and the prefetch engine is configured to prefetch data that needs to be read by any one of at least two processing units from the memory pool and cache the data in the cache unit.

[0079] Optionally, the prefetch engine may implement intelligent data prediction by using a specified algorithm or an AI algorithm to further improve the efficiency of accessing memory data by the processing unit.

[0080] In some possible implementation manners, the memory sharing control device further includes a quality of service (QoS) engine.

[0081] The QoS engine is configured to realize optimized storage of data that needs to be cached by any one of at least two processing units in the cache unit. By using the QoS engine, different capabilities of caching memory data accessed by different processing units can be realized in the cache unit 304. For example, a memory access request initiated by a processing unit with high priority has exclusive cache space within the cache unit 304. In this way, it can be ensured that the data accessed by the processing unit can be cached in a timely manner, thereby ensuring the service processing quality of this type of processing unit.

[0082] In some possible implementation manners, the memory sharing control device further includes a compression / decompression engine.

[0083] The compression / decompression engine is configured to compress or decompress data related to memory access.

[0084] Optionally, the function of the compression / decompression engine may be disabled.

[0085] Optionally, the compression / decompression engine may compress the data written to the memory by the processing unit at a granularity of 4KB per page by using a compression ratio algorithm, and then write the compressed data to the memory, or when the processing unit reads the compressed data in the memory, it may decompress the data to be read and then send the decompressed data to the processor. In this way, the data transmission rate can be improved, and the efficiency of accessing memory data by the processing unit can be further improved. Optionally, the compression / decompression engine may be disabled.

[0086] Optionally, the memory sharing control device may further include a storage unit, and the storage unit includes software code of at least one of a QoS engine, a prefetch engine, and a compression / decompression engine. The memory sharing control device may read the code in the storage unit to implement the corresponding function.

[0087] Optionally, at least one of the QoS engine, the prefetch engine, and the compression / decompression engine may be implemented by using the control logic of the memory sharing control device.

[0088] According to a fourth aspect, this application provides a memory sharing control method, which is applied to a computer device. The computer device includes at least two processing units, a memory sharing control device, and a memory pool. The memory pool includes one or more memories. The method includes the following.

[0089] The memory sharing control device receives a first memory access request sent by a first processing unit in at least two processing units. The processing unit is a processor, a core in the processor, or a combination of cores in the processor.

[0090] The memory sharing control device allocates a first memory from the memory pool to the first processing unit. The first memory can be accessed by a second processing unit in at least two processing units during other periods.

[0091] The first processing unit accesses the first memory through the memory sharing control device.

[0092] According to this method, different processing units access at least one memory in the memory pool during different periods, so that the memory resource requirements of the processing units can be met, and the utilization rate of memory resources is improved.

[0093] In a possible implementation manner, the method further includes the following.

[0094] The memory sharing control device receives, via a serial bus, a first memory access request transmitted in the form of a serial signal by a first processing unit in at least two processing units, and the first memory access request is used to access a first memory assigned to the first processing unit.

[0095] In a possible implementation manner, the method further includes the following.

[0096] The memory sharing control device converts the first memory access request into a second memory access request in the form of a parallel signal and accesses the first memory based on the second memory access request.

[0097] In a possible implementation manner, the method further includes the following.

[0098] The memory sharing control device establishes a correspondence relationship between the memory address of the first memory in the memory pool and the first processing unit in at least two processing units.

[0099] In a possible implementation manner, the method further includes the following.

[0100] The memory sharing control device virtualizes a plurality of virtual memory devices from the memory pool, and the physical memory corresponding to the first virtual memory device in the plurality of virtual memory devices is the first memory.

[0101] The memory sharing control device further assigns the first virtual memory device to the first processing unit.

[0102] In a possible implementation manner, the method further includes the following.

[0103] When a preset condition is satisfied, the memory sharing control device cancels the correspondence relationship between the first virtual memory device and the first processing unit, and establishes a correspondence relationship between the first virtual memory device and a second processing unit among at least two processing units.

[0104] In a possible implementation manner, the method further includes the following.

[0105] The memory sharing control device caches data read by any one of at least two processing units from a memory pool, or caches data evicted by any one of at least two processing units.

[0106] In a possible implementation manner, the method further includes the following.

[0107] The memory sharing control device prefetches data that needs to be read by any one of at least two processing units from a memory pool and caches the data.

[0108] In a possible implementation manner, the method further includes the following.

[0109] The memory sharing control device controls optimized storage of data that needs to be cached by any one of at least two processing units in a cache storage medium.

[0110] In a possible implementation manner, the method further includes compressing or decompressing data related to memory access.

[0111] According to a fifth aspect, an embodiment of this application further provides a chip, and the chip is configured to implement the functions implemented by the memory sharing control device according to the third aspect.

[0112] According to a sixth aspect, an embodiment of this application further provides a computer-readable storage medium including program code. The program code includes instructions used to execute some or all of the steps in any of the methods provided in the fourth aspect.

[0113] According to a seventh aspect, an embodiment of this application further provides a computer program product. When the computer program product operates on a computer, any of the methods according to the fourth aspect can be executed.

[0114] It can be understood that any of the memory sharing control devices, computer-readable storage media, or computer program products provided above are configured to execute the corresponding methods provided above. Therefore, for the advantageous effects that can be achieved by the memory sharing control device, computer-readable storage medium, or computer program product, refer to the advantageous effects in the corresponding methods. Details will not be described again here.

Brief Description of the Drawings

[0115] The following briefly describes the accompanying drawings necessary for explaining the embodiments. The accompanying drawings in the following description merely show some embodiments of the present invention, and it is obvious that those skilled in the art can derive other drawings from these accompanying drawings without creative efforts.

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7A

Figure 7B

Figure 7C

Figure 7D

Figure 7E

Figure 7F

Figure 8A-1

Figure 8A-2

Figure 8B-1

Figure 8B-2

Figure 9A

Figure 9B

Figure 9C

Figure 10

Figure 11

Figure 12

Mode for Carrying Out the Invention

[0116] Embodiments of the present invention will be described below with reference to the accompanying drawings.

[0117] In the specification, claims and accompanying drawings of this application, terms such as "first", "second", etc. are intended to distinguish between similar objects and do not necessarily indicate a specific order or sequence. Data thus named can be exchanged in appropriate circumstances, whereby it should be understood that the embodiments described herein can be realized in an order other than the order illustrated or described herein. Furthermore, the terms "first" and "second" are merely intended for the purpose of explanation and should not be understood as an indication of relative importance or an implication thereof, or as an implicit indication of the number of technical features shown. Therefore, features limited by "first" or "second" may explicitly or implicitly include one or more features.

[0118] In the description and claims of this application, the terms "include", "have" and any other variations thereof are meant to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to these explicitly listed steps or modules, and may also include other steps or modules not explicitly listed or specific to such a process, method, product or device. The names or numbers of the steps in this application do not mean that the steps in the method procedure need to be executed in the time / logical order indicated by the name or number. On the condition that the same or similar technical effects can be achieved, the execution order of the steps in the named or numbered procedure can be changed based on the technical purpose to be achieved. The division of units in this application is a logical division, and other divisions are also possible in the actual implementation method. For example, multiple units may be combined or integrated into other systems, or some features may be ignored or not executed. Furthermore, the indicated mutual connection, direct connection or communication connection may be realized through some interfaces. The indirect connection or communication connection between units may be realized in an electronic or other similar form. This is not limited in this application. Additionally, the units or subunits described as separate components may or may not be physically separate, may or may not be physical units, or may be distributed among multiple circuit units. Some or all of the units may be selected depending on the actual requirements to achieve the purpose of the solution of this application.

[0119] It should be understood that the terms used in the description and claims of this application in the description of various examples are merely intended to illustrate specific examples and are not intended to limit the examples. The singular terms "one" ("a" and "an") and "the" used in the description of various examples and the appended claims are also intended to include the plural unless clearly specified otherwise in the context.

[0120] The term "and / or" used in the description and claims of this application indicates any one or all possible combinations of one or more of the related listed items and is also to be understood as including. The term "and / or" describes the associative relationship between related objects and indicates that three relationships may exist. For example, A and / or B may indicate the following three cases, namely, only A exists, both A and B exist, and only B exists. Further, the character " / " in this application usually indicates an "or" relationship between related objects.

[0121] It should be understood that determining B based on A does not mean that B is determined only based on A. B may alternatively be determined based on A and / or other information.

[0122] The term "include" (also referred to as "includes", "including", "comprises" and / or "comprising") used in this specification specifies the presence of the described features, integers, steps, operations, elements and / or components, but it should be further understood that it does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0123] It should be further understood that the term "if" may be interpreted to mean "when" (either "when" or "upon"), "in response to a determination", or "in response to a detection". Similarly, depending on the context, the phrases "when determined" or "when (the stated condition or event) is detected" may be interpreted to mean "when determined", "in response to a determination", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".

[0124] It should be understood that the "one embodiment", "embodiment", and "possible implementation" referred to throughout this specification mean that a particular feature, structure, or characteristic related to the embodiment or implementation is included in at least one embodiment of this application. Therefore, the phrases "in one embodiment", "in an embodiment", or "in a possible implementation" that appear throughout this specification do not necessarily mean the same embodiment. Furthermore, these specified features, structures, or characteristics may be combined in one or more embodiments in any suitable manner.

[0125] Preferably, for ease of understanding, some terms and related technologies in this application are explained and described.

[0126] The memory controller is an important component for controlling the internal memory of a computer system and realizing data exchange between the memory and the processor, and it is a bridge for communication between the central processing unit and the memory. The memory controller is mainly configured to execute read and write operations on the memory, and can be roughly classified into a conventional memory controller and an integrated memory controller. In a conventional computer system, the memory controller is located in the north bridge chip of the main board chipset. In this structure, any data transmission between the CPU and the memory passes through the path "CPU - north bridge - memory - north bridge - CPU". When the CPU reads / writes data from / to the memory, multi-level data transmission is required. Therefore, a long latency is caused. The integrated memory controller is located inside the CPU, and any data transmission between the CPU and the memory needs to pass through the path "CPU - memory - CPU". Compared with the conventional memory controller, the latency of data transmission is significantly reduced.

[0127] DRAM is a widely used memory medium. Different from sequential access to disk media, DRAM enables the central processing unit to directly and randomly access any byte of the disk media. DRAM has a simple memory structure, and each memory structure mainly includes a capacitor and a transistor. When the capacitor is charged, this indicates that the data "1" is stored. The state after the capacitor is completely discharged represents the data "0".

[0128] PCM is a non-volatile memory that stores information based on phase change memory materials. Each memory unit in PCM includes a phase change material (such as sulfide glass) and two electrodes. The phase change material can be converted between a crystalline state and an amorphous state by changing the voltage of the electrode and the power-on time. When in the crystalline state, the medium has a low resistance. When in the amorphous state, the medium has a high resistance. Therefore, data may be stored by changing the state of the phase change material. The most typical characteristic of PCM is non-volatility.

[0129] A serializer / deserializer (Serdes) converts parallel data into serial data at the transmitting end and then transmits the serial data through a transmission line to the receiving end, or converts serial data into parallel data at the receiving end, thereby reducing the number of transmission lines and the system cost. Serdes is a time division multiplexing (TDM) and point-to-point communication technology. Specifically, multiple low-speed parallel signals (i.e., parallel data) at the transmitting end are converted into a high-speed serial signal (i.e., serial data), and the high-speed serial signal is reconverted into a low-speed parallel signal at the receiving end through a transmission medium. Serdes uses differential signals for transmission, thereby enabling the interference and noise captured by two differential transmission lines to cancel each other out. This improves the transmission speed and the signal transmission quality. The parallel interface technology represents the parallel transmission of multi-bit data, and a synchronous clock is transmitted to divide the data bytes. Therefore, this method is simple and easy to implement, but because there are a large number of signal lines, it is usually used for short-distance data transmission. The serial interface technology is widely applied to long-distance data communication to transmit byte data bit by bit.

[0130] With the continuous improvement of the technical level of integrated circuits, especially the continuous improvement of the architecture design of processors, the performance of processors has been gradually improved. Compared with processors, the improvement of memory performance is much slower. With the cumulative increase of the gap, the result of the unbalanced cumulative increase is that the memory access speed lags significantly behind the computing speed of the processor, and the bottleneck formed by the memory makes it difficult to exert the advantages of high-performance processors. For example, the memory access speed is greatly restricted for increasing high performance computing (HPC).

[0131] Furthermore, multi-core processors are gradually replacing single-core processors, and the parallel execution of multiple cores within the processor significantly increases the number of accesses to memory (also called off-chip memory or main memory). This also results in a corresponding increase in the bandwidth requirements between the processor and the memory.

[0132] The access speed and bandwidth between the processor and the memory are usually improved by sharing memory resources.

[0133] Depending on whether there are differences in accesses from the processor to the memory, an architecture in which multiple processors share memory may be divided into a centralized memory sharing system and a distributed memory sharing system. The centralized memory sharing system has the characteristics of a small number of processors and a single interconnection method, and the memory is connected to all processors via a cross-switch or a shared bus. FIG. 1A is a typical architecture of a centralized memory sharing system. Since all memory accesses are equal or symmetric for all processors, this type of architecture is also called a unified memory architecture (UMA) or a symmetric multiprocessing (SMP) architecture.

[0134] The centralized memory sharing system has a single memory system and thus faces the problem that after the number of processors reaches a specified scale, the required access memory bandwidth cannot be provided. This becomes a bottleneck limiting performance. The distributed memory sharing system effectively solves this problem. FIG. 1B is a schematic diagram of the structure of the distributed memory sharing system. As shown in FIG. 1B, in the system, the memory is globally shared, uniformly addressed, and distributed to the processors. The address space of the memory is divided into several parts, each of which is managed by a processor. For example, when processor 1 needs to access the memory address space managed by a processor, processor 1 does not need to cross the processor or the interconnect bus. When processor 1 accesses the memory space managed by another processor, for example, when it needs to access the memory address space managed by processor N, processor 1 needs to cross the interconnect bus. The distributed memory sharing system is also called a non-uniform memory access (NUMA) system.

[0135] In a NUMA system, the address space of the shared memory is managed by each processor. Since there is no unified memory management mechanism, when a processor needs to use the memory space managed by another processor, the memory resources are not flexible enough to be shared, and there is a low utilization rate of the memory resources. Furthermore, when a processor accesses a memory address space not managed by a processor, since the processor crosses the interconnect bus, a relatively long latency is usually caused.

[0136] Embodiments of this application provide a memory sharing control device, a chip, a computer device, a system, and a method, and provide a new memory access architecture in which a bridge for access between a processor and a shared memory pool (which may also be simply called a memory pool in the embodiments) is established through the memory sharing control device to improve the utilization of memory resources.

[0137] Figure 2A is a schematic diagram of the structure of the memory sharing control device 200 according to an embodiment of this application. As shown in Figure 2A, the memory sharing control device 200 includes a control unit 201, a processor interface 202, and a memory interface 203.

[0138] The memory sharing control device 200 may be a chip located between a processor (CPU or a core within the CPU) and a memory (also called the main memory) in a computer device, for example, an FPGA chip.

[0139] Figure 2B is a schematic diagram of the connection relationship between the memory sharing control device 200 and each of the processor and the memory. As shown in Figure 2B, the processor 210 is connected to the memory sharing control device 200 through the processor interface 202, and the memory 220 is connected to the memory sharing control device 200 through the memory interface 203. The processor 210 may be a CPU or a CPU including a plurality of cores. The memory 220 includes, but is not limited to, DRAM, PCM, flash memory, SCM, SRAM, PROM, EPROM, STT-RAM, or RRAM. SCM is a composite storage technology that combines the characteristics of conventional storage devices and memories. Storage class memory provides a higher read / write speed than a hard disk, but can provide a lower operating speed and a lower cost than DRAM. In this embodiment of this application, the memory 220 may further include a DIMM or an SSD.

[0140] The processor interface 202 is an interface through which the memory sharing control device 200 is connected to the processor 210. The interface can receive a serial signal transmitted by the processor and convert the serial signal into a parallel signal. Based on the processor interface 202, the memory sharing control device 200 may be connected to the processor 210 via a serial bus. The serial bus has characteristics of high bandwidth and low latency to ensure the efficiency of data transmission between the processor 210 and the memory sharing control device 200. For example, the processor interface 202 may be a low-latency-based Serdes interface. The Serdes interface functioning as the processor interface 202 is connected to the processor via a serial bus to realize the conversion between the serial signal and the parallel signal based on serial-parallel logic. The serial bus may be a memory semantic bus. The memory semantic bus includes, but is not limited to, a bus based on QPI, PCIe, HCCS, or CXL protocol interconnect.

[0141] In a specific implementation manner, the processor 210 may be connected to a serial bus through a Serdes interface and is connected to the processor interface 202 (for example, a Serdes interface) of the memory sharing control device 200 via the serial bus. The memory access request initiated by the processor 210 is a memory access request in parallel signal form. The memory access request in parallel signal form is converted into a memory access request in serial signal form through the Serdes interface in the processor 210, and the memory access request in serial signal form is transmitted via the serial bus. After receiving the memory access request in serial signal form from the processor 210 via the serial bus, the processor interface 202 converts the memory access request in serial signal form into a memory access request in parallel signal form and transmits the memory access request obtained through the conversion to the control unit 301. The control unit 301 may access the corresponding memory based on the memory access request in parallel signal form. For example, the corresponding memory may be accessed in parallel mode. In this embodiment of this application, the parallel signal may be a signal that transmits multiple bits at a time, and the serial signal may be a signal that transmits one bit at a time.

[0142] Similarly, when the memory sharing control device 200 returns a response message of a memory access request to the processor 210, the response message in parallel signal form is converted into a response message in serial signal form through the processor interface 202 (for example, a Serdes interface), and the response message in serial signal form is transmitted to the processor 210 via the serial bus. After receiving the response message in serial signal form, the processor 210 converts the response message in serial signal form into a parallel signal and then executes subsequent processing.

[0143] The memory sharing control device 200 may access the corresponding memory in the memory 220 through the memory interface 203 used as a memory controller. For example, when the memory 220 is a shared memory pool including DRAM, the memory interface 203 is a DDR controller having a DRAM control function and is configured to realize the interface control of the DRAM storage medium. When the memory 220 is a shared memory pool including PCM, the memory interface 203 is a memory controller having a PCM control function and is configured to realize the interface control of the PCM storage medium.

[0144] One processor 210 shown in FIG. 2B is merely an example, and it should be noted that the processor connected to the memory sharing control device 200 may alternatively be a multi-core processor or a processor resource pool. The processor resource pool includes at least two processing units, and each processing unit may be a processor, a core in the processor, or a combination of cores in the processor. The processing units in the processor resource pool may be a combination of different cores in the same processor or a combination of different cores in different processors. When the processor executes different tasks, multiple cores need to execute computing tasks in parallel, or cores in different processors need to execute computing tasks in parallel. When these cores execute computing tasks in parallel, the combination of these cores may be used as processing units for accessing the same memory in the shared memory pool.

[0145] One memory 220 shown in FIG. 2B is merely an example, and the memory 220 connected to the memory sharing control device 200 may alternatively be a shared memory pool including a plurality of memories. At least one memory in the shared memory pool is accessible by different processing units at different times. The memories in the shared memory pool include, but are not limited to, DRAM, PCM, flash memory, STT-RAM, or RRAM. Similarly, the memories in the shared memory pool may be the memories of one computer device, or may be the memories of different computer devices. The computer device may be a device such as a computer (desktop computer or portable computer) or a server that requires a processor to access the memory, or may include a terminal device such as a mobile phone terminal. It can be understood that the specific form of the device is not limited in this embodiment of this application.

[0146] The control unit 201 is configured to control memory access based on a memory access request, and includes, but is not limited to, dividing the memory resources in the shared memory pool into a plurality of independent memory resources and separately allocating the plurality of independent memory resources to the processing units in the processor resource pool (for example, allocating on demand). The independent memory resources obtained through the division by the control unit 201 may be a memory storage space corresponding to a segment of the physical address in the shared memory pool. The physical addresses of the memory resources may be continuous or discontinuous. For example, the memory sharing control device 200 may virtualize a plurality of virtual memory devices based on the shared memory pool, and each virtual memory device corresponds to or manages several memory resources. The control unit 201 allocates each of the plurality of independent memory resources obtained through the division in the shared memory pool to the processing units in the processor resource pool by establishing a correspondence between different virtual memory devices and the processing units.

[0147] However, the correspondence between the processing unit and the memory resources is not fixed. When certain conditions are met, the correspondence may be adjusted. That is, the correspondence between the processing unit and the memory resources may be dynamically adjusted. The control unit 201 adjusting the correspondence between the processing unit and the memory resources may include receiving a control instruction sent by a driver in the operating system and adjusting the correspondence based on the control instruction. The control instruction includes information regarding deletion, modification, or addition of the correspondence.

[0148] For example, a computer device 20 (not shown in the drawings) includes a processor 210, a memory sharing device 200, and a memory 220 shown in FIG. 2B. The processor 210 executes an operating system that the computer device 20 needs to execute in order to control the computer device 20. Assume that the computer device 20 is a server that provides cloud services and the processor 210 has 8 cores. Core A provides cloud services to user A, and core B provides cloud services to user B. Based on the service requirements of user A and user B, the operating system of the computer device separately allocates memory resource A in the memory 220 to core A as a memory access resource, and allocates memory resource B in the memory 220 to core B as a memory access resource. The operating system may send control instructions to the memory sharing control device 200 to establish the correspondence between core A and memory resource A and the correspondence between core B and memory resource B. The memory sharing control device 200 establishes the correspondence between core A and memory resource A and the correspondence between core B and memory resource B based on the control instructions of the operating system. In this way, when core A starts a memory access request, the memory sharing control device 200 may determine the memory resource (i.e., memory resource A) accessible by core A based on the information carried in the access request, thereby enabling core A to access memory resource A. When user A needs to rest for reasons such as service requirements or time zones, does not need to use cloud services, and the requirement for memory resources is reduced, and when user B needs to use more cloud services and requires more memory resources for reasons such as service requirements or time zones, the operating system of the computer device 20 may cancel the correspondence between core A and memory resource A based on the change in the service requirements of user A and user B, and send control instructions to allocate memory resource A to core B for use.Specifically, the operating system may also be a driver within the operating system, which deletes the correspondence between core A and memory resource A and sends a control command to establish the correspondence between core B and memory resource A. Based on the control command sent by the driver within the operating system, the memory sharing control device 200 reconstructs the correspondence between core B and memory resource A, deletes the correspondence between core A and memory resource A, and establishes the correspondence between core B and memory resource A. In this way, memory resource A can be used as the memory of core A and core B during different periods, so that the requirements of different cores for different services can be met, and the utilization rate of the memory resource can be improved.

[0149] The driver within the operating system may send a control command to the memory sharing control device 200 on a dedicated or specified channel. Specifically, when the processor executing the operating system is in the privilege mode, the driver within the operating system can send a control command to the memory sharing control device 200 on a dedicated or specified channel. In this way, the driver within the operating system may send a control command to delete, change or add the correspondence on the dedicated channel.

[0150] The memory sharing control device 200 may be connected to the processor 210 through an interface that supports serial - parallel (for example, a Serdes interface). The processor 210 can communicate with the memory sharing control device 200 via a serial bus. Based on the characteristics of the high bandwidth and low latency of the serial bus, even when the communication distance between the processor 210 and the memory sharing control device 200 is relatively large, the access rate for the processor 210 to access the shared memory pool can also be ensured.

[0151] Furthermore, the control unit 201 may be further configured to implement data buffering control, data compression control, data priority control, etc. Therefore, the efficiency and quality of accessing the memory by the processor are further improved.

[0152] Hereinafter, an example in which the FPGA is used as a chip for implementing the memory sharing control device 200 will be used to describe an example of the implementation method of the memory sharing control device 200 provided in this embodiment of this application.

[0153] As a programmable logic device, the FPGA may be classified into three types according to different principles of programmability, namely, a static random access memory (SRAM)-based SRAM type FPGA, an anti-fuse type FPGA, and a flash type FPGA. Due to the erasability and volatility of SRAM, the SRAM type FPGA can be programmed repeatedly, but the configuration data is lost due to power failure. The anti-fuse type FPGA can be programmed only once. After programming, the circuit function is fixed and cannot be modified again. Therefore, the circuit function does not change even when power is not supplied.

[0154] Hereinafter, the SRAM type FPGA will be used as an example for explaining the internal structure of the FPGA. FIG. 3 is a schematic diagram of the internal structure of the SRAM type FPGA. As shown in FIG. 3, the FAGA internally includes at least the following parts.

[0155] The configurable logic block (CLB) mainly includes programmable resources such as a lookup table (LUT), a multiplexer, a carry chain, and a D flip-flop inside, and is configured to implement different logic functions, and is the core of the entire FPGA chip.

[0156] The programmable input / output block (IOB) provides an interface between the FPGA and the external circuit, and provides appropriate driving for input / output signals to achieve alignment when the internal and external electrical characteristics of the FPGA are different. Electronic design automation (EDA) software is configured to configure different electrical standards and physical information as needed, for example, to adjust the value of the drive current and change the resistance of pull-up resistors and pull-down resistors. Usually, several IOBs are grouped into a bank. Different series of FPGA chips have different numbers of IOBs included in each group.

[0157] The block random access memory (BRAM) is configured to store data with a large amount of data. To meet different data read / write requirements, the BRAM may be configured as a common memory structure such as a single-port RAM, a dual-port RAM, a content addressable memory (CAM), and a first in first out (FIFO) cache queue, and the memory bit width and depth can be changed based on the design requirements. The BRAM can expand the application range of the FPGA and improve the flexibility of the FPGA.

[0158] The switch matrix (SM) is an important part of the interconnect resource (IR) inside the FPGA. It is mainly distributed at the left end of each resource module. The switch matrices at the left end of different modules are very similar but also different. It is configured to connect module resources. The other part of the interconnect resource inside the FPGA is the wire segment. The wire segment and the SM are used together to connect the overall chip resources.

[0159] FIG. 3 shows only some of the main components related to the implementation method of the memory sharing control device 200 in this embodiment of this application in the FPGA chip. In a specific implementation method, in addition to the components shown in FIG. 3, the FPGA may further include other components or embedded functional units. For example, it may further include a digital signal processor (DSP), a phase locked loop (PLL), or a multiplier (MUL).

[0160] The control unit 201 in FIG. 2A or FIG. 2B may be implemented by using the CLB in FIG. 3. Specifically, the shared memory pool connected to the memory sharing control device 200 is controlled by using the CLB. For example, the memory resources in the shared memory pool are divided into multiple blocks, and one or more memory resources are allocated to one processing unit. Alternatively, a plurality of virtual memory devices are virtualized based on the memory resources in the shared memory pool. Each virtual memory device corresponds to the physical address space in a segment of the shared memory pool, and one or more virtual memory devices are allocated to one processing unit, thereby establishing a correspondence table between the allocated virtual memory device and the corresponding processing unit, etc.

[0161] The processor interface 202 in FIG. 2A or FIG. 2B may be implemented by using the IOB in FIG. 3B. Specifically, an interface having a serial - parallel function may be implemented by using the IOB. For example, a Serdes interface is implemented by using the IOB.

[0162] Figure 4 is a schematic diagram of the specific structure of the Serdes interface. As shown in Figure 4, the Serdes interface mainly includes a transmission channel and a reception channel. In the transmission channel, the encoder encodes the input parallel data, then the parallel-serial module converts the encoded input parallel data into a serial signal, and then the transmitter (Tx) drives to output the serial data. In the reception channel, the receiver and the clock recovery circuit recover the sampling clock and data, then the serial-parallel module finds the byte boundary and converts the byte boundary into parallel data, and finally the decoder completes the recovery of the original parallel data.

[0163] The encoder and the decoder complete the functions of encoding and decoding the data so as to ensure the DC balance of the serial data stream and as many data jumps as possible. For example, 8b / 10b and irregular Scrambling / Descrambling encoding / decoding solutions may be used. The parallel-serial module and the serial-parallel module are configured to complete the conversion of data between the parallel format and the serial format. The clock generation circuit generates a conversion clock for the parallel-serial circuit, which is usually realized by a phase-locked loop. The clock generation circuit and the clock recovery circuit provide a conversion control signal for the serial-parallel circuit, which is usually realized by a phase-locked loop, but alternatively, it may be realized by a phase interpolator or the like.

[0164] The above only describes an example of the implementation method of the Serdes interface. That is, the Serdes interface shown in Figure 4 may be realized based on the IOB in Figure 3. Obviously, the function of the Serdes interface may alternatively be realized based on other hardware components, for example, other dedicated hardware components in the FPGA. The specific implementation form of the Serdes interface is not limited in this embodiment of this application.

[0165] The memory interface 203 in FIG. 2A or FIG. 2B may be implemented based on an IOB or other dedicated circuits. For example, when the memory interface 203 is implemented by using a DDR controller, the logical structure of the DDR controller may be as shown in FIG. 5.

[0166] FIG. 5 is a schematic diagram of the internal structure of the memory controller 500. Referring to FIG. 5, the memory controller 500 includes a receiving module 501 configured to record information regarding an access request, where the access request mainly includes a request type and a request address. From these two pieces of information, it can be known that a specific operation should be executed on a specific memory address based on the access request. The information regarding the access request recorded by the receiving module 501 may include the request type and the request address. Further, the information recorded by the receiving module 501 may further include some auxiliary information for estimating system performance, such as the arrival time and completion time of the access request.

[0167] The control module 502 is configured to control the initialization of the memory, power-off, etc. Further, the control module 502 may further perform operations such as controlling the depth of the memory queue used to control memory access, determining whether the memory queue is empty or full, determining whether a memory request is completed, determining the arbitration resolution policy to be used, and determining the scheduling method to be used.

[0168] The address mapping module 503 is configured to implement the conversion between the address of the access request and the address identifiable by the memory. For example, the memory address of a DDR4 memory system includes six parts, namely, Channel, Rank, Bankgroup, Bank, Row, and Column. Different address mapping methods have different access efficiencies.

[0169] The refresh module 504 is configured to implement a scheduled refresh for the memory. DRAM includes many repeating cells, and each cell includes a transistor (Mosfet) and a capacitor. The capacitor is configured to accumulate charge and determine whether the logical state of the DRAM unit is 1 or 0. However, since the capacitor is prone to leakage, the charge is sometimes lost, and as a result, the data is lost. Therefore, the refresh module 504 needs to perform a scheduled refresh.

[0170] The scheduling module 505 is configured to schedule access requests to different queues separately based on the access requests and request types sent by the address mapping module 503. For example, the scheduling module schedules access requests to a queue with a high priority, and to complete one scheduling, it may select the request with the highest priority from the queue with the highest priority according to a preset scheduling policy. Here, the queue is a memory access control queue, and the scheduling policy may be determined based on the time series of the arrival of the requests, where an earlier arrival time indicates a higher priority, or the scheduling policy may be determined based on the request that is prepared first.

[0171] It should be noted that FIG. 5 only shows some components or functional modules of the memory controller 500. In a specific implementation, the memory controller 500 may further include other components or functional modules. For example, the memory controller 500 may further include a Villa engine used for multi-threaded computing, a direct memory access (DMA) module for direct memory access, etc. Details will not be described one by one.

[0172] The above describes an implementation method of the memory sharing control device 200 by using an FPGA as an example. Among specific implementation methods, the memory sharing control device 200 may alternatively be implemented by using other chips or other devices that can realize similar chip functions. For example, the memory sharing control device 200 may alternatively be implemented by using an ASIC. The circuit function of the ASIC is defined at the beginning of the design, and the ASIC is characterized by high chip integration, ease of realizing a large number of tape-outs, low cost of a single tape-out, small size, etc. The specific hardware implementation method of the memory sharing control device 200 is not limited in this embodiment of this application.

[0173] In this embodiment of this application, the processor connected to the memory sharing control device 200 may be any processor that realizes processor functions. FIG. 6 is a schematic diagram of the structure of the processor 210 according to an embodiment of this application. As shown in FIG. 6, the processor 210 includes a kernel 601, a memory 602, a peripheral interface 603, etc. The kernel 601 may include at least one core, and the processor 210It is configured to realize the functions. In FIG. 6, two cores (core 1 and core 2) are used as examples for explanation. However, the number of cores in the processor 600 is not limited. The processor 600 may further include 4, 8, or 16 cores. The memory 602 includes a cache or SRAM and is configured to cache the read / write data of core 1 or core 2. The peripheral interface 603 includes a Serdes interface 6031, a memory controller 6032, an input / output interface, a power supply, a clock, etc. The Serdes interface 6031 is an interface for connecting the processor 210 and a serial bus. After the memory access request in parallel signal form started by the processor 210 is converted into a serial signal through the Serdes interface 6031, the serial signal is transmitted to the memory sharing control device 200 via the serial bus. The memory controller 6032 may be a memory controller having the same functions as those of the memory controller shown in FIG. 5. When the processor 210 has a local memory controlled by the processor 210, the processor 210 may realize access control to the local memory via the memory controller 6032 .

[0174] It can be understood that FIG. 6 is merely an example of a schematic diagram of the structure of the implementation manner of the processor. The specific structure or form of the processor connected to the memory sharing control device 200 is not limited in this embodiment of this application as long as any processor capable of realizing a specific computing or control function falls within the scope disclosed in this embodiment of this application.

[0175] Next, the specific implementation manner of the memory sharing control device provided in this embodiment of this application will be further described.

[0176] FIG. 7A is a schematic diagram of the structure of the memory sharing control device 300 according to an embodiment of this application. As shown in FIG. 7A, the memory sharing control device 300 includes a control unit 301, a processor interface 302, and a memory interface 303. For the specific implementation manner of the memory sharing control device 300 shown in FIG. 7, refer to the implementation manner of the memory sharing control device 200 in FIG. 2A or FIG. 2B, or refer to the implementation manner of the FPGA shown in FIG. 3. Specifically, the control unit 301 may be implemented by referring to the implementation manner of the control unit 201 in FIG. 2A or FIG. 2B, or may be implemented by using the CLB shown in FIG. 3. The processor interface 302 may be implemented by referring to the Serdes interface shown in FIG. 4, and the memory interface 303 may be implemented by referring to the memory controller shown in FIG. 5. Details will not be described again.

[0177] Specifically, the control unit 301 in FIG. 7A may implement the following functions through its configuration.

[0178] 1. Virtualize a plurality of virtual memory devices based on the memory resources connected to the memory sharing control device 300.

[0179] The memory resources connected to the memory sharing control device 300 form a shared memory pool. The control unit 301 executes unified address assignment for the memory resources in the shared memory pool, and after the unified address assignment, the physical memory address space may be divided into several address segments, and each address segment corresponds to one virtual memory device. The size of the address space corresponding to the address segments obtained through the division may be the same or different. In other words, the sizes of the virtual memory devices may be the same or different.

[0180] The virtual memory device is not an actual existing device, but a segment of the memory address space that is within the shared memory pool and configured to be identified by the control unit 301. The segment of the address space is allocated to a processing unit (which may be a processor, a core within a processor, a combination of different cores within the same processor, or a combination of cores within different processors) for memory access (e.g., data reading / writing), and thus is called a virtual memory device. For example, each virtual memory device corresponds to a segment of a memory area having consecutive physical addresses. Optionally, one virtual memory device may alternatively correspond to a discontinuous physical address space.

[0181] The control unit 301 may assign one identifier to each virtual memory device to identify different virtual memory devices. FIG. 7A shows an example of two virtual memory devices, namely virtual memory device a and virtual memory device b. Virtual memory device a and virtual memory device b respectively correspond to different memory address spaces within the shared memory pool. 2. Allocate a plurality of virtualized virtual memory devices to the processing unit connected to the memory sharing control device 300.

[0182] The control unit 301 may allocate virtual memory devices to the processing unit. To avoid possible complex logic or possible traffic storms, when allocating virtual memory devices, the control unit 301 avoids allocating one virtual memory device to multiple processors or allocating one virtual memory device to multiple cores within one processor. However, for some services, when different cores within the same processor need to execute computing tasks in parallel, or when different cores within different processors need to execute computing tasks in parallel, the memory corresponding to the virtual memory device is allocated to a combination of cores through complex logic, so as to improve the service processing efficiency during parallel computing.

[0183] The method by which the control unit 301 allocates virtual memory devices may also be to establish a correspondence between the identifier of the virtual memory device and the identifier of the processing unit. For example, the control unit 301 establishes a correspondence between a virtual memory device and a different processing unit based on the number of processing units connected to the memory sharing control device 300. Optionally, the control unit 301 may alternatively establish a correspondence between the processing unit and the virtual memory device and then establish a correspondence between the virtual memory device and a different memory resource in order to establish a correspondence between the processing unit and a different memory resource.

[0184] 3. Record the correspondence between the virtual memory device and the allocated processing unit.

[0185] In a specific implementation manner, the control unit 301 may maintain an access control table (also called a mapping table) used to record the correspondence between the virtual memory device and the processing unit. The implementation manner of the access control table may be as shown in Table 1.

Table 1

[0186] In Table 1, Device_ID represents the identifier of the virtual memory device, Address represents the start address of the physical memory address managed or accessible by the virtual memory device, Size represents the size of the memory managed or accessible by the virtual memory device, Access Attribute represents the access method, specifically, a read operation or a write operation. Resource_ID represents the identifier of the processing unit.

[0187] In Table 1, Resource_ID usually corresponds to one processing unit. Since the processing unit may be a processor, a core within a processor, a combination of multiple cores within a processor, or a combination of multiple cores within different processors, the control unit 301 may further maintain a correspondence table between the Resource_ID and the combination of cores to determine information regarding the core or processor corresponding to each processing unit. For example, Table 2 shows an example of the correspondence relationship between Resource_ID and cores.

Table 2

[0188] In a computer device, cores within different processors have a unified identifier. Therefore, the core IDs in Table 2 can be used to distinguish different cores within different processors. It can be understood that Table 2 merely shows an example of the correspondence relationship between the Resource_ID of the processing unit and the corresponding core or the corresponding processor. The manner in which the memory sharing control device 300 determines the correspondence relationship between the Resource_ID and the corresponding core or the corresponding processor is not limited in this embodiment of this application.

[0189] In other implementation manners, when the memory connected to the memory sharing control device 300 includes DRAM and PCM, due to the non-persistent characteristics of the DRAM storage medium and the persistent characteristics of the PCM storage medium, the access control table maintained by the control unit 301 may further include whether each virtual memory device is a persistent virtual memory device or a non-persistent virtual memory device.

[0190] Table 3 shows another implementation manner of the access control table according to the embodiment of this application.

Table 3

[0191] In Table 3, the Persistent Attribute represents the persistent attributes of the virtual memory device. In other words, it represents whether the memory address space corresponding to the virtual memory device is persistent or non-persistent.

[0192] Optionally, the access control table maintained by the control unit 301 may further include other information for further memory access control. For example, the access control table may further include permission information for the processing unit to access the virtual memory device, and the permission information includes, but is not limited to, read-only access or write-only access.

[0193] When a memory access request sent by the processing unit is received, based on the correspondence between the virtual memory device recorded in the access control table and the processing unit, determine the virtual memory device corresponding to the processing unit that sent the memory access request, and access the corresponding memory based on the determined virtual memory device.

[0194] For example, a memory access request includes information such as the RESOURCE_ID of the processing unit, address information, and access attributes. The RESOURCE_ID is the ID of a combination of cores, the address information is the address information of the memory to be accessed, and the access attributes indicate whether the memory access request is a read request or a write request. The control unit 301 may query an access control table (e.g., Table 1) based on the RESOURCE_ID to determine at least one virtual memory device corresponding to the RESOURCE_ID. For example, the determined virtual memory device is the virtual memory device a shown in FIG. 7A, and from Table 2, it can be determined that all cores corresponding to the RESOURCE_ID can access the memory resources that can be managed and accessed by the virtual memory device a. Next, the control unit 301 refers to the address information and access attributes in the access request and controls the memory access request so as to implement access control of the memory within the memory address space that can be managed or accessed by the virtual memory device a. Optionally, when the access control table records permission information, the control unit 301 may further control the access of the corresponding processing unit to the memory based on the permission information recorded in the access control table.

[0195] The access control executed by the control unit 301 on the virtual memory device is the access control implemented on the memory for the physical address space of the memory resources corresponding to the virtual memory device.

[0196] 5. Dynamically adjust the correspondence between the virtual memory device and the processing unit.

[0197] The control unit 301 may dynamically adjust the virtual memory device by changing the correspondence between the processing unit and the virtual memory device in the access control table based on preset conditions (for example, different processing units have different requirements for memory resources). For example, the control unit 301 may delete the correspondence between the virtual memory device and the processing unit, in other words, release the memory resources corresponding to the virtual memory device, and the released memory resources may be allocated to other processing units for memory access. Specifically, this may be realized by referring to the method in which the control unit 201 dynamically adjusts the correspondence to delete, modify, or add the correspondence in FIG. 2B.

[0198] In an optional implementation method, the modulation of the correspondence between the processing unit and the memory resources in the shared memory pool may alternatively be realized by changing the memory resources corresponding to each virtual memory device. For example, when the service processed by the processing unit is in a suspended state and does not need to occupy too much memory, the memory resources managed by the virtual memory device corresponding to the processing unit may be allocated to the virtual memory device corresponding to other processing units, so that the same memory resources are accessed by different processing units during different periods.

[0199] For example, when the memory sharing control device 300 is realized by using the FPGA chip shown in FIG. 3, the function of the control unit 301 may be realized by configuring the CLB in FIG. 3.

[0200] The control unit 301 may virtualize a plurality of virtual memory devices, allocate the virtualized plurality of virtual memory devices to a processing unit connected to the memory sharing control device 300, and dynamically adjust the correspondence relationship between the virtual memory devices and the processing unit. It should be noted that this may be realized based on the received control instructions transmitted by a driver in the operating system on a dedicated channel. In other words, a driver in the operating system of the computer device where the memory sharing control device 300 is located transmits instructions for virtualizing a plurality of virtual memory devices, allocating the virtual memory devices to the processing unit, and dynamically adjusting the correspondence relationship between the virtual memory devices and the processing unit to the memory sharing control device 300 on a dedicated channel, and the control unit 301 realizes the corresponding functions based on the received control instructions.

[0201] The memory sharing control device 300 is connected to the processor via a serial bus through a serial - parallel interface (for example, a Serdes interface), thereby ensuring the speed at which the processor accesses the memory while enabling long - distance transmission between the memory sharing control device 300 and the processor. Therefore, the processor can quickly access the memory resources in the shared memory pool. Since the memory resources in the shared memory pool can be allocated to different processing units at different times for memory access, the utilization rate of the memory resources is improved.

[0202] For example, the control unit 301 in the memory sharing control device 300 dynamically adjusts the correspondence between the virtual memory device and the processing unit. When the processing unit requires more memory space, it adjusts an unoccupied virtual memory device or a virtual memory device that is allocated to another processing unit but is temporarily idle to the processing unit that requires more memory, that is, it can establish the correspondence between these idle virtual memory devices and the processing unit that requires more memory. In this way, the existing memory resources can be effectively utilized to meet the different service requirements of the processing units. This not only ensures the requirements of the processing unit for the memory space in different service scenarios, but also improves the utilization rate of the memory resources.

[0203] FIG. 7B is a schematic diagram of the structure of another memory sharing control device 300 according to an embodiment of this application. Based on FIG. 7A, the memory sharing control device 300 shown in FIG. 7B further includes a cache unit 304.

[0204] The cache unit 304 may be a random access memory (RAM) and is configured to cache data that needs to be accessed by the processing unit during memory access. For example, the data that needs to be read by the processing unit is read from the shared memory pool in advance and cached in the cache unit 304, so that the processing unit can quickly access the data and further improve the rate at which the processing unit reads the data. Alternatively, the cache unit 304 may cache the data evicted by the processing unit, such as Cacheline data evicted by the processing unit. The speed at which the processing unit accesses the memory data can be further improved by using the cache unit 304.

[0205] In an optional implementation method, the cache unit 304 may include a level 1 cache and a level 2 cache. As shown in FIG. 7C, the cache unit 304 in the memory sharing control device 300 further includes a level 1 cache 3041 and a level 2 cache 3042.

[0206] The level 1 cache 3041 may be a cache with a small capacity (for example, a capacity at the 100MB level), or may be an SRAM medium at the nanosecond level, and caches the Cacheline data evicted from the processing unit.

[0207] The level 2 cache 3042 may be a cache with a large capacity (for example, a capacity at the 1GB level), or may be a DRAM medium. The level 2 cache 3042 may cache the Cacheline data evicted from the level 1 cache and the data prefetched from the memory 220 (for example, DDR or PCM medium) at a granularity of 4KB per page. The Cacheline data is the data in the cache. For example, the cache in the cache unit 304 includes three parts, namely, significant bits, flag bits, and data bits. Each row includes these three types of data, and the data of one row forms one Cacheline. When starting a memory access request, the processing unit matches the data in the memory access request with the corresponding bits in the cache to read the Cacheline data in the cache or write data to the cache.

[0208] For example, when the memory sharing control device 300 is implemented by using the FPGA chip shown in FIG. 3, the function of the cache unit 304 may be implemented by configuring the BRAM in FIG. 3, or the functions of the level 1 cache 3041 and the level 2 cache 3042 may be implemented by configuring the BRAM in FIG. 3.

[0209] The cache unit 304 further includes a level 1 cache 3041 and a level 2 cache 3042, whereby the data access speed of the processing unit can be improved by using the cache, while the cache space can be increased, and the range within which the processing unit can quickly access the memory by using the cache is extended, so that the memory access rate of the processor resource pool is generally further improved.

[0210] FIG. 7D is a schematic diagram of the structure of another memory sharing control device 300 according to an embodiment of this application. As shown in FIG. 7D, the memory sharing control device 300 further includes a storage unit 305. The storage unit 305 may be a volatile memory, such as RAM, or may include a non-volatile memory, such as a read-only memory (ROM) or a flash memory. The storage unit 305 stores a program or instructions that can be read by the control unit 301, such as program code including at least one process or program code including at least one thread. The control unit 301 executes the program code in the storage unit 305 to implement the corresponding control.

[0211] The program code stored in the storage unit 305 may include at least one of a QoS engine 306, a prefetch engine 307, and a compression / decompression engine 308. FIG. 7D is used to conveniently display the functions related to the QoS engine 306, the prefetch engine 307, and the compression / decompression engine 308. These engines are shown outside the control unit 301 and the storage unit 305, but this does not mean that these engines are located outside the control unit 301 and the storage unit 305. In a specific implementation manner, the control unit 301 executes the corresponding code stored in the storage unit 305 to implement the corresponding functions of these engines.

[0212] The QoS engine 306 is configured to control the storage area of data to be accessed by the processing unit in the cache unit 304 (level 1 cache 3041 or level 2 cache 3042) based on the RESOURCE_ID in the memory access request, so that the memory data accessed by different processing units has different cache capabilities in the cache unit 304. For example, a memory access request initiated by a processing unit with a high priority has exclusive cache space in the cache unit 304. In this way, it can be ensured that the data accessed by the processing unit can be cached in a timely manner, thereby ensuring the service processing quality of this type of processing unit.

[0213] The prefetch engine 307 is configured to prefetch memory data based on a specific algorithm and prefetch the data to be read by the processing unit. Different prefetch methods affect the prefetch accuracy and memory access efficiency. The prefetch engine 307 realizes prefetch with higher accuracy based on a specified algorithm to further improve the hit rate when the processing unit accesses memory data. For example, the prefetch realized by the prefetch engine 307 includes, but is not limited to, prefetching Cacheline from the level 2 cache to the level 1 cache, or prefetching data from the external DRAM or PCM to the cache.

[0214] The compression / decompression engine 308 compresses or decompresses memory access data. For example, by using a compression ratio algorithm, it compresses the data written to the memory by the processing unit at a granularity of 4KB per page, and then writes the compressed data to the memory, or when the processing unit reads the compressed data in the memory, it decompresses the data to be read and then sends the decompressed data to the processing unit. Optionally, the compression / decompression engine may be disabled. In this way, the compression / decompression engine 308 is disabled and does not perform compression or decompression when the processing unit accesses the data in the memory.

[0215] The above QoS engine 306, prefetch engine 307, and compression / decompression engine 308 are stored in the storage unit 305 as software modules, and the control unit 301 reads the corresponding code in the storage unit to implement the corresponding functions. In an optional implementation method, at least one of the QoS engine 306, prefetch engine 307, and compression / decompression engine 308 may alternatively be directly configured in the control unit 301, and this function is implemented through the control logic of the control unit 301. In this way, the control unit 301 may execute the relevant control logic to implement the relevant functions and does not need to read the code in the storage unit 305. For example, when the memory sharing control device 300 is implemented by using the FPGA chip shown in FIG. 3, the relevant functions of the QoS engine 306, prefetch engine 307, and compression / decompression engine 308 may be implemented by configuring the CLB in FIG. 3.

[0216] Some of the QoS engine 306, the prefetch engine 307, and the compression / decompression engine 308 may be directly implemented by using the control unit 301, and some are stored in the storage unit 305. It can be understood that the control unit 301 reads the software code in the storage unit 305 to execute the corresponding functions. For example, the QoS engine 306 and the prefetch engine 307 are directly implemented through the control logic of the control unit 301. The compression / decompression engine 308 is software code stored in the storage unit 305, and the control unit 301 reads the software code of the compression / decompression engine 308 in the storage unit 305 to implement the functions of the compression / decompression engine 308.

[0217] For example, when the memory sharing control device 300 is implemented by using the FPGA chip shown in FIG. 3, the function of the storage unit 305 may be implemented by configuring the BRAM in FIG. 3.

[0218] An example in which the memory resources connected to the memory sharing control device 300 include DRAM and PCM is used below to illustrate an example of the implementation method of the memory access executed by the memory sharing control device 300.

[0219] FIG. 7E and FIG. 7F show two implementation methods in which DRAM and PCM are used as storage media for a shared memory pool. In the implementation method shown in FIG. 7E, DRAM and PCM are different types of memories included in the shared memory pool and have no hierarchical levels. When data is stored in the shared memory pool shown in FIG. 7E, the type of memory is not distinguished. Further, in FIG. 7E, the DDR controller 3031 controls the DRAM storage medium, and the PCM controller 3032 controls the PCM storage medium. The control unit 301 may access the DRAM via the DDR controller 3031 and access the PCM via the PCM controller 3032. However, in the implementation method shown in FIG. 7F, since DRAM has a higher speed and higher performance than PCM, DRAM may be used as the first-level memory, and data with a high access frequency may be preferentially stored in the DRAM. PCM is used as the second-level memory and is configured to store data that is not frequently accessed or data evicted from the DRAM. In FIG. 7F, the memory controller 303 includes two parts, namely, PCM control logic and DDR control logic. After receiving a memory access request, the control unit 301 accesses the PCM storage medium through the PCM control logic. When it is predicted that data will be accessed based on a preset algorithm or policy, the data may be cached in the DRAM in advance. In this way, subsequent access requests received by the control unit 301 may hit the corresponding data from the DRAM through the DDR control logic, thereby further improving the memory access efficiency.

[0220] In an optional implementation method, in the horizontal architecture shown in FIG. 7E, DRAM and PCM correspond to different memory spaces. For this architecture, the control unit 301 may store frequently accessed hot data in DRAM. In other words, it may establish a correspondence relationship between the processing unit that starts the process of accessing frequently accessed hot data and the virtual memory device corresponding to the DRAM memory. In this way, the read / write speed of memory data and the service life of the main memory system can be improved. The control unit 301 may establish a correspondence relationship between the processing unit that starts the process of accessing infrequently accessed cold data and the virtual memory device corresponding to the PCM memory in order to store infrequently accessed cold data in PCM. In this way, the security of important data can be ensured based on the non-volatile feature of PCM. In the vertical architecture shown in FIG. 7F, the features of high integration of PCM and low read / write latency of DRAM are used. The PCM-based main memory with a larger capacity may be configured to mainly store various types of data in order to reduce the number of disk accesses. Furthermore, DRAM is used as a cache to further improve memory access efficiency and performance.

[0221] Although a cache (Level 1 cache 3041 or Level 2 cache 3042), QoS engine 306, prefetch engine 307, and compression / decompression engine 308 are included in FIG. 7E or FIG. 7F, it should be noted that all of these components are optional in the implementation method. That is, in FIG. 7E or FIG. 7F, alternatively, these components may not be included, or at least one of these components may be included.

[0222] Based on the characteristics of different architectures in FIGS. 7E and 7F, the control unit 301 in the memory sharing control device 300 may create virtual memory devices with different characteristics under different architectures, and allocate the virtual memory devices to processing units with different service requirements, so that the requirements for the processing units to access memory resources can be more flexibly met, and the memory access efficiency of the processing units can be further improved.

[0223] FIG. 8A-1 is a schematic diagram of the structure of a computer device 80a according to an embodiment of this application. As shown in FIG. 8A-1, the computer device 80a includes a plurality of processors (processor 810a to processor 810a+N), a memory sharing control device 800a, a shared memory pool including N memories 820a, and a bus 840a. The memory sharing control device 800a is separately connected to the processors (processor 810a to processor 810a+N) via the bus 840a, and the shared memory pool (memories 820a to memories 820a+N) is connected to the bus 840a. In this embodiment of this application, N is a positive integer greater than or equal to 1.

[0224] In FIG. 8A-1, each processor 810a has its own local memory. For example, processor 810a has local memory 1. Each processor 810a may access its local memory, and when it is necessary to expand the memory resource to be accessed, it may also access the shared memory pool via the memory sharing control device 800a. The unified shared memory pool is shared by any of the processors 810a to 810a+N for use, so that not only can the utilization rate of memory resources be improved, but also the overly long latency caused by inter-processor access when processor 810a accesses the local memory controlled by other processors can be avoided.

[0225] Optionally, the memory sharing control device 800a in FIG. 8A-1 may alternatively include the logical functions of a network adapter, thereby enabling the memory sharing control device 800a to further access the memory resources of other computer devices through the network. This can further expand the scope of shared memory resources and improve the utilization rate of memory resources.

[0226] FIG. 8A-2 is a schematic diagram of the structure of another computer device 80a according to an embodiment of this application. As shown in FIG. 8A-2, the computer device 80a further includes a network adapter 830a, and the network adapter 830a is connected to a bus 840a through a Serdes interface. The memory sharing control device 800a may access the memory resources of other computer devices via the network adapter 830a. In the computer device 80a shown in FIG. 8A-2, the memory sharing control device 800a may not have the functions of a network adapter.

[0227] FIG. 8B-1 is a schematic diagram of the structure of a computer device 80b according to an embodiment of this application. As shown in FIG. 8B-1, the computer device 80b includes a processor resource pool including a plurality of processors 810b (processor 810b to processor 810b+N), a memory sharing control device 800b, a shared memory pool including a plurality of memories 820b (memory 820b to memory 820b+N), and a bus 840b. The processor resource pool is connected to the bus 840b via the memory sharing control device 800b, and the shared memory pool is connected to the bus 840b. Different from FIG. 8A-1 or FIG. 8A-2, each processor (any one of processors 810b to 810b+N) in FIG. 8B-1 or FIG. 8B-2 does not have its own local memory, and any memory access request of the processor is realized by using the memory sharing control device 800b within the shared memory pool including the memories 820b (memory 820b to memory 820b+N).

[0228] Optionally, the memory sharing control device 800a in FIG. 8B-1 may alternatively include the logical functions of a network adapter, thereby enabling the memory sharing control device 800b to further access the memory resources of other computer devices through the network. This can further expand the scope of shared memory resources and improve the utilization rate of memory resources.

[0229] FIG. 8B-2 is a schematic diagram of the structure of another computer device 80b according to an embodiment of this application. As shown in FIG. 8B-2, the computer device 80b further includes a network adapter 830b, and the network adapter 830b is connected to a bus 840b through a Serdes interface. The memory sharing control device 800b may access the memory resources of other computer devices through the network adapter 830b. In the computer device 80b shown in FIG. 8B-2, the memory sharing control device 800b may not have the functions of a network adapter.

[0230] The memory sharing control device 80a in FIGS. 8A-1 and 8A-2, and the memory sharing control device 80b in FIGS. 8B-1 and 8B-2 may be realized with reference to the implementation manner of the memory sharing control device 200 in FIG. 2A or FIG. 2B, or the memory sharing control device 300 in FIGS. 7A-7F. The processor 810a or the processor 810b may be realized with reference to the implementation manner of the processor in FIG. 6, and the memory 820a or the memory 820b may be a memory resource such as DRAM or PCM. The network adapter 830a is connected to the bus 840a through a serial interface, for example, a Serdes interface. The network adapter 830b is connected to the bus 840b through a serial interface, for example, a Serdes interface. The bus 840a or the bus 840b may be a PCIe bus.

[0231] In computer device 80a or computer device 80b, multiple processors may quickly access a shared memory pool via a memory sharing control device, which can improve the utilization rate of memory resources within the shared memory pool. Further, network adapter 830 is connected to the bus through a serial interface, and the data transmission latency between the processor and the network adapter does not increase significantly as the distance increases. Therefore, computer device 80a or computer device 80b may allocate the memory resources accessible by the processor via the memory sharing control device and the network adapter to other devices connected to computer device 80a or computer device 80b. Accordingly, the range of memory resources that can be shared by the processors is further extended, so that the memory resources are shared over a larger range and the utilization rate of the memory resources is further improved.

[0232] It can be understood that computer device 80a may alternatively include a processor without local memory. These processors access the shared memory pool via memory sharing control device 800a to realize memory access. Computer device 80b may alternatively include a processor with local memory. The processor may access the local memory or may access the memory within the shared memory pool via memory sharing control device 800b. Optionally, when some processors of computer device 80b have local memory, most of the memory accesses of these processors are realized in the local memory.

[0233] FIG. 9A is a schematic diagram of the structure of system 901 according to an embodiment of this application. As shown in FIG. 9A, system 901 includes M computer devices, for example, devices such as computer device 80a, computer device 81a, and computer device 82a. In this embodiment of this application, M is a positive integer greater than or equal to 3. The M computer devices are connected to each other through network 910a, and network 910a may be an Ethernet-based network or a USB-based network. Computer device 81a has the same structure as computer device 80a. The structure includes a processor resource pool including a plurality of processors (processor 8012a to processor 8012a+N), a memory sharing control device 8011a, a shared memory pool including a plurality of memories (memory 8013a to memory 8013a+N), a network adapter 8014a, and a bus 8015a. In computer device 81a, the processor resource pool, the memory sharing control device 8011a, and the network adapter 8014a are separately connected to bus 8015a, and the shared memory pool (memory 8013a to memory 8013a+N) is connected to the memory sharing control device 8011a. A memory access request initiated by a processor in the processor resource pool accesses the shared memory pool through the memory sharing control device 8011a. The memory sharing control device 8011a may be realized with reference to the realization manner of the memory sharing control device 200 in FIG. 2A or FIG. 2B, or the memory sharing control device 300 in FIGS. 7A-7F. Processor 8012a may be realized with reference to the realization manner of the processor in FIG. 6, and memory 8013a may be a memory resource such as DRAM or PCM. The network adapter 8014a is connected to the bus 8015a through a serial interface, for example, a Serdes interface. The bus 8015a may be a PCIe bus.

[0234] In FIG. 9A, each processor has its own local memory, and the local memory is the main memory resource for the processor to access the memory. Processor 8012a is used as an example. Processor 8012a may directly access local memory 1 of processor 8012a, and most of the memory access requests of processor 8012a may be realized in memory 1 of processor 8012a. When processor 8012a needs more memory to process a traffic burst, processor 8012a may access the memory resources in the shared memory pool via memory sharing control device 8011a to meet the memory resource requirements of processor 8012a. Optionally, processor 8012a may alternatively access the local memory of other processors, for example, access the local memory (memory N) of processor N. In other words, processor 8012a may alternatively access the local memory of other processors in a memory sharing manner in a NUMA system.

[0235] Computer device 82a and other computer devices M may have a structure similar to that of computer device 80a. Details will not be described again.

[0236] In system 901, processor 80a may access a shared memory pool including memory 8013a via memory sharing control device 800a, network adapter 830a, network 910a, network adapter 8014a, and memory sharing control device 8011a. In other words, the memory resources accessible by processor 810a include the memory resources within computer device 80a and the memory resources within computer device 81a. Similarly, processor 810a may alternatively access the memory resources of all computer devices within system 901. Thus, when a processor 8012a operating on a computer device, for example, computer device 81a, has a low service load and has a large amount of idle memory 8013a, but the processor 810a within computer device 80a requires a large amount of memory resources to execute an application such as HPC, the memory resources within computer device 81a may be allocated to the processor 810a within computer device 80a via memory sharing control device 800a. In this way, the memory resources within system 901 are effectively utilized. This not only meets the memory requirements of different computer devices for processing services, but also improves the utilization rate of memory resources within the entire system, thereby making it clearer how to improve the utilization rate of memory resources to reduce TCO.

[0237] In the system 901 shown in FIG. 9A, it should be noted that the computer device 80a includes a network adapter 830a. In a specific implementation manner, alternatively, the computer device 80a may not include the network adapter 830a, and the memory sharing control device 800a may include the control logic of the network adapter. Thus, the processor 810a may access other memory resources in the network via the memory sharing control device 8011a. For example, the processor 80a may access a shared memory pool including the memory 8013a via the memory sharing control device 800a, the network 910a, the network adapter 8014a, and the memory sharing control device 8011a. Alternatively, when the computer device 81a does not include the network adapter 8014a and the memory sharing control device 8011a realizes the function of the network adapter, the processor 80a may access a shared memory pool including the memory 8013a via the memory sharing control device 800a, the network 910a, and the memory sharing control device 8011a.

[0238] FIG. 9B is a schematic diagram of the structure of the system 902 according to an embodiment of this application. As shown in FIG. 9B, the system 902 includes M computer devices, for example, devices such as computer device 80b, computer device 81b, and computer device 82b, where M is a positive integer greater than or equal to 3. The M computer devices are connected to each other through a network 910b, and the network 910b may be an Ethernet-based network or a USB bus-based network. The computer device 81b has a structure similar to that of the computer device 80b. The structure includes a processor resource pool including a plurality of processors (processor 8012b to processor 8012b+N), a memory sharing control device 8011b, a shared memory pool including a plurality of memories (memory 8013b to memory 8013b+N), a network adapter 8014b, and a bus 8015b. The processor resource pool is connected to the bus 8015b through the memory sharing control device 8011b, and each of the shared memory pool and the network adapter 8014b is also connected to the bus 8015b. The memory sharing control device 8011b may be realized with reference to the implementation manner of the memory sharing control device 200 in FIG. 2A or FIG. 2B or the memory sharing control device 300 in FIGS. 7A to 7F. The processor 8012b may be realized with reference to the implementation manner of the processor in FIG. 6, and the memory 8013b may be a memory resource such as DRAM or PCM. The network adapter 8014b is connected to the bus 8015b through a serial interface, for example, a Serdes interface. The bus 8015b may be a PCIe bus.

[0239] The computer device 82b and the other computer devices M may have a structure similar to that of the computer device 80b. Details will not be described again.

[0240] In system 902, processor 810b may access a shared memory pool including memory 8013b via memory sharing control device 800b, network adapter 830b, network 910b, and network adapter 8014b. In other words, the memory resources accessible by processor 810b include the memory resources within computer device 80b and the memory resources within computer device 81b. Similarly, processor 810b may alternatively access the memory resources within all computer devices within system 902, thereby enabling the memory resources within system 902 to be used as shared memory resources. Thus, when a processor 8012 operating on a computer device, such as computer device 81b, has a low service load and has a large amount of idle memory 8013b, but the processor 810b within computer device 80b requires a large amount of memory resources to execute an application such as HPC, the memory resources within computer device 81b may be allocated to the processor 810b within computer device 80b via memory sharing control device 800b. In this way, the memory resources within system 902 are effectively utilized. This becomes clearer the aspect of satisfying the memory requirements of different computer devices for processing services, improving the utilization rate of the memory resources within system 902, and thereby improving the utilization rate of memory resources to reduce TCO.

[0241] In system 902 shown in FIG. 9B, it should be noted that computer device 80b includes network adapter 830b. In a specific implementation, alternatively, computer device 80b may not include network adapter 830b, and memory sharing control device 800b may include the control logic of the network adapter. Thus, processor 810b may access other memory resources in the network via memory sharing control device 8011b. For example, processor 80b may access a shared memory pool including memory 8013b via memory sharing control device 800b, network 910b, network adapter 8014b, and memory sharing control device 8011b. Alternatively, when computer device 81b does not include network adapter 8014b and memory sharing control device 8011b implements the function of the network adapter, processor 80b may access a shared memory pool including memory 8013b via memory sharing control device 800b, network 910b, and memory sharing control device 8011b.

[0242] FIG. 9C is a schematic diagram of the structure of system 903 according to an embodiment of this application. As shown in FIG. 9C, system 903 includes computer device 80a, computer device 81b, and computer devices 82c to M. The implementation of 80a in system 903 is the same as that of 80a in system 901, and the implementation of 81b in system 903 is the same as that of 81b in system 902. Computer devices 82C to M may be computer devices similar to computer device 80a, or may be computer devices similar to computer device 81b. System 903 integrates computer device 80a in system 910 and computer device 81b in system 920, and can improve the utilization rate of memory resources in the system through memory sharing.

[0243] In systems 901 to 903, it should be noted that the computer device needs to transmit a memory access request through the network. Since the network adapter 830b is connected to the memory sharing control device via a serial bus through a Serdes interface, the transmission rate and bandwidth of the serial bus can ensure the data transmission rate. Therefore, although network transmission affects the data transmission rate to a certain extent, from the perspective of improving the utilization rate of memory resources, in this way, while considering the memory access rate of the processor, the utilization rate of memory resources can be improved.

[0244] FIG. 10 is a schematic logic diagram showing how the computer device 80a shown in FIG. 8A-1 or FIG. 8A-2, or the computer device 80b shown in FIG. 8B-1 or FIG. 8B-2 realizes memory sharing, or may be a schematic logic diagram showing how the system 900 shown in FIGS. 9A to 9C realizes memory sharing.

[0245] As an example, the schematic logic diagram showing how the computer device 80a shown in FIG. 8A-1 realizes memory sharing is used. Processors 1 to 4 are any four processors (or cores within a processor) within the processor resource pool including the processor 810a, the memory sharing control device 1000 is the memory sharing control device 800a, and memories 1 to 4 are any four memories within the shared memory pool including the memory 820a. The memory sharing control device 1000 virtualizes four virtual memory devices (i.e., virtual memory 1 to virtual memory 4 shown in FIG. 10) based on memories 1 to 4, and the access control table 1001 records the correspondence between the virtual memory devices and the processors. When receiving a memory access request sent by any one of processors 1 to 4, the memory sharing control device 1000 obtains information about the virtual memory device corresponding to the processor that sent the memory access request based on the access control table 1001, and accesses the corresponding memory via the memory controller 1002 based on the obtained information about the virtual memory device.

[0246] The schematic logic diagram showing that the system 902 shown in FIG. 9B realizes memory sharing is used as an example. Processors 1 to 4 are four processors (or cores within a processor) of any one or more computer devices within the system 902, the memory sharing control device 1000 is a memory sharing control device of any computer device, and memories 1 to 4 are four memories of any one or more computer devices within the system 900. The memory sharing control device 1000 virtualizes four virtual memory devices (that is, virtual memories 1 to 4 shown in FIG. 10) based on memories 1 to 4, and the access control table 1001 records the correspondence between the virtual memory devices and the processors. When receiving a memory access request sent by any one of processors 1 to 4, the memory sharing control device 1000 obtains information about the virtual memory device corresponding to the processor that sent the memory access request based on the access control table 1001, and accesses the corresponding memory via the memory controller 1002 based on the obtained information about the virtual memory device.

[0247] FIG. 11 is a schematic diagram of the structure of a computer device 1100 according to an embodiment of this application. As shown in FIG. 11, the computer device 1100 includes at least two processing units 1102, a memory sharing control device 1101, and a memory pool. The processing unit is a processor, a core within a processor, or a combination of cores within a processor. The memory pool includes one or more memories 1103. At least two processing units 1102 are coupled to the memory sharing control device 1101. The memory sharing control device 1101 is configured to separately allocate the memory from the memory pool to at least two processing units 1102, and at least one memory within the memory pool is accessible by different processing units during different periods. At least two processing units 1102 are configured to access memory allocated via a memory sharing control device 1101.

[0248] The coupling of at least two processing units 1102 to a memory sharing control device 1101 means that at least two processing units 1102 are separately connected to the memory sharing control device 1101, and any one of the at least two processing units 1102 may be directly connected to the memory sharing control device 1101, or may be connected to the memory sharing control device 1101 via other hardware components (e.g., other chips).

[0249] For the specific implementation of the computer device 1100 shown in FIG. 11, refer to the implementation in FIGS. 8A-1, 8A-2, 8B-1, and 8B-2, or refer to the implementation of the computer device (e.g., computer device 80a or computer device 80b) in FIGS. 9A-9C, or refer to the implementation shown in FIG. 10. The memory sharing control device 1101 within the computer device 1100 may alternatively be implemented with reference to the implementation of the memory sharing control device 200 in FIG. 2A or FIG. 2B or the memory sharing control device 300 in FIGS. 7A-7F. Details will not be described again.

[0250] At least two processing units 1102 within the computer device 1100 shown in FIG. 11 can access at least one memory within the memory pool during different periods via the memory sharing control device 1101, thereby satisfying the memory resource requirements of the processing units and improving the utilization rate of the memory resources.

[0251] FIG. 12 is a schematic flowchart of a memory sharing control method according to an embodiment of this application. The method may be applied to the computer device shown in FIGS. 8A-1, 8A-2, 8B-1, or 8B-2, or may be applied to the computer devices in FIGS. 9A-9C (for example, computer device 80a or computer device 80b). The computer device includes at least two processing units, a memory sharing control device, and a memory pool. The memory pool includes one or more memories. As shown in FIG. 12, the method includes the following steps.

[0252] Step 1200: The memory sharing control device receives a first memory access request sent by a first processing unit among at least two processing units, and the processing unit is a processor, a core in the processor, or a combination of cores in the processor.

[0253] Step 1202: The memory sharing control device allocates a first memory from the memory pool to the first processing unit, and the first memory is accessible by a second processing unit among at least two processing units during other periods.

[0254] Step 1204: The first processing unit accesses the first memory via the memory sharing control device.

[0255] Based on the method shown in FIG. 12, different processing units access at least one memory in the memory pool during different periods, so that the memory resource requirements of the processing units can be satisfied, and the utilization rate of memory resources is improved.

[0256] Specifically, the method shown in FIG. 12 may be implemented with reference to the implementation manner of the memory sharing control device 200 in FIG. 2A or FIG. 2B or the memory sharing control device 300 in FIGS. 7A-7F. Details will not be described again.

[0257] Those skilled in the art may recognize that, referring to the examples described in the embodiments disclosed in this specification, the steps of the units and methods may be implemented by electronic hardware, computer software, or a combination thereof. In order to clearly explain the compatibility between hardware and software, the above generally describes the configurations and steps of each example according to functions. Whether a function is executed by hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but the implementation method should not be considered to exceed the scope of the present invention.

[0258] In some embodiments provided in this application, the described device embodiments are merely illustrative. For example, the division of units is merely a logical function division, and other divisions may be used in actual implementations. For example, a plurality of units or components may be combined or integrated into other systems, or some features may be ignored or not executed. Furthermore, the mutual coupling, direct coupling, or communication connection shown or discussed may be realized through some interfaces. The indirect coupling or communication connection between devices or units may be realized in electrical, mechanical, or other forms.

[0259] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. In other words, they may be located in one place or may be distributed on multiple network units. Some or all of the units may be selected depending on actual requirements to achieve the objectives of the solutions of the embodiments of the present invention.

[0260] The above description is merely a specific embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification or substitution that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention shall fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall comply with the protection scope of the claims.

Claims

1. A computer device including at least two processing units, a memory sharing control device, and a memory pool, wherein the processing unit is a processor, a core within the processor, or a combination of cores within the processor, and the memory pool includes one or more memories, the computer device being The at least two processing units are coupled to the memory sharing control device, The memory sharing control device is Connected to the at least two processing units via a serial bus, configured to receive a memory access request transmitted from the at least two processing units via the serial bus, and convert the memory access request into a parallel memory access request; an interface component A control unit connected to the interface component and configured to allocate memory from the memory pool to the at least two processing units, wherein at least one memory in the memory pool is accessible by different processing units during different periods; a control unit A cache stage connected to the control unit A plurality of memory controllers connected upstream to the cache stage and downstream to the one or more memories in the memory pool, configured to receive the parallel memory access request, each memory controller being configured to receive the parallel memory access request and access the corresponding memory in the memory pool; a plurality of memory controllers A computer device comprising.

2. The plurality of memory controllers are of different memory types, and the control unit is Configured to establish a correspondence between the memory address of the first memory and the first processing unit in order to allocate the first memory in the memory pool to the first processing unit among the at least two processing units. The computer device according to claim 1.

3. The control unit is Configured to virtualize a plurality of virtual memory devices from the memory pool, and the physical memory corresponding to the first virtual memory device among the plurality of virtual memory devices is the first memory. The computer device according to claim 2, configured to allocate the first virtual memory device to the first processing unit.

4. The control unit When a preset condition is satisfied, cancels the correspondence between the first virtual memory device and the first processing unit, The computer device according to claim 3, further configured to establish a correspondence between the first virtual memory device and a second processing unit among the at least two processing units.

5. The cache stage is configured to cache data read by any one of the at least two processing units from the memory pool, or to cache data evicted by any one of the at least two processing units, the computer device according to any one of claims 1 to 4.

6. The cache stage includes a prefetch engine, and the prefetch engine is configured to prefetch data that needs to be read by any one of the at least two processing units from the memory pool and cache the data in a cache within the cache stage, the computer device according to claim 5.

7. The cache stage further includes a quality of service (QoS) engine, and the QoS engine is configured to realize optimized storage of the data that needs to be cached by any one of the at least two processing units in the cache stage, the computer device according to claim 5.

8. The cache stage further includes a compression / decompression engine, and the compression / decompression engine is configured to compress or decompress data related to memory access, the computer device according to claim 1.

9. A system including at least two computer devices according to any one of claims 1 to 8, The at least two computer devices according to any one of claims 1 to 8 are connected to each other through a network, the system.

10. A memory sharing control device including a control unit, an interface component, a cache stage, and a plurality of memory controllers, The interface component is connected to at least two processing units via a serial bus, receives a memory access request transmitted from the at least two processing units via the serial bus, and is configured to convert the memory access request into a parallel memory access request. The processing unit is a processor, a core in the processor, or a combination of cores in the processor. The control unit is connected to the interface component and is configured to allocate memory from a memory pool to the at least two processing units. At least one memory in the memory pool is accessible by different processing units at different times. The cache stage is connected to the control unit. The plurality of memory controllers are connected upstream to the cache stage and downstream to the corresponding memories in the memory pool, and are configured to receive the parallel memory access requests. Each memory controller is a memory sharing control device configured to receive a parallel memory access request and access the corresponding memory in the memory pool.

11. The plurality of memory controllers are of different memory types. The control unit is configured to establish a correspondence between the memory address of the first memory and the first processing unit in order to allocate the first memory in the memory pool to the first processing unit. The memory sharing control device according to claim 10.

12. A memory sharing control method, which is applied to a computer device. The computer device includes at least two processing units, a memory sharing control device, and a memory pool. The memory pool includes one or more memories. The memory sharing control device includes an interface component connected to the at least two processing units via a serial bus, a control unit connected to the interface component, a cache stage connected to the control unit, and a plurality of memory controllers connected upstream to the cache stage and downstream to the one or more memories in the memory pool. The method comprises: A step of receiving, by the interface component of the memory sharing control device, a memory access request transmitted from the at least two processing units via the serial bus, wherein the processing unit is a processor, a core in the processor, or a combination of cores in the processor; A step of converting, by the interface component of the memory sharing control device, the memory access request into a parallel memory access request; A step of allocating, by the control unit of the memory sharing control device, memory from the memory pool to the at least two processing units, wherein at least one memory in the memory pool is accessible by different processing units at different times; A step of receiving, by each of the plurality of memory controllers, a parallel memory access request; A step of accessing, by each of the plurality of memory controllers, the corresponding memory in the memory pool A method comprising.

Citation Information

Patent Citations

  • Tajuenzanshorihoshiki

    JP1976080138A

  • Electronic filing device

    JP1992205571A

  • Memory control system

    JP2011070253A

  • Controller management of memory array of storage device using magnetic random access memory (MRAM)

    JP2015026379A

  • Memory control circuit, memory system, and processor system

    JP2018049381A