Smart memory data storage or loading method and device
Through location-aware memory mapping technology, the data storage fragmentation problem caused by physical memory interleaving in intelligent memory systems is solved, performance improvement and continuous processing of data sets are achieved, and the advantages of hardware interleaving are maintained.
Patent Information
- Application Number
- CN201810565251.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-06-30
- Filing Date
- 2018-06-04
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2038-06-04
AI Technical Summary
In existing intelligent memory systems, data storage fragmentation caused by physical memory interleaving affects the realization of performance improvements and the difficulty of sequentially processing data sets. Disabling interleaving will lead to overall performance loss.
Using location-aware memory mapping technology, the memory controller and hardware data processing logic blocks are used to selectively store or load data continuously into multiple memory units, reducing or eliminating data storage fragmentation caused by physical memory interleaving.
Effectively reduce or eliminate data storage fragmentation, maintain the performance benefits of hardware interleaving configurations, and support both sequential processing and parallel operations on data sets.
Smart Images

Figure CN109213697B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of computing and memory, and more particularly to a method and apparatus for storing or loading data in a smart memory arrangement. Background Art
[0002] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art by inclusion in this section.
[0003] An emerging generation of advanced computing systems may include smart memory. Smart memory is a term currently used to refer to memory subsystems that have near data processing (NDP) or in-memory processing (IMP) capabilities in addition to the memory or storage module itself. The memory is intelligent because the complementary NDP or IMP capabilities can process data after it is received by the memory subsystem.
[0004] Typically, for performance reasons, memory controllers that support smart memories also include hardware-implemented physical memory interleaving. If the software is well designed to read / write data on cache line boundaries, physical memory interleaving can provide a performance improvement in the range of 20%-30%. However, physical memory interleaving can have the disadvantage of potentially fragmenting the data being stored. As a result, processing the fragmented data may require another layer of merge operations (merging the results from two or more interleaved memory / storage modules), which reduces the performance gain that should be achieved. It is also very difficult or impossible to perform merge operations for some cases where complete data sets must be processed together or in sequence. And if physical memory interleaving is disabled to prevent data storage fragmentation, the entire performance gain will be lost. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The embodiments of the smart memory data storage / loading technology disclosed herein will be readily understood by the following detailed description in conjunction with the accompanying drawings. To facilitate this description, like reference numerals denote like structural elements. In the figures of the accompanying drawings, embodiments are illustrated by way of example and not by way of limitation.
[0006] Figure 1 A computing device with the disclosed smart memory data storage / loading techniques is illustrated, according to various embodiments.
[0007] Figure 2Illustrated is a page table for mapping memory pages of a virtual address space to interleaved memory frames of a smart memory according to various embodiments.
[0008] Figure 3 A driver configured to support location-aware specific references for memory mapped device accesses is illustrated in accordance with various embodiments.
[0009] Figure 4 An example location-aware smart memory data storage / loading process is illustrated in accordance with various embodiments.
[0010] Figure 5 An example computer system suitable for use in implementing aspects of the present disclosure is illustrated in accordance with various embodiments.
[0011] Figure 6 FIGURES 1 and 2 illustrate various embodiments of the present invention. Figure 1-4 A storage medium containing instructions for the method described. DETAILED DESCRIPTION
[0012] Disclosed herein are apparatuses, methods, and storage media associated with smart memory data storage / loading technology. In an embodiment, an apparatus may include: a processor; a plurality of memory units; a memory controller coupled to the processor and the plurality of memory units, the memory controller being configured to control access to the plurality of memory units, the control of access to the plurality of memory units including hardware physical memory interleaving support; and one or more hardware data processing logic blocks coupled to the plurality of memory units, the one or more hardware data processing logic blocks being configured to provide near data processing of data received by the plurality of memory units. The apparatus may further include a driver operated by one or more processors to support an application operated by the processor to perform location-aware memory-mapped device access to selectively store or load data continuously to a plurality of selected memory units or an aggregation of a plurality of selected memory units in the plurality of memory units, thereby reducing or eliminating memory fragmentation. These and other aspects will be further described below.
[0013] In the following detailed description, reference is made to the accompanying drawings forming a part hereof, wherein like reference numerals indicate like parts throughout, and wherein illustrative embodiments are shown by way of illustration. It should be understood that other embodiments may be employed, and structural or logical changes may be made without departing from the scope of this disclosure. Therefore, the following detailed description is not intended to be limiting, and the scope of the embodiments is defined by the appended claims and their equivalents.
[0014] The accompanying description discloses various aspects of the present disclosure. Alternative embodiments of the present disclosure and their equivalents may be designed without departing from the spirit or scope of the present disclosure. It should be noted that the same elements disclosed below are indicated by the same reference numerals in the accompanying drawings.
[0015] In a manner that is most helpful in understanding the claimed subject matter, each operation can be described in sequence as a plurality of discrete actions or operations. However, the order of description should not be interpreted as implying that the operations are necessarily dependent on the order. Specifically, the operations may not be performed in the order presented. The described operations may be performed in an order different from that of the described embodiments. In additional embodiments, various additional operations may be performed and / or the described operations may be omitted.
[0016] For the purposes of this disclosure, the phrase "A and / or B" means (A), (B), or (A and B). For the purposes of this disclosure, the phrase "A, B and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B and C).
[0017] The description may use the phrases "in an embodiment" or "in embodiments," each of which may refer to one or more of the same or different embodiments. Furthermore, the terms "comprising," "including," "having," and the like used with respect to the embodiments of the present disclosure are synonymous.
[0018] As used herein, the term "module" may refer to, be part of, or include an application specific integrated circuit (ASIC), an electronic circuit, a programmable combinatorial logic circuit (such as a field programmable gate array (FPGA)), a processor (shared, dedicated, or group) and / or memory (shared, dedicated, or group) that executes one or more software or firmware programs, and / or other suitable components that provide the described functionality.
[0019] Now refer to Figure 1, which shows a computing device with smart memory data storage / loading technology of the present disclosure according to various embodiments. As shown, in an embodiment, computing device 100 may include hardware 101, firmware (FW) / basic input / output services (BIOS) 106, operating system (OS) 112, and application 114, which are operably coupled to each other as shown. Hardware 101 may include one or more processors 102 (each processor having one or more processor cores and one or more memory controllers 117), conventional memory 104, and a number of smart memories 105. In an embodiment, each smart memory 105 may include a number of memory blocks 118 and NDP / IMP 116. In some embodiments, the number of smart memories 105 may be 1, 2, 3, 4, 5, 6, etc. In an embodiment, the memory controller(s) 117 may include hardware-implemented physical memory interleaving. For ease of understanding, the description will be presented in the context of a processor 102 with a memory controller 117, but this description is not intended to be construed as limiting. It is contemplated that the present disclosure may be implemented utilizing one or more processors, each having one or more memory controllers.
[0020] The storage locations of conventional memory 104 and memory block 118 can be organized into several memory frames. Typically, the size of a memory page, and therefore the size of a memory frame, can be processor and memory controller implementation dependent. In an embodiment, conventional memory 104 can be organized into 4KB memory pages. The size of the interleaved rows of a frame can be the same as or a multiple of the memory page size, and they are aligned on boundaries. For some embodiments where the memory page size is 4KB and data is interleaved across two smart memories 105, the size of a frame row can be 4KB, 8KB, and so on. In other words, for embodiments where a frame row is 4KB, addresses 0 to 4KB-1 come from the first smart memory 105, addresses 4KB to 8KB-1 come from the second smart memory 105, addresses 8KB to 12KB-1 come from the first smart memory 105, and so on. In alternative embodiments, the size of the interleaved rows can be smaller than the size of the memory page with additional support from processor 102, such as a mapping table of pages to interleaved rows.
[0021] OS 112 may include a kernel 130 that may be configured to perform core operating system functions / services, such as task scheduling, memory management / allocation, etc. Figure 2, applications 114 may be executed in their respective virtual address spaces 202, each of which has a number of virtual memory pages, e.g., virtual memory pages 204a-204f. Virtual memory pages 204a-204f may be selectively mapped to physical memory frames of conventional memory 104 or physical memory frames of smart memory 105. In embodiments, both mappings may be performed using paging tables associated with respective memories 104 and 105, such as example page table 206, which may be configured to map virtual memory pages 204a-204f of application 114 to interleaved memory frames 212-218 of smart memory 105 (when application 114 is active). In embodiments, when execution context is switched to application 112, a page table may be created in an area of conventional memory 104 allocated to OS 112 and loaded by OS 112 into the allocated area. In an embodiment, processor 102 may include a translation lookaside buffer (TLB) (not shown) to cache portions of multiple page tables, including the active application's page table 206. When a contiguous virtual address range is read or written, processor 102 may use the cached portion of the read page table 206 in the TLB and obtain the physical frame number of smart memory 105 interleaved in the hardware configuration.
[0022] Return Reference Figure 1 The kernel 130 may include several drivers 132 to interact with various devices of the computing system 100. In one embodiment, the drivers 132 may be configured with a memory-mapped device access driver (SMEM) 132 to facilitate memory-mapped device access by applications 114 using conventional memory 104 or smart memory 105. Specifically, SMEM 132 may be configured to populate page table 206 with mappings of virtual memory pages 204a-204f of the virtual address space 202 of the application 114 to memory frames of the smart memory 105. Furthermore, SMEM 132 may be combined with the smart memory data storage / loading techniques of the present disclosure. SMEM 132 may be configured to respond to a request from the application 114 dedicated to a plurality of selected smart memories or an aggregate of a plurality of selected smart memories in the smart memory 105 and return to the application 114 a range of virtual addresses of a specific smart memory 105 or a specific aggregate of memory cells 105 opened by the application 114.
[0023] Additional references Figure 3In example embodiments, the hardware 101 can include six (6) smart memories 105a-105f, the SMEM 132 can be configured to recognize invocations of SMEM0-SMEM5 132a-132f as corresponding to specific requests by the application 114 for the virtual address ranges of the corresponding smart memories 105a-105f. As a result, the application 114 can perform location-aware data storage or loading to selectively store or load data contiguously into multiple selected smart memories of the smart memories 105a-105f to reduce or eliminate data storage fragmentation problems caused by hardware-implemented physical memory interleaving. In alternative embodiments, the SMEM 132 can also be configured to recognize invocations of SMEM01, SMEM23, SMEM45 as corresponding to specific requests by the application 114 for the virtual address ranges of the three corresponding combined / aggregated smart memory units 105a-105b, 105c-105d, and 105e-105f. As a result, the application 114 can perform location-aware data storage or loading to selectively store or load data contiguously into selected pairs of the three combined / aggregated smart memories 105a-105b, 105c-105d, and 105e-105f to reduce or eliminate data storage fragmentation problems caused by hardware-implemented physical memory interleaving. In alternative embodiments, the combined / aggregated smart memory units 105a-105f can be more than 2 units and can be memory units that are not adjacent to each other.
[0024] In embodiments, the SMEM 132 can gather the mapping information of the smart memories 105 from a configuration table of the computer device 100 (e.g., an Advanced Configuration and Power Interface (ACPI) table maintained by the FW / BIOS 106) that has the mapping information of the smart memories 105 and compute the virtual address ranges of the smart memories 105 from the obtained mapping information. In alternative embodiments, the SMEM 132 can obtain the mapping information by discovering the memory channels in which the plurality of memories 105 are connected. In embodiments, the SMEM 132 can use serial presence detection or other known detection methods to discover the smart memory 105 information, such as the channel, size, etc. of the smart memories 105. Upon discovery, the SMEM 132 can compute the physical memory mapping information. Further, the SMEM driver 132 can form the post-interleaving physical address mapping for the smart memories 105 by detecting hardware-level registers to understand the configured interleaving settings.
[0025] Referring back to Figure 1Processor 102 can be any of several processors known in the art having one or more processor cores. In alternative embodiments, memory controller 117 can be externally located, separate from processor 102. Similarly, conventional memory 104 can be any volatile or non-volatile memory known in the art suitable for storing data. Memory 104 can include a hierarchy of cache memory and system memory. Both cache memory and system memory can be organized into cache pages and memory pages, respectively. Similarly, smart memory 105 can be any number of conventional memory / storage modules, including but not limited to dual inline memory modules (DIMMs), solder-down synchronous random access memory (SDRM) modules on the system board, shared memory, storage as memory, non-volatile DIMMs, and the like. NDP / IMP 116 can be implemented as an ASIC or any of several programmable circuits, including but not limited to FPGAs. Furthermore, in alternative embodiments, NDP / IMP 116 can be located external to smart memory 105, for example, along with processor 102, and can service multiple instances of smart memory 105.
[0026] In an embodiment, as shown in the figure, the hardware 101 may further include input / output (I / O) devices 108 or other elements (not shown). Examples of I / O devices may include Ethernet, WiFi, 3G / 4G, Communication or networking interfaces such as near field communication, universal serial bus (USB), storage devices such as solid-state, magnetic and / or optical drives, input devices such as keyboards, mice, touch-sensitive screens, and output devices such as display devices, printers, etc.
[0027] FW / BIOS 106 may be any of several FW / BIOS known in the art. In addition to the teachings of this disclosure, OS 112 may also be any of several OS known in the art, such as, for example, Windows OS, or Linux. Although for ease of understanding, the SMEM 132 has been described as being configured to recognize requests for virtual address ranges of a particular smart memory 105 or an aggregation of smart memories 105, in alternative embodiments, the SMEM 132 can be configured to recognize requests for virtual address ranges of a particular portion of a smart memory 105, thereby allowing the application 114 to implement contiguous location-aware data storage / loading at a particular group of storage blocks 118 of the smart memory 105. Thus, as used herein, the term memory can refer to a storage block or group of storage blocks within the smart memory 105, or the smart memory 105. In addition to the implementation of location-aware data storage / loading by the application 114 using the smart memory data storage / loading techniques according to the present disclosure to reduce / eliminate data storage fragmentation when using the smart memory 105, the application 114 can likewise be any of a number of applications known in the art.
[0028] Reference is now made to Figure 4 An example smart memory data storage / loading process according to embodiments is shown therein. As illustrated, the process 400 for storing data into a smart memory can include operations performed at blocks 402-408. The operations at blocks 402-408 can be performed by the OS 112 (in particular, by the SMEM 132) and the application 114 as previously described.
[0029] The process 400 can begin at block 402. At block 402, the number of smart memories can be detected, and a configuration for supporting contiguous location-aware data storage / loading can be determined, e.g., by the SMEM 132. The configuration can be determined according to input from a system administrator or according to a default specification. The configuration can be made available to applications in any of a number of known ways, e.g., through an ACPI configuration table. At block 404, upon determination, the memory mapped device access driver (e.g., the SMEM 132) can be configured (or self-configured) to recognize requests from applications for virtual address ranges of a particular smart memory (or aggregation of smart memories), or in some embodiments, portions of a particular smart memory.
[0030] Next, at block 406, the application 114 may selectively open specific smart memories or aggregates of smart memories (or, in alternative embodiments, portions of smart memories) in the smart memories and specifically request virtual address ranges for the individual smart memories or aggregates of smart memories (or, in alternative embodiments, portions of smart memories) in the smart memories. In a Linux embodiment, the application 114 may open a specific smart memory or aggregate of smart memories in the smart memories and invoke memory mapping using an Open&MMap call to specifically request the virtual address ranges of the smart memory or aggregate (or portion) to be opened. At block 408, the application 114 may perform location-aware data stores / loads on the specific smart memory or aggregate (or portion) to reduce or eliminate data storage fragmentation issues caused by hardware-implemented physical memory interleaving.
[0031] The following example illustrates the use of the disclosed smart memory data storage / loading techniques to reduce or eliminate data fragmentation issues caused by hardware-implemented physical memory interleaving when storing data in smart memories. Consider an example student database with one or more tables having columns for storing student name, birthday, height, weight, and parent income. For a current location-agnostic programming approach, in a Linux environment, an application would issue open / dev / mem to open the smart memory and then call mmap to obtain a memory mapping. The application could then write to the mapped memory: [student1][name], [student1][birthday], [student1][height], [student1][weight], [student1][parent income], [student2][name], [student2][birthday], [student2][height], [student2][weight], [student2][parent income], [student3][name], and so on. When data is stored in the smart memories, the hardware-interleaved data is thus fragmented across multiple smart memories.
[0032] However, using the smart memory data storage technology disclosed herein, applications can implement location-aware programming methods. For an example Linux embodiment in which smart memory units are individually identified, for as many smart memories as needed and available, the application can issue open / dev / smem1 to specifically open smart memory 0, then call mmap to obtain the memory mapping, issue open / dev / smem2 to specifically open smart memory 1, then call mmap to obtain the memory mapping. The application can then store [student 1][name], [student 2][name], [student 3][name], etc. in the mapped first smart memory (corresponding to smem1), [student 1][birthday], [student 2][birthday], [student 3][birthday], etc. in the mapped second smart memory (corresponding to smem2), and so on. Data can be interleaved into multiple data sets. The page table described above ensures that each data set is stored in a continuous order in the smart memory. In effect, the process defragments the memory fragments caused by the physical memory interleaving implemented by the smart memory hardware.
[0033] In this example, each individually identified smart memory stores columns of the database. When performing a search of the data, for example, to determine who are students with a height between 4 feet and 4.5 feet, a third smart memory can perform the search. Different search requests can also be sent to different smart memories in parallel. Thus, the smart memories can benefit from the hardware interleaving configuration while other operations in the same system continue to benefit from the hardware interleaving configuration.
[0034] In a real database environment, even if the platform setup is interleaved, the application can have as many or all columns as the database requires reach the smart memory in a contiguous manner (as seen by the NDP / IMP of the smart memory). Another smart memory can be used to accommodate another set of columns or another database in a contiguous manner as needed, as seen by the accelerator logic.
[0035] In some embodiments, if NDP is not required for certain operations, the smart memory can operate as normal memory. The basic concept is that when a location-agnostic application issues open / dev / mem to open the smart memory and then mmap it, the smart memory can operate as normal memory. Hardware implementation of interleaving is preferred to maintain the performance of normal memory applications in this operation. In other embodiments, different operating areas can be implemented for the smart memory, for example, one operating area for normal operations and another operating area for smart operations. In yet other embodiments, separate smart memories and separate normal memories can be interleaved together. The present disclosure enables the coexistence of interleaving and smart memories. When using smart memories, users do not lose the benefits of the hardware interleaving configuration.
[0036] The following are examples of different usage areas of smart memory. An example system may have 6 DIMMs with hardware-implemented interleaving, each DIMM may have 64GB of storage, providing a total of 384GB of storage. The user may decide that the system only needs 64GB for OS and application operations. The remaining 320GB of memory may be used for the in-memory database (IMDB). Therefore, the user may set certain system configuration parameters to set aside 64GB of smart memory for the OS and applications. After booting, the OS and applications will operate within the boundaries of the set aside 64GB of storage. From a physical perspective, each smart DIMM contributes a portion of the memory (10.66GB) for the OS and applications to run and use. The IMDB application understands that the 320GB portion is free, so it places the database to reside in that 320GB.
[0037] Now, the IMDB application will call the memory-mapped device access driver to gain access to the 320GB of storage. As data is stored into this 320GB portion, the hardware interleaves the data across the six SmartDIMMs. Without the techniques of this disclosure, the continuous stream of data would be split and distributed across the six SmartDIMMs.
[0038] However, using the present disclosure, the application can call specific smart memories: in the example Linux environment, call / dev / smem1, / dev / smem2 to / dev / smem6. Thereafter, the IMDB application can store / load data to / from smem1 in a continuous manner, store / load data to / from smem2 in a continuous manner..., and store / load data to / from smem6 in a continuous manner. The present disclosure enables the IMDB application to interleave data at the software level, which means that the IMDB application can generate many data sets and store one or more data sets to smem1 or smem2 to smem6 accordingly. A complete data set can be stored to smem1 or smem2 to smem6, or processed by smem1 or smem2 to smem6, or loaded from smem1 or smem2 to smem6. Data storage or processing or loading can also be completed in parallel with smem1 or smem2 until smem6, or completed by smem1 or smem2 until smem6, or completed from smem1 or smem2 until smem6.
[0039] In other words, in general, a smart memory arrangement having n dual in-line memory modules (DIMMs), each having a size of m GB, can each contribute k / n of the m GB memory locations for data storage that does not require NDP / IMP (such as, but not limited to, an operating system, or memory holes for memory-mapped I / O devices), and the remaining (mk / n) GB of memory locations of each of the n DIMMs can be set aside to meet the high-speed, low-latency data storage needs of applications that do require NDP / IMP (such as, but not limited to, in-memory databases).
[0040] Figure 5An example computer system that can be suitable for implementing selected aspects of the present disclosure is illustrated. As shown, computer system 500 can include one or more processors or processor cores 502 with application execution enclave support, read only memory (ROM) 503, and system memory 504. Additionally, computer system 500 can include mass storage devices 506. Examples of mass storage devices 506 can include, but are not limited to, tape drives, hard disk drives, compact disc read only memories (CD-ROMs), and the like. Moreover, computer system 500 can include input / output devices 508 (such as displays, keyboards, cursor control devices, etc.) and communication interfaces 510 (such as network interface cards, modems, etc.). These elements can be coupled to one another via system bus 512, which can represent one or more buses. In the case of multiple buses, they can be bridged by one or more bus bridges (not shown).
[0041] Each of these elements can perform its conventional functions known in the art. In particular, ROM 503 can comprise a Basic Input / Output System service (BIOS) 505. System memory 504 and mass storage devices 506 can be employed to store a working copy of the programming instructions collectively referred to as computing logic 522, which implement operations associated with application 114 and / or OS 112, as previously described, including kernel 130 with drivers (in particular, SMEM) 132. The various elements can be implemented by assembly instructions supported by processor(s) 502 or high-level languages, such as C, that can be compiled into such instructions, for example.
[0042] The number, capability, and / or capacity of these elements 510-512 can vary depending on whether computer system 500 is used as a mobile device, smart phone, computer tablet, laptop computer, etc., or as a stationary device, such as a desktop computer, server, game console, set-top box, infotainment console, etc. In other respects, the constitution of elements 510-512 is known and will not be further described accordingly.
[0043] As will be appreciated by those skilled in the art, the present disclosure can be embodied as a method or a computer program product. Accordingly, the present disclosure can take the form of an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects that can all generally be referred to as a "circuit," "module" or "system." Furthermore, the present disclosure can take the form of a computer program product embodied in any tangible or non-transitory medium having computer usable program code embodied in it. Figure 6An example computer-readable non-transitory storage medium is illustrated, which is suitable for storing instructions that, in response to the execution of these instructions by a device, cause the device to practice selected aspects of the present disclosure. As shown, a non-transitory computer-readable storage medium 602 may include several programming instructions 604. The programming instructions 604 may be configured to enable a device (e.g., a computer system 500) to implement (aspects of) an OS 112 and / or application 114 including a kernel 130 having a driver (specifically, SMEM) 132 in response to the execution of these programming instructions. In an alternative embodiment, on the contrary, these programming instructions 604 may be provided on multiple computer-readable non-transitory storage media 602. In yet other embodiments, the programming instructions 604 may be provided on a computer-readable transient storage medium 602 (such as a signal).
[0044] Any combination of one or more computer-usable or computer-readable media may be utilized. A computer-usable or computer-readable medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, apparatus, or propagation medium. More specific examples (a non-exclusive list) of computer-readable media would include the following: an electrical connector having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a transmission medium such as a transmission medium supporting the Internet or an intranet, or a magnetic storage device. Note that the computer-usable or computer-readable medium may even be paper or another suitable medium on which the program is printed, as the program may be captured electronically, for example, by optical scanning of the paper or other medium, and then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory. In the context of this document, a computer-usable or computer-readable medium can be any medium that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-usable medium may include a propagated data signal with the computer-usable program code embodied therewith in baseband or as part of a carrier wave. The computer-usable program code may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber cable, RF, etc.
[0045] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0046] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0047] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0048] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0049] The flowcharts and block diagrams in the figure illustrate the architecture, functions and operations of the possible implementations of the system, method and computer program product according to the various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a code module, a code segment or a code portion comprising one or more executable instructions for implementing (multiple) specified logical functions. It should also be noted that in some alternative implementations, the multiple functions marked in the box may not occur in the order marked in the figure. For example, depending on the functions involved, the two boxes shown in succession can actually be executed substantially simultaneously, or these boxes can sometimes be executed in the opposite order. It will also be noted that each box in the block diagram and / or flowchart illustration and the combination of multiple boxes in the block diagram and / or flowchart illustration can be realized by a special hardware-based system or a variety of combinations of special hardware and computer instructions that perform the specified functions or actions.
[0050] The terms used herein are used only to describe particular embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprise" and / or "comprising" are used in this specification, they specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0051] The embodiments may be implemented as a computer process, a computing system, or as an article of manufacture, such as a computer program product on a computer-readable medium. The computer program product may be a computer storage medium that can be read by a computer system and encode computer program instructions for executing a computer process.
[0052] The corresponding structures, materials, acts, and equivalents of all means or step-plus-function elements in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements to which the rights are expressly claimed. The description of the present disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or to limit the disclosure to the form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the disclosure. The embodiments were chosen and described in order to best explain the principles and practical application of the disclosure and to enable others skilled in the art to understand the disclosure of the embodiments with various modifications as are applicable to the particular use contemplated.
[0053] Return Reference Figure 5For one embodiment, at least one of the processors 502 may be packaged together with a memory having aspects of an OS 112, including a kernel 130 having drivers (specifically, SMEM) 132. For one embodiment, at least one of the processors 502 may be packaged together with a memory having aspects of an OS 112, including a kernel 130 having drivers (specifically, SMEM) 132, to form a system-in-package (SiP). For one embodiment, at least one of the processors 502 may be integrated on the same die with a memory having aspects of an OS 112, including a kernel 130 having drivers (specifically, SMEM) 132. For one embodiment, at least one of the processors 502 may be packaged together with a memory having aspects of an OS 112, including a kernel 130 having drivers (specifically, SMEM) 132, to form a system-on-chip (SoC). For at least one embodiment, the SoC may be utilized, but is not limited to being utilized, in, for example, a smartphone or a computing tablet.
[0054] Thus, various example embodiments of the present disclosure have been described, which include, but are not limited to:
[0055] Example 1 may be an apparatus for computing, comprising: a processor; a plurality of memory units; a memory controller coupled to the processor and the plurality of memory units, the memory controller being used to control access to the plurality of memory units, the control of access to the plurality of memory units including hardware physical memory interleaving support; one or more hardware data processing logic blocks coupled to the plurality of memory units, the one or more hardware data processing logic blocks being used to provide near data processing of data received by the plurality of memory units; and a driver operated by the processor, the driver being used to support an application operated by the processor to perform location-aware memory-mapped device access to selectively store data contiguously to a plurality of selected memory units or an aggregate of a plurality of selected memory units among the plurality of memory units, or to load data from a plurality of selected memory units or an aggregate of a plurality of selected memory units among the plurality of memory units.
[0056] Example 2 may be Example 1, wherein the driver may obtain, for the application, consecutive virtual addresses mapped for memory cells or an aggregate of a plurality of selected memory cells among the plurality of memory cells.
[0057] Example 3 may be Example 2, wherein for location-aware memory-mapped device access by an application, the application calls a driver to open one of multiple memory units or an aggregate of multiple memory units, and in response to the application's call, the driver specifically obtains a continuous virtual address mapped for the opened one of the memory units or the aggregate, and stores the selected data group in the mapped memory unit or the aggregate of the mapped memory units.
[0058] Example 4 may be Example 1, wherein the driver may further obtain a physical memory map of a memory unit in the plurality of memory units.
[0059] Example 5 may be Example 4, wherein the driver may obtain the physical memory map of the memory unit from an Advanced Configuration Power Interface (ACPI) table.
[0060] Example 6 may be Example 4, wherein the driver may further process physical address to virtual address translation, select physical frame numbers of the plurality of memory cells to map to virtual page numbers, and populate the mapping of virtual page numbers to physical frame numbers into a page table associated with the plurality of memory cells.
[0061] Example 7 may be Example 1, wherein the size of the interleaved rows of memory cells is equal to or a multiple of the size of the virtual memory page.
[0062] Example 8 may be Example 1, wherein the plurality of memory units is a dual inline memory module (DIMM), a plurality of DIMMs, a nonvolatile DIMM, or a memory module within a synchronous dynamic random access memory (SDRAM).
[0063] Example 9 may be example 1, further comprising an operating system having a kernel, the kernel including the driver.
[0064] Example 10 may be Example 1, wherein the memory unit may include n dual in-line memory modules (DIMMs), each DIMM having a size of m GB, and each DIMM contributes k / n of the m GB of memory locations to storage of data that does not use near data processing (NDP) or in-memory processing (IMP), and the remaining (mk / n) GB of memory locations of each of the n DIMMs are used for storage of data that uses the NDP or IMP.
[0065] Example 11 may be Example 1, wherein the memory controller is part of the processor.
[0066] Example 12 may be any of Examples 1-11, wherein the plurality of memory units and the one or more hardware data processing logic blocks may be components of one or more smart memory units.
[0067] Example 13 can be a method for computing, comprising: obtaining, by a driver operated by a processor of a computing device, one or more physical memory mappings of a plurality of memory units of a smart memory arrangement; populating, by the driver, a mapping of virtual page numbers to physical frame numbers into a page table associated with the memory units; and in response to a request by the computing device specific to one of the memory units or an aggregate of the plurality of memory units, providing, by the driver, a virtual address range of the one of the memory units or the aggregate of the memory units for application implementation of a location-aware memory mapping device access to selectively store data contiguously into the one of the memory units or the aggregate of the memory units; and wherein the memory units of the smart memory arrangement comprise hardware physical memory interleaving and near data processing logic, or are complementary to hardware physical memory interleaving and near data processing logic.
[0068] Example 14 can be Example 13, further comprising: an application invoking the driver to open an aggregate of one of the plurality of memory units or the plurality of memory units, in response to the application's invocation, the driver specifically obtaining contiguous virtual addresses mapped for the opened one of the memory units or the aggregate, and storing selected groups of data into the mapped memory unit or aggregate.
[0069] Example 15 can be Example 13, wherein the obtaining step can comprise: the driver obtaining the physical memory mappings of the memory units from an Advanced Configuration and Power Interface (ACPI) table.
[0070] Example 16 can be any of Examples 13-15, further comprising: the driver selecting physical frame numbers of the plurality of memory units to map to virtual page numbers; and handling translation of physical addresses to virtual addresses.
[0071] Example 17 can be one or more transitory or non-transitory computer-readable media (CRM) comprising instructions that, in response to execution of the instructions by a processor of a computing device, cause the computing device to provide a driver to: obtain one or more physical memory mappings of a plurality of memory units of a smart memory arrangement; populate a mapping of virtual page numbers to physical frame numbers into a page table associated with the memory units; and in response to a request by the computing device specific to one of the memory units or an aggregate of the memory units, provide a virtual address range of the one of the memory units or the aggregate for application implementation of a location-aware memory mapping device access to selectively store data contiguously into the one of the memory units or the aggregate; and wherein the plurality of memory units of the smart memory arrangement comprise near data processing logic and hardware physical memory interleaving, or are complementary to near data processing logic and hardware physical memory interleaving.
[0072] Example 18 may be example 17, wherein the obtaining may include obtaining a physical memory map of the memory unit from an Advanced Configuration Power Interface (ACPI) table.
[0073] Example 19 may be Example 17 or 18, wherein the driver may further select physical frame numbers of the plurality of memory cells to map to virtual page numbers and handle conversion of physical addresses to virtual addresses.
[0074] Example 20 may be a device for computing, comprising: a processor; a plurality of memory units; a memory controller coupled to the processor and the plurality of memory units, the memory controller for controlling access to the plurality of memory units, the controlling access to the plurality of memory units including hardware physical memory interleaving support; one or more hardware data processing logic blocks coupled to the plurality of memory units, the one or more hardware data processing logic blocks for providing near data processing of data received by the plurality of memory units; and a device for supporting an application operated by one or more processors to perform location-aware memory-mapped device access to selectively store data contiguously to a plurality of selected memory units among the plurality of memory units or an aggregated plurality of memory units, or to load data from a plurality of selected memory units among the plurality of memory units or an aggregated plurality of memory units.
[0075] Example 21 may be Example 20, wherein the means for supporting may include means for obtaining, for an application, contiguous virtual addresses mapped for a memory cell or an aggregate of a plurality of selected memory cells in a plurality of memory cells.
[0076] Example 22 may be example 21, wherein to perform location-aware memory-mapped device access, the application may call a supporting device to open a memory cell or an aggregate of multiple memory cells among multiple memory cells, call a supporting device to specifically obtain a contiguous virtual address mapped for the opened memory cell or aggregate of memory cells, and store the selected data group in the mapped memory cell or aggregate of memory cells among the memory cells.
[0077] Example 23 may be Example 20, wherein the means for supporting may include means for obtaining a physical memory map of a memory unit in the plurality of memory units.
[0078] Example 24 may be example 23, wherein the means for obtaining may include means for retrieving a physical memory mapping of the memory unit from an Advanced Configuration Power Interface (ACPI) table.
[0079] Example 25 may be Example 23, wherein the apparatus for supporting may include: an apparatus for processing conversion of physical addresses to virtual addresses; an apparatus for selecting physical frame numbers of a plurality of memory cells to map to virtual page numbers; and an apparatus for populating the mapping of virtual page numbers to physical frame numbers into a page table associated with the plurality of memory cells.
[0080] Example 26 may be Example 20, wherein a size of the interleaved rows of memory cells is equal to or a multiple of a size of the virtual memory page.
[0081] Example 27 may be Example 20, wherein the plurality of memory units is a dual inline memory module (DIMM), a plurality of DIMMs, a nonvolatile DIMM, or a memory module within a synchronous dynamic random access memory (SDRAM).
[0082] Example 28 may be example 20, further comprising an operating system having a kernel comprising means for supporting.
[0083] Example 29 may be Example 20, wherein the memory unit includes n dual in-line memory modules (DIMMs), each DIMM having a size of m GB, and each DIMM contributes k / n of the m GB of memory locations to storage of data that does not use near data processing (NDP) or in-memory processing (IMP), and the remaining (mk / n) GB of memory locations of each of the n DIMMs are used for storage of data that uses the NDP or IMP.
[0084] Example 30 may be example 20, wherein the memory controller is part of the processor.
[0085] Example 31 may be any of Examples 20-30, wherein the plurality of memory units and the one or more hardware data processing logic blocks are components of one or more smart memory units.
[0086] It will be apparent to those skilled in the art that various modifications and variations may be made in the disclosed embodiments of the disclosed apparatus and associated methods without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is intended to cover modifications and variations of the above-disclosed embodiments if they fall within the scope of any claims and their equivalents.
Claims
1. A device for computing, comprising: processor; a plurality of memory cells; a memory controller coupled to the processor and the plurality of memory units, the memory controller being configured to control access to the plurality of memory units, wherein controlling access to the plurality of memory units includes hardware physical memory interleaving support; one or more hardware data processing logic blocks coupled to the plurality of memory units, the one or more hardware data processing logic blocks for providing near data processing of data received by the plurality of memory units; as well as A driver, operated by the processor, for supporting an application operated by the processor to perform location-aware memory-mapped device access to selectively store data contiguously to or load data from a plurality of selected ones of the plurality of memory cells or an aggregate of the plurality of selected ones of the plurality of memory cells, to allow processing of the loaded back data without first reorganizing the data.
2. The device according to claim 1, wherein The driver is configured to obtain, for the application, consecutive virtual addresses mapped to memory units among the plurality of memory units.
3. The device according to claim 2, wherein In order to perform the location-aware memory-mapped device access of the application, the application calls the driver to open one of the multiple memory units or an aggregate of the multiple memory units. In response to the call of the application, the driver specifically obtains the consecutive virtual addresses mapped for the opened memory unit or the aggregate of memory units, and stores the selected data group in the mapped memory unit or the aggregate of memory units.
4. The device according to claim 1, wherein The driver is further configured to obtain a physical memory map of a memory unit in the plurality of memory units.
5. The device according to claim 4, wherein The driver is used for obtaining a physical memory mapping of a memory unit from an Advanced Configuration Power Interface (ACPI) table.
6. The device according to claim 4, wherein The driver is for further processing physical address to virtual address translation, selecting physical frame numbers of the plurality of memory units to map to virtual page numbers, and populating the mapping of virtual page numbers to physical frame numbers into a page table associated with the plurality of memory units.
7. The device according to claim 1, wherein The size of the interleaved rows of memory cells is equal to, or a multiple of, the size of a virtual memory page.
8. The device according to claim 1, wherein The plurality of memory units are dual inline memory modules (DIMMs), a plurality of DIMMs, non-volatile DIMMs, or memory modules within synchronous dynamic random access memory (SDRAMs).
9. The apparatus of claim 1, further comprising an operating system having a kernel, the kernel including the driver.
10. The device of claim 1, wherein: The plurality of memory units includes n dual in-line memory modules (DIMMs), each DIMM having a size of m GB, and each DIMM contributes k / n of the memory locations of the m GB to storage of data that does not use a near data processing (NDP) or an in-memory processing (IMP), and the remaining (mk / n) GB of memory locations of each of the n DIMMs are used for storage of data that uses the NDP or the IMP.
11. The device according to claim 1, wherein The memory controller is part of the processor.
12. The device according to any one of claims 1 to 11, wherein: The plurality of memory units and the one or more hardware data processing logic blocks are components of one or more smart memory units.
13. A method for computing, comprising: obtaining, by a driver operated by a processor of the computing device, one or more physical memory maps of a plurality of memory cells of the smart memory arrangement; populating, by the driver, a mapping of virtual page numbers to physical frame numbers into a page table associated with the plurality of memory cells; as well as providing, by the driver, a virtual address range of a selected one of the plurality of memory cells for the application to implement location-aware memory mapped device access to selectively store data contiguously to the one of the plurality of memory cells or the aggregate of the plurality of memory cells in response to a request by the computing device dedicated to the selected one of the plurality of memory cells or the aggregate of the plurality of memory cells to allow processing of the loaded data without first reassembling the data; The memory unit of the smart memory arrangement includes or is complementary to hardware physical memory interleaving and near data processing logic.
14. The method of claim 13, further comprising: The application calls the driver to open one of the multiple selected memory units or an aggregate of multiple selected memory units in the memory units. In response to the call of the application, the driver specifically obtains the continuous virtual addresses mapped for the opened memory units and stores the selected data group in the mapped memory unit or the aggregate of multiple memory units.
15. The method of claim 13, wherein: The acquiring step includes: the driver acquiring the physical memory mapping of the plurality of memory units from an Advanced Configuration Power Interface (ACPI) table.
16. The method of claim 13, further comprising: The driver selects physical frame numbers of the plurality of memory cells to map to virtual page numbers; And handles the translation of physical addresses to virtual addresses.
17. One or more computer readable media (CRMs) comprising instructions that, in response to execution of the instructions by a processor of a computing device, cause the computing device to provide a driver to perform any one of the methods of claims 13-16.
18. A device for computing, comprising: means for controlling access to a plurality of memory units, controlling access to the plurality of memory units comprising hardware physical memory interleaving support; means for providing near data processing of data received by said plurality of memory units; as well as Means for supporting an application operated by one or more processors to perform location-aware memory mapped device access to selectively store data contiguously to or load data from a plurality of selected ones of the plurality of memory cells or an aggregate of the plurality of selected ones of the plurality of memory cells to allow processing of the loaded back data without first reorganizing the data.
19. The apparatus of claim 18, wherein: The means for supporting includes means for obtaining, for the application, contiguous virtual addresses mapped for memory units in the plurality of memory units.
20. The apparatus of claim 19, wherein: In order to perform location-aware memory-mapped device access by an application, the application calls a supporting device to open one of the multiple memory units. In response to the call by the application, the supporting device specifically obtains a continuous virtual address mapped for the opened memory unit and stores the selected data group in the mapped memory unit.
21. The apparatus of claim 18, wherein: The means for supporting includes means for obtaining a physical memory map of the memory unit.
22. The apparatus of claim 21, wherein: The means for obtaining includes means for retrieving a physical memory mapping of the memory unit from an Advanced Configuration Power Interface (ACPI) table.
23. The apparatus of claim 21, wherein: The apparatus for supporting includes: an apparatus for handling translation of physical addresses to virtual addresses; an apparatus for selecting physical frame numbers of the plurality of memory cells to map to virtual page numbers; and an apparatus for populating a mapping of virtual page numbers to physical frame numbers into a page table associated with the plurality of memory cells.
24. The apparatus of any one of claims 18 to 23, wherein: The size of the interleaved rows of memory cells is equal to, or a multiple of, the size of a virtual memory page.
25. The apparatus of any one of claims 18 to 23, wherein: The plurality of memory units are dual inline memory modules (DIMMs), a plurality of DIMMs, non-volatile DIMMs, or memory modules within synchronous dynamic random access memory (SDRAMs).
Citation Information
Patent Citations
Systems and methods for implementing a stride value for accessing memory
US20080229035A1
Allocating and configuring persistent memory
US20160179375A1