Apparatus and method for user-by-user security access control with fine granularity
Through the fine-grained access control list and user-by-user authentication circuit supported by hardware extension, fine-grained security access control for large-scale data sets is realized, the problems of security vulnerabilities in the existing technology are solved, and the security and reliability of the system are improved.
Patent Information
- Application Number
- CN202411699700.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is vulnerable to security vulnerabilities caused by software defects and abuse of license inspection mechanisms when implementing fine-grained user-by-user secure access control for large-scale data sets.
Through hardware extensions supported fine-grained access control list (ACL) definitions and enforcement of memory access on a user-by-user-based basis, each user-by-job security indexing data structure and user-by-user authentication circuitry ensures that each user only accesses a subset of their authorized data.
It realizes fine-grained user-by-user access control for large-scale data sets, enhances security, and reduces the security risks brought by software defects and license abuse.
Smart Images

Figure CN120180460A_ABST
Abstract
Description
Background Art
[0001] This invention was made with government support under contract number W911NF-22-C-0081 awarded by the Army Research Office and IARPA. The government has certain rights in this invention. Technical Field
[0002] This invention generally relates to the field of computer processors. More specifically, this invention relates to apparatus and methods for fine-grained per-user secure access control.
[0004] For both high availability and utility purposes, large-scale data sets (up to petabytes in size) typically must remain resident in memory. Such data sets must be available to a large number of potential concurrent users while ensuring that users or groups of users are only permitted to access or modify specific subsets of the entire data set to enforce licensing or privacy issues.
[0005] Currently, per-user access is controlled through software abstractions that define a formalism for representing data and APIs for accessing and manipulating that data. The software is then responsible for enforcing the permissions for accessing the data. Although these methods are effective, they are inherently vulnerable to software bugs and security vulnerabilities that abuse the permission checking mechanism. Brief Description of the Drawings
[0006] A better understanding of the present invention can be obtained from the following detailed description in conjunction with the following drawings, in which:
[0007] Figure 1 Illustrative example computer system architecture;
[0008] Figure 2 Illustrative processor including multiple cores;
[0009] Figure 3(A) illustrates multiple stages of a processing pipeline;
[0010] Figure 3(B) illustrates details of one embodiment of a core;
[0011] Figure 4 Illustrative execution circuitry according to one embodiment;
[0012] Figure 5 Illustrative embodiment of a register architecture;
[0013] Figure 6 Illustrative example of an instruction format;
[0014] Figure 7 Illustrative addressing techniques according to one embodiment;
[0015] Figure 8 An embodiment of an illustrated instruction prefix;
[0016] Figure 9(A)-Figure 9(D) An embodiment of how to use the R, X, and B fields of the prefix;
[0017] Figure 10(A)-Figure 10(B) An example of an illustrated second instruction prefix;
[0018] Figure 11 The payload bytes of an embodiment of an illustrated instruction prefix;
[0019] Figure 12 Illustrating instruction conversion and binary translation implementations;
[0020] Figure 13-Figure 14 Is a schematic diagram of a virtual extension to Global Address Space (VEGAS) system according to one or more example embodiments of the present disclosure;
[0021] Figure 15 A flowchart of a process for an illustrative VEGAS system according to one or more example embodiments of the present disclosure;
[0022] Figure 16 Illustrating an example of a computing device or computing system according to one or more example embodiments of the present disclosure, on which any one of one or more techniques (e.g., methods) can be executed;
[0023] Figure 17 Illustrating a per-user authentication circuit according to some embodiments of the present disclosure;
[0024] Figure 18 Illustrating a secure index data structure according to some embodiments;
[0025] Figure 19 Illustrating example values of the relative amounts of memory used by embodiments of the present invention;
[0026] Figure 20 Illustrating an authentication circuit for providing access to internal and external memory access requests via an on-chip or off-chip network;
[0027] Figure 21 Illustrating a method according to some embodiments of the present invention. Detailed Description
[0028] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the invention described below. It will be apparent, however, to one of ordinary skill in the art that embodiments of the invention may be practiced without some of these specific details. In other instances, well-known structures and devices are shown in block diagram form to avoid obscuring the fundamental principles of embodiments of the invention.
[0029] Exemplary Computer Architecture
[0030] A description of an exemplary computer architecture is detailed below. Other system designs and configurations known in the art for laptop computers, desktop computers, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular telephones, portable media players, handheld devices, and a wide variety of other electronic devices are also suitable. In general, a variety of systems or electronic devices capable of incorporating the processors and / or other execution logic disclosed herein are suitable.
[0031] Figure 1 An embodiment of an exemplary system is illustrated. Multiprocessor system 100 is a point-to-point interconnect system and includes a plurality of processors, including a first processor 170 and a second processor 180 coupled via a point-to-point interconnect 150. In some embodiments, first processor 170 and second processor 180 are homogeneous. In some embodiments, first processor 170 and second processor 180 are heterogeneous.
[0032] Processors 170 and 180 are shown as including integrated memory controller (IMC) unit circuits 172 and 182, respectively. Processor 170 also includes point-to-point (P-P) interfaces 176 and 178 as part of its interconnect controller unit; similarly, second processor 180 includes P-P interfaces 186 and 188. Processors 170, 180 may exchange information via P-P interface circuits 178, 188 through P-P interconnect 150. IMCs 172 and 182 couple processors 170, 180 to respective memories, namely memory 132 and memory 134, which may be part of the main memory locally attached to each processor.
[0033] The processors 170, 180 may each utilize point-to-point interface circuits 176, 194, 186, 198 to exchange information with the chipset 190 via respective P-P interconnections 152, 154. The chipset 190 may optionally exchange information with the coprocessor 138 via a high-performance interface 192. In some embodiments, the coprocessor 138 is a dedicated processor, such as a high-throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, and so on.
[0034] A shared cache (not shown) may be included in either processor 170, 180, or outside of both processors but connected to these processors via a P-P interconnection, such that if a processor is placed in a low-power mode, the local cache information of either or both processors may also be stored in the shared cache.
[0035] The chipset 190 may be coupled to a first interconnect 116 via an interface 196. In some embodiments, the first interconnect 116 may be a Peripheral Component Interconnect (PCI) interconnect, or an interconnect such as a PCI Express interconnect or another I / O interconnect. In some embodiments, one of these interconnections is coupled to a power control unit (PCU) 117, which may include circuitry, software, and / or firmware to perform power management operations regarding the processors 170, 180 and / or the coprocessor 138. The PCU 117 provides control information to a voltage regulator such that the voltage regulator generates an appropriate regulated voltage. The PCU 117 also provides control information to control the generated operating voltage. In various embodiments, the PCU 117 may include various power management logic units (circuits) to perform hardware-based power management. Such power management may be fully controlled by the processor (e.g., controlled by various processor hardware and may be triggered by workload and / or power constraints, thermal constraints, or other processor constraints), and / or the power management may be performed in response to an external source (e.g., a platform or power management source or system software).
[0036] The PCU 117 is illustrated as existing as separate logic from the processor 170 and / or the processor 180. In other cases, the PCU 117 may execute on one or more given cores (not shown) within the core of the processor 170 or 180. In some cases, the PCU 117 may be implemented as a microcontroller (either dedicated or general-purpose) or other control logic that is configured to execute its own dedicated power management code (sometimes referred to as P-code). In still other embodiments, the power management operations to be performed by the PCU 117 may be implemented external to the processor, such as by a separate power management integrated circuit (PMIC) or another component external to the processor. In still other embodiments, the power management operations to be performed by the PCU 117 may be implemented within the BIOS or other system software.
[0037] A variety of I / O devices 114 and an interconnect (bus) bridge 118 may be coupled to the first interconnect 116, and the interconnect (bus) bridge 118 couples the first interconnect 116 to the second interconnect 120. In some embodiments, one or more additional processors 115 are coupled to the first interconnect 116, such as a coprocessor, a high-throughput MIC processor, a GPGPU, an accelerator (such as, for example, a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array (FPGA), or any other processor. In some embodiments, the second interconnect 120 may be a low pin count (LPC) interconnect. A variety of devices may be coupled to the second interconnect 120, such devices including, for example, a keyboard and / or a mouse 122, a communication device 127, and a storage unit circuit 128. The storage unit circuit 128 may be a disk drive or other mass storage device that may include instructions / code and data 130 in some embodiments. Additionally, audio I / O 124 may be coupled to the second interconnect 120. Note that other architectures are possible in addition to the point-to-point architecture described above. For example, a system such as the multiprocessor system 100 may implement a multi-drop interconnect or other such architecture instead of a point-to-point architecture.
[0038] Exemplary Core Architectures, Processors, and Computer Architectures
[0039] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, the implementation of these cores can include: 1) general-purpose in-order cores, for general computing purposes; 2) high-performance general-purpose out-of-order cores, for general computing purposes; 3) specialized cores, mainly for graphics and / or scientific (throughput) computing purposes. The implementation of different processors can include: 1) a CPU, including one or more general-purpose in-order cores for general computing purposes and / or one or more general-purpose out-of-order cores for general computing purposes; and 2) a coprocessor, including one or more specialized cores mainly for graphics and / or scientific (throughput) purposes. These different processors result in different computer system architectures, which can include: 1) the coprocessor and the CPU on separate chips; 2) the coprocessor and the CPU on separate dies within the same package; 3) the coprocessor and the CPU on the same die (in this case, such a coprocessor is sometimes referred to as specialized logic, such as integrated graphics and / or scientific (throughput) logic, or as a specialized core); and 4) a system-on-chip, which can include the above coprocessor and additional functions on the same die as the described CPU (sometimes referred to as the (one or more) application core or (one or more) application processor). An exemplary core architecture is described next, followed by a description of exemplary processors and computer architectures.
[0040] Figure 2 A block diagram illustrating an embodiment of an example processor 200 is shown. The processor 200 can have more than one core, can have an integrated memory controller, and can have an integrated graphics device. The processor 200 illustrated by the solid-line block diagram has a single core 202(A), a system agent 210, and a group of one or more interconnect controller unit circuits 216, while the optionally added dashed-line block diagram illustrates an alternative processor 200 as having multiple cores 202(A)-(N), a group of one or more integrated memory control unit circuits 214 in the system agent unit circuit 210, specialized logic 208, and a group of one or more interconnect controller unit circuits 216. Note that the processor 200 can be Figure 1 one of the processors 170 or 180 or the coprocessor 138 or 115.
[0041] Accordingly, different implementations of the processor 200 can include: 1) a CPU, where the dedicated logic 208 is integrated graphics and / or scientific (throughput) logic (which can include one or more cores, not shown), and the cores 202(A)-(N) are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor, where the cores 202(A)-(N) are a large number of dedicated cores mainly for graphics and / or scientific (throughput) purposes; and 3) a coprocessor, where the cores 202(A)-(N) are a large number of general-purpose in-order cores. Accordingly, the processor 200 can be a general-purpose processor, a coprocessor, or a special-purpose processor, such as a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit circuit), a high-throughput many integrated core (MIC) coprocessor (including 30 or more cores), an embedded processor, and so on. The processor can be implemented on one or more chips. The processor 200 can be part of one or more substrates and / or can be implemented on one or more substrates using any of a variety of process technologies, such as BiCMOS, CMOS, or NMOS.
[0042] The memory hierarchy includes one or more levels of cache unit circuits 204(A)-(N) within the cores 202(A)-(N), a group of one or more shared cache unit circuits 206, and external memory (not shown) coupled to the group of integrated memory controller unit circuits 214. The group of one or more shared cache unit circuits 206 can include one or more intermediate-level caches, such as a second-level (L2), third-level (L3), fourth-level (L4) or other levels of cache, such as a last level cache (LLC), and / or a combination of these. Although in some embodiments the ring-based interconnect network circuit 212 interconnects the dedicated logic 208 (e.g., integrated graphics logic), the group of shared cache unit circuits 206, and the system agent unit circuit 210, alternative embodiments use any number of well-known techniques to interconnect these units. In some embodiments, coherence is maintained between one or more of the circuits in the shared cache unit circuits 206 and the cores 202(A)-(N).
[0043] In some embodiments, one or more of cores 202(A)-(N) have multithreading capabilities. The system agent unit circuitry 210 includes those components that coordinate and operate the cores 202(A)-(N). The system agent unit circuitry 210 may include, for example, a power control unit (PCU) circuitry and / or a display unit circuitry (not shown). The PCU may be (or may include) the logic and components required to regulate the power states of the cores 202(A)-(N) and / or the dedicated logic 208 (e.g., integrated graphics logic). The display unit circuitry is used to drive one or more externally connected displays.
[0044] The cores 202(A)-(N) may be homogeneous or heterogeneous with respect to the architectural instruction set; that is, two or more of the cores 202(A)-(N) may be capable of executing the same instruction set, while other cores may be capable of executing only a subset of that instruction set or may be capable of executing a different ISA.
[0045] Exemplary Core Architectures
[0046] In-Order and Out-of-Order Core Block Diagrams
[0047] The block diagram of FIG. 3(A) illustrates an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline both in accordance with an embodiment of the present invention. The block diagram of FIG. 3(B) illustrates an exemplary embodiment of an in-order architecture core to be included in a processor and an exemplary register renaming, out-of-order issue / execution architecture core both in accordance with an embodiment of the present invention. The solid blocks in FIGS. 3(A)-(B) illustrate the in-order pipeline and the in-order core, while the optionally added dashed blocks illustrate the register renaming, out-of-order issue / execution pipeline and core. Given that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.
[0048] In FIG. 3(A), the processor pipeline 300 includes a fetch stage 302, an optional length decoding stage 304, a decode stage 306, an optional allocation stage 308, an optional rename stage 310, a schedule (also referred to as dispatch or issue) stage 312, an optional register read / memory read stage 314, an execution stage 316, a write-back / memory write stage 318, an optional exception handling stage 322, and an optional commit stage 324. One or more operations may be performed in each of these processor pipeline stages. For example, during the fetch stage 302, one or more instructions are fetched from an instruction memory. During the decode stage 306, the one or more fetched instructions may be decoded, an address using a forwarding register port (e.g., a load store unit (LSU) address) may be generated, and branch forwarding (e.g., immediate offset or link register (LR)) may be performed. In one embodiment, the decode stage 306 and the register read / memory read stage 314 may be combined into one pipeline stage. In one embodiment, during the execution stage 316, the decoded instructions may be executed, LSU address / data pipelining to an Advanced Microcontroller Bus (AMB) interface may be performed, multiplication and addition operations may be performed, arithmetic operations with branch results may be performed, and so on.
[0049] As an example, an exemplary register renaming, out-of-order issue / execution core architecture may implement the pipeline 300 as follows: 1) The instruction fetch 338 performs the fetch and length decoding stages 302 and 304; 2) The decode unit circuit 340 performs the decode stage 306; 3) The rename / allocator unit circuit 352 performs the allocation stage 308 and the rename stage 310; 4) The (one or more) scheduler unit circuits 356 perform the schedule stage 312; 5) The (one or more) physical register file unit circuits 358 and the memory unit circuit 370 perform the register read / memory read stage 314; The execution cluster 360 performs the execution stage 316; 6) The memory unit circuit 370 and the (one or more) physical register file unit circuits 358 perform the write-back / memory write stage 318; 7) Various units (unit circuits) may be involved in the exception handling stage 322; and 8) The retirement unit circuit 354 and the (one or more) physical register file unit circuits 358 perform the commit stage 324.
[0050] FIG. 3(B) shows that the processor core 390 includes a front-end unit circuit 330 coupled to an execution engine unit circuit 350, and both are coupled to a memory unit circuit 370. The core 390 can be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 390 can be a specialized core, such as a network or communication core, a compression engine, a co-processor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.
[0051] The front-end unit circuit 330 may include a branch prediction unit circuit 332, which is coupled to an instruction cache unit circuit 334, which is coupled to an instruction translation lookaside buffer (TLB) 336, which is coupled to an instruction fetch unit circuit 338, which is coupled to a decoding unit circuit 340. In one embodiment, the instruction cache unit circuit 334 is included in the memory unit circuit 370 rather than the front-end unit circuit 330. The decoding unit circuit 340 (or decoder) may decode the instruction and generate one or more micro-operations, microcode entry points, micro-instructions, other instructions, or other control signals as output, which are decoded from the original instruction, or otherwise reflect the original instruction, or are derived from the original instruction. The decoding unit circuit 340 may further include an address generation unit circuit (AGU, not shown). In one embodiment, the AGU uses the forwarded register ports to generate LSU addresses and may further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc.). The decoding unit circuit 340 may be implemented using various different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), microcode read only memories (ROMs), etc. In one embodiment, the core 390 includes a microcode ROM (not shown) or other medium that stores microcode for certain macro-instructions (e.g., in the decoding unit circuit 340 or otherwise within the front-end unit circuit 330). In one embodiment, the decoding unit circuit 340 includes a micro-operation (micro-op) or operation cache (not shown) to save / cache the decoded operations, micro-tags, or micro-operations generated during the decoding or other stages of the processor pipeline 300. The decoding unit circuit 340 may be coupled to a rename / allocator unit circuit 352 in the execution engine unit circuit 350.
[0052] The execution engine circuit 350 includes a rename / allocator unit circuit 352, which is coupled to a retirement unit circuit 354 and a set of one or more scheduler circuits 356. The scheduler circuits 356 represent any number of different schedulers, including reservation stations, a central instruction window, and so on. In some embodiments, the (one or more) scheduler circuits 356 may include an arithmetic logic unit (ALU) scheduler / scheduling circuit, an ALU queue, an address generation unit (AGU) scheduler / scheduling circuit, an AGU queue, and so on. The (one or more) scheduler circuits 356 are coupled to the (one or more) physical register file circuits 358. Each of the (one or more) physical register file circuits 358 represents one or more physical register files, and different physical register files among these physical register files store one or more different data types, such as scalar integer, scalar floating point, packed integer, packed floating point, vector integer, vector floating point, status (e.g., instruction pointer, i.e., the address of the next instruction to be executed), and so on. In one embodiment, the (one or more) physical register file unit circuits 358 include a vector register unit circuit, a write mask register unit circuit, and a scalar register unit circuit. These register units may provide architected vector registers, vector mask registers, general-purpose registers, and so on. The (one or more) physical register unit file circuits 358 are overlapped by the retirement unit circuit 354 (also referred to as a retirement queue) to illustrate various ways that can be used to implement register renaming and out-of-order execution (e.g., using the (one or more) reorder buffer (ROB) and the (one or more) retirement register file; using the (one or more) future file, the (one or more) history buffer, and the (one or more) retirement register file; using a register map and a pool of registers; and so on). The retirement unit circuit 354 and the (one or more) physical register file circuits 358 are coupled to the (one or more) execution clusters 360. The (one or more) execution clusters 360 include a set of one or more execution unit circuits 362 and a set of one or more memory access circuits 364. The execution unit circuits 362 may perform various arithmetic, logical, floating-point, or other types of operations (e.g., shift, add, subtract, multiply) on various types of data (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). Although some embodiments may include several execution units or execution unit circuits dedicated to a particular function or set of functions, other embodiments may include only one execution unit circuit or multiple execution units / execution unit circuits that perform all functions.(One or more) scheduler circuits 356, (one or more) physical register file unit circuits 358, and (one or more) execution clusters 360 are shown as potentially being plural because some embodiments create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating point / tight integer / tight floating point / vector integer / vector floating point pipelines, and / or memory access pipelines, each of which has its own scheduler circuit, (one or more) physical register file unit circuits, and / or execution cluster—and in the case of a separate memory access pipeline, in some embodiments only the execution cluster of that pipeline has (one or more) memory access unit circuits 364). It should also be understood that in cases where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution while the rest are in-order.
[0053] In some embodiments, the execution engine unit circuit 350 may perform load / store unit (LSU) address / data pipelining to a high-level microcontroller bus (AMB) interface (not shown), as well as address phase and write-back, data phase load, store, and branch.
[0054] A set of memory access circuits 364 is coupled to a memory unit circuit 370, which includes a data TLB unit circuit 372, which is coupled to a data cache circuit 374, which is coupled to a level 2 (L2) cache circuit 376. In one exemplary embodiment, the memory access unit circuit 364 may include a load unit circuit, a store address unit circuit, and a store data unit circuit, each of which is coupled to the data TLB circuit 372 in the memory unit circuit 370. An instruction cache circuit 334 is further coupled to the level 2 (L2) cache unit circuit 376 in the memory unit circuit 370. In one embodiment, the instruction cache 334 and the data cache 374 are combined into a single instruction and data cache (not shown) in the L2 cache unit circuit 376, a level 3 (L3) cache unit circuit (not shown), and / or main memory. The L2 cache unit circuit 376 is coupled to one or more other levels of cache and ultimately to main memory.
[0055] The core 390 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with newer versions); the MIPS instruction set; the ARM instruction set (with optional additional extensions such as NEON)), which includes the (one or more) instructions described herein. In one embodiment, the core 390 includes logic to support SIMD (Single Instruction, Multiple Data) instruction set extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using SIMD data.
[0056] Exemplary Execution Unit Circuit
[0057] Figure 4 Illustrates an embodiment of an (one or more) execution unit circuit, such as the (one or more) execution unit circuits 362 of FIG. 3(B). As shown, the (one or more) execution unit circuits 362 may include one or more ALU circuits 401, vector / SIMD unit circuits 403, load / store unit circuits 405, and / or branch / jump unit circuits 407. The ALU circuit 401 performs integer arithmetic and / or Boolean operations. The vector / SIMD unit circuit 403 performs vector / SIMD operations on packed data (e.g., SIMD / vector registers). The load / store unit circuit 405 executes load and store instructions to load data from memory into registers or store data from registers to memory. The load / store unit circuit 405 may also generate addresses. The branch / jump unit circuit 407 causes a branch or jump to a certain memory address depending on the instruction. The floating-point unit (FPU) circuit 409 performs floating-point arithmetic. The width of the (one or more) execution unit circuits 362 varies depending on the embodiment and may range from 16 bits to 1024 bits. In some embodiments, two or more smaller execution units are logically combined to form a larger execution unit (e.g., two 128-bit execution units are logically combined to form a 256-bit execution unit).
[0058] Exemplary Register Architecture
[0059] Figure 5 Is a block diagram of a register architecture 500 according to some embodiments. As shown, there are vector / SIMD registers 510, whose widths vary from 128 bits to 1024 bits. In some embodiments, the vector / SIMD registers 510 are physically 512 bits, and depending on the mapping, only some of the lower bits are used. For example, in some embodiments, the vector / SIMD register 510 is a 512-bit ZMM register: the lower 256 bits are used for YMM registers, and the lower 128 bits are used for XMM registers. Thus, there is register coverage. In some embodiments, the vector length field selects between a maximum length and one or more other shorter lengths, where each such shorter length is half the length of the previous length. Scalar operations are operations performed on the lowest-order data element positions in the ZMM / YMM / XMM registers; the higher-order data element positions are either kept the same as they were before the instruction or zeroed, depending on the embodiment.
[0060] In some embodiments, the register architecture 500 includes write mask / predicate registers 515. For example, in some embodiments, there are 8 write mask / predicate registers (sometimes referred to as k0 through k7), each of which is 16 bits, 32 bits, 64 bits, or 128 bits in size. The write mask / predicate registers 515 can allow merging (e.g., allowing any set of elements in the destination to be protected from update during the execution of any operation) and / or zeroing (e.g., a zeroing vector mask allows any set of elements in the destination to be zeroed) during the execution of any operation. In some embodiments, each data element position in a given write mask / predicate register 515 corresponds to a data element position in the destination. In other embodiments, the write mask / predicate registers 515 are scalable and consist of a set number of enable bits for a given vector element (e.g., 8 enable bits for each 64-bit vector element).
[0061] The register architecture 500 includes a plurality of general-purpose registers 525. These registers can be 16 bits, 32 bits, 64 bits, etc., and are capable of being used for scalar operations. In some embodiments, these registers are named RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.
[0062] In some embodiments, the register architecture 500 includes scalar floating-point registers 545, which are used for scalar floating-point operations on 32 / 64 / 80-bit floating-point data using the x87 instruction set extension, or as MMX registers to perform operations on 64-bit packed integer data, and to hold operands for some operations performed between MMX and XMM registers.
[0063] One or more flag registers 540 (e.g., EFLAGS, RFLAGS, etc.) store status and control information for arithmetic, comparison, and system operations. For example, one or more flag registers 540 can store condition code information such as carry, parity, auxiliary carry, zero, sign, and overflow. In some embodiments, one or more flag registers 540 are referred to as program status and control registers.
[0064] The segment registers 520 contain segment points for accessing memory. In some embodiments, these registers are named CS, DS, SS, ES, FS, and GS.
[0065] A machine-specific register (MSR) 535 controls and reports processor performance. Most MSRs 535 handle system-related functions and are not accessible to applications. The machine check register 560 consists of control, status, and error-reporting MSRs for detecting and reporting hardware errors.
[0066] One or more instruction pointer registers 530 store instruction pointer values. One or more control registers 555 (e.g., CR0-CR4) determine the operating mode of the processor (e.g., processors 170, 180, 138, 115, and / or 200) and the characteristics of the currently executing task. Debug registers 550 control and allow monitoring of the debug operations of the processor or core.
[0067] Memory management registers 565 specify the locations for data structures in protected mode memory management. These registers can include the GDTR, IDTR, task register, and LDTR registers.
[0068] Alternative embodiments of the present invention can use wider or narrower registers. Additionally, alternative embodiments of the present invention can use more, fewer, or different register banks and registers.
[0069] Instruction Set
[0070] An instruction set architecture (ISA) can include one or more instruction formats. A given instruction format can define various fields (e.g., number of bits, position of bits) to specify the operation to be performed (e.g., opcode) and the operand(s) and / or other data field(s) (e.g., mask) on which the operation is to be performed, etc. Some instruction formats are further decomposed by the definition of instruction templates (or sub-formats). For example, an instruction template of a given instruction format can be defined as having different subsets of the fields of that instruction format (the fields included are typically in the same order, but at least some have different bit positions since fewer fields are included) and / or be defined as having a given field interpreted in a different way. Thus, each instruction of the ISA is expressed using a given instruction format (and if defined, using a given instruction template in that instruction format's instruction templates), and includes fields for specifying the operation and operands. For example, an exemplary ADD instruction has a specific opcode and an instruction format that includes an opcode field to specify the opcode and an operand field to select the operands (source 1 / destination and source 2); and the occurrence of this ADD instruction in the instruction stream will have specific contents in the operand fields that select the specific operands.
[0071] Exemplary Instruction Formats
[0072] Embodiments of the (one or more) instructions described herein may be implemented in different formats. Additionally, exemplary systems, architectures, and pipelines are detailed below. Embodiments of the (one or more) instructions may be executed on these systems, architectures, and pipelines, but are not limited to those detailed.
[0073] Figure 6 Embodiments of the instruction format are illustrated. As shown, an instruction may include multiple components, which include but are not limited to one or more fields for the following: one or more prefixes 601, an opcode 603, addressing information 605 (e.g., register identifiers, memory addressing information, etc.), a displacement value 607, and / or an immediate value 609. Note that some instructions utilize some or all of the fields of this format, while other instructions may only use the fields of the opcode 603. In some embodiments, the illustrated order is the order in which these fields are to be encoded, however it should be understood that in other embodiments, these fields may be encoded in a different order, combined, etc.
[0074] The (one or more) prefix fields 601 modify the instruction when used. In some embodiments, one or more prefixes are used for repeat string instructions (e.g., 0xF0, 0xF2, 0xF3, etc.), provide section overrides (e.g., 0x2E, 0x36, 0x3E, 0x26, 0x64, 0x65, 0x2E, 0x3E, etc.), perform bus lock operations, and / or change operand (e.g., 0x66) and address sizes (e.g., 0x67). Certain instructions require mandatory prefixes (e.g., 0x66, 0xF2, 0xF3, etc.). Some of these prefixes may be considered “traditional” prefixes. Other prefixes (one or more examples of which are detailed herein) indicate and / or provide further capabilities, such as specifying particular registers, etc. These other prefixes typically follow the “traditional” prefixes.
[0075] The opcode field 603 is used to at least partially define the operation to be performed upon decoding of the instruction. In some embodiments, the length of the main opcode encoded in the opcode field 603 is 1, 2, or 3 bytes. In other embodiments, the main opcode may be of other lengths. An additional 3-bit opcode field is sometimes encoded in another field.
[0076] The addressing field 605 is used to address one or more operands of the instruction, such as a location in memory or one or more registers. Figure 7Illustrates an embodiment of the addressing field 605. In this illustration, an optional MOD R / M byte 702 and an optional Scale, Index, Base (SIB) byte 704 are shown. The MOD R / M byte 702 and the SIB byte 704 are used to encode up to two operands of an instruction, and each operand is either a direct register or an effective memory address. Note that each of these fields is optional, i.e., not all instructions include one or more of these fields. The MOD R / M byte 702 includes a MOD field 742, a register field 744, and an R / M field 746.
[0077] The content of the MOD field 742 differentiates between memory access and non-memory access modes. In some embodiments, when the MOD field 742 has the value b11, register direct addressing mode is utilized, otherwise register indirect addressing is used.
[0078] The register field 744 can encode a destination register operand or a source register operand, or can also encode an opcode extension without being used to encode any instruction operand. The content of the register index field 744 directly specifies or specifies through address generation the location of the source or destination operand (in a register or in memory). In some embodiments, the register field 744 is supplemented with additional bits from a prefix (e.g., prefix 601) to allow for greater addressing.
[0079] The R / M field 746 can be used to encode an instruction operand that references a memory address, or can be used to encode a destination register operand or a source register operand. Note that in some embodiments, the R / M field 746 can be combined with the MOD field 742 to specify an addressing mode.
[0080] The SIB byte 704 includes a scale field 752, an index field 754, and a base field 756 for address generation. The scale field 752 indicates a scaling factor. The index field 754 specifies the index register to be used. In some embodiments, the index field 754 is supplemented with additional bits from a prefix (e.g., prefix 601) to allow for greater addressing. The base field 756 specifies the base register to be used. In some embodiments, the base field 756 is supplemented with additional bits from a prefix (e.g., prefix 601) to allow for greater addressing. In practice, the content of the scale field 752 allows the content of the index field 754 to be scaled for memory address generation (e.g., for address generation using 2 缩放 * index + base).
[0081] Some addressing forms utilize displacement values to generate memory addresses. For example, according to 2缩放 * Index + base address + displacement, index * scale + displacement, r / m + displacement, instruction pointer (RIP / EIP) + displacement, register + displacement, etc. to generate a memory address. The displacement can be a value of 1 byte, 2 bytes, 4 bytes, etc. In some embodiments, the displacement field 607 provides this value. Additionally, in some embodiments, the use of the displacement factor is encoded in the MOD field of the addressing field 605, which indicates a compressed displacement scheme for which the displacement value is calculated by multiplying disp8 with a scaling factor N determined based on the vector length, a value of b bits, and the input element size of the instruction. The displacement value is stored in the displacement field 607.
[0082] In some embodiments, the immediate field 609 specifies an immediate value for the instruction. The immediate value can be encoded as a 1-byte value, 2-byte value, 4-byte value, etc.
[0083] Figure 8 Illustrates an embodiment of the first prefix 601(A). In some embodiments, the first prefix 601(A) is an embodiment of the REX prefix. Instructions using this prefix can specify general-purpose registers, 64-bit packed data registers (e.g., single instruction multiple data (SIMD) registers or vector registers), and / or control registers and debug registers (e.g., CR8 - CR15 and DR8 - DR15).
[0084] Instructions using the first prefix 601(A) can specify up to three registers using a 3-bit field, depending on the format: 1) using the reg field 744 and R / M field 746 of the MOD R / M byte 702; 2) using the MOD R / M byte 702 and the SIB byte 704, including using the reg field 744 and the base address field 756 and index field 754; or 3) using the register field of the opcode.
[0085] In the first prefix 601(A), bit positions 7:4 are set to 0100. Bit position 3 (W) can be used to determine the operand size but cannot determine the operand width alone. Thus, when W = 0, the operand size is determined by the code segment descriptor (CS.D), and when W = 1, the operand size is 64 bits.
[0086] Note that adding another bit allows addressing of 16 (2 4 ) registers, while the separate MOD R / M reg field 744 and the R / M field 746 of MOD R / M can each address only 8 registers.
[0087] In the first prefix 601(A), bit position 2 (R) can be an extension of the reg field 744 of MOD R / M, and can be used to modify the reg field 744 of MOD R / M when this field encodes a general-purpose register, a 64-bit packed data register (e.g., an SSE register), or a control or debug register. When the MOD R / M byte 702 specifies other registers or defines an extended opcode, R is ignored.
[0088] Bit position 1 (X) The X bit can modify the SIB byte index field 754.
[0089] Bit position B (B) B can modify the base address in the R / M field 746 of MOD R / M or the SIB byte base address field 756; or it can modify the opcode register field for accessing a general-purpose register (e.g., general-purpose register 525).
[0090] Figures 9(A)-(D) illustrate examples of how the R, X, and B fields of the first prefix 601(A) are used. Figure 9(A) illustrates that when the SIB byte 704 is not used for memory addressing, R and B from the first prefix 601(A) are used to extend the reg field 744 and the R / M field 746 of the MOD R / M byte 702. Figure 9(B) illustrates that when the SIB byte 704 is not used, R and B from the first prefix 601(A) are used to extend the reg field 744 and the R / M field 746 of the MOD R / M byte 702 (register-register addressing). Figure 9(C) illustrates that when the SIB byte 704 is used for memory addressing, R, X, and B from the first prefix 601(A) are used to extend the reg field 744, the index field 754, and the base address field 756 of the MOD R / M byte 702. Figure 9(D) illustrates that when a register is encoded in the opcode 603, B from the first prefix 601(A) is used to extend the reg field 744 of the MOD R / M byte 702.
[0091] Figures 10(A)-(B) illustrate examples of the second prefix 601(B). In some embodiments, the second prefix 601(B) is an embodiment of the VEX prefix. The second prefix 601(B) encodes to allow an instruction to have more than two operands, and allows SIMD vector registers (e.g., vector / SIMD register 510) to be longer than 64 bits (e.g., 128 bits and 256 bits). The use of the second prefix 601(B) provides a syntax for three-operand (or more) operations. For example, a previous two-operand instruction performed an operation such as A = A + B, which overwrote the source operand. The use of the second prefix 601(B) enables the operands to perform non-destructive operations, such as A = B + C.
[0092] In some embodiments, the second prefix 601(B) has two forms - a two-byte form and a three-byte form. The two-byte second prefix 601(B) is mainly used for 128-bit, scalar, and some 256-bit instructions; while the three-byte second prefix 601(B) provides a compact replacement for 3-byte opcode instructions and the first prefix 601(A).
[0093] Figure 10(A) illustrates an embodiment of the two-byte form of the second prefix 601(B). In one example, the format field 1001 (byte 0 1003) contains the value C5H. In one example, byte 1 1005 includes the value "R" in bit [7]. This value is the complement of the same value of the first prefix 601(A). Bit [2] is used to specify the length (L) of the vector (where a value of 0 is a scalar or 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). The bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones' complement) form and is valid for instructions with two or more source operands; 2) encoding the destination register operand, which is specified in ones' complement form for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, such as 1111b.
[0094] Instructions using this prefix can use the R / M field 746 of MOD R / M to encode the instruction operand that references a memory address, or to encode the destination register operand or source register operand.
[0095] Instructions using this prefix can use the reg field 744 of MOD R / M to encode the destination register operand or source register operand, which is treated as an opcode extension and not used to encode any instruction operand.
[0096] For instruction syntax that supports four operands, vvvv, the R / M field 746 of MOD R / M, and the reg field 744 of MOD R / M encode three of the four operands. Then bits [7:4] of the immediate 609 are used to encode the third source register operand.
[0097] Figure 10(B) illustrates an embodiment of the three - byte form of the second prefix 601(B). In one example, the format field 1011 (byte 0 1013) contains the value C4H. Byte 1 1015 includes "R", "X", and "B" in bits [7:5], which are the complements of these values of the first prefix 601(A). The bits [4:0] of byte 1 1015 (shown as mmmmm) include content for encoding one or more implicit leading opcode bytes as needed. For example, 00001 means a 0FH leading opcode, 00010 means a 0F38H leading opcode, 00011 means a 0F3AH leading opcode, and so on.
[0098] The use of bit [7] of byte 2 1017 is similar to that of W in the first prefix 601(A), including helping to determine the size of the operand that can be promoted. Bit [2] is used to specify the length (L) of the vector (where a value of 0 is a scalar or a 128 - bit vector, and a value of 1 is a 256 - bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). The bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones - complement) form and is valid for instructions with two or more source operands; 2) encoding the destination register operand, which is specified in ones - complement form for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, such as 1111b.
[0099] Instructions using this prefix can use the R / M field 746 of MOD R / M to encode instruction operands that reference a memory address, or to encode a destination register operand or a source register operand.
[0100] Instructions using this prefix can use the reg field 744 of MOD R / M to encode a destination register operand or a source register operand, or be treated as an opcode extension and not be used to encode any instruction operand.
[0101] For instruction syntax that supports four operands, vvvv, the R / M field 746 of MOD R / M, and the reg field 744 of MOD R / M encode three of the four operands. Then the bits [7:4] of the immediate 609 are used to encode the third source register operand.
[0102] Figure 11Illustrates an embodiment of a third prefix 601(C). In some embodiments, the first prefix 601(A) is an embodiment of an EVEX prefix. The third prefix 601(C) is a four-byte prefix.
[0103] The third prefix 601(C) can encode 32 vector registers (e.g., 128-bit, 256-bit, and 512-bit registers) in 64-bit mode. In some embodiments, instructions that utilize a write mask / operation mask (see the discussion of registers in the previous figures, e.g., Figure 5 ) or a predicate utilize this prefix. The operation mask register allows conditional processing or select control. Operation mask instructions - whose source / destination operand is an operation mask register and treats the contents of the operation mask register as a single value - are encoded using the second prefix 601(B).
[0104] The third prefix 601(C) can encode function-specific to instruction classes (e.g., packed instructions with "load + operation" semantics can support an embedded broadcast function, floating-point instructions with rounding semantics can support a static rounding function, floating-point instructions with non-rounding arithmetic semantics can support a "suppress all exceptions" function, etc.).
[0105] The first byte of the third prefix 601(C) is the format field 1111, which has a value of 62H in one example. The subsequent bytes are referred to as payload bytes 1115 - 1119 and together form a 24-bit value of P[23:0], providing specific capabilities in the form of one or more fields (detailed herein).
[0106] In some embodiments, P[1:0] of payload byte 1119 is the same as the two lowermost mm bits. In some embodiments, P[3:2] is reserved. Bit P[4] (R') allows access to the upper 16 vector register set when combined with P[7] and the reg field 744 of MOD R / M. When SIB type addressing is not required, P[6] can also provide access to the upper 16 vector registers. P[7:5] consists of R, X, and B, which are operand specifier modifier bits for vector registers, general-purpose registers, and memory addressing operands, and when combined with the MOD R / M register field 744 and the R / M field 746 of MOD R / M, allow access to the next set of 8 registers beyond the lower 8 registers. P[9:8] provides opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). P
[10] is a fixed value 1 in some embodiments. P[14:11], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones' complement) form and is valid for instructions with 2 or more source operands; 2) encoding the destination register operand, which is specified in ones' complement form for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, e.g., 1111b.
[0107] P
[15] is similar to the W of the first prefix 601(A) and the second prefix 611(B), and can serve as an opcode extension bit or an operand size promotion.
[0108] P[18:16] specifies the index of the register in the operation mask (write mask) register (e.g., write mask / predicate register 515). In one embodiment of the present invention, the specific value aaa = 000 has a special behavior, implying that no operation mask is used for this particular instruction (which can be achieved in various ways, including using a hard-wired all-ones operation mask or hardware that bypasses the masking hardware). When combined, the vector mask allows any set of elements in the destination to be protected from update during the execution of any operation (specified by the base and enhanced operations); in another embodiment, the old value of each element of the destination is retained (if the corresponding mask bit has a value of 0). In contrast, when zeroing, the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base and enhanced operations); in one embodiment, the elements of the destination are set to 0 when the corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span of the elements being modified, from the first to the last); however, the elements being modified do not have to be contiguous. Thus, the operation mask field allows for partial vector operations, including loads, stores, arithmetic, logic, etc. Although in the described embodiments of the present invention, the content of the operation mask field selects which one of several operation mask registers contains the operation mask to be used (so the content of the operation mask field indirectly identifies the masking to be performed), alternatively or additionally, alternative embodiments allow the content of the mask write field to directly specify the masking to be performed.
[0109] P
[19] can be combined with P[14:11] to encode the second source vector register in a non-destructive source syntax that can utilize P
[19] to access the upper 16 vector registers. P
[20] encodes various functions that vary among different classes of instructions and can affect the meaning of the vector length / rounding control specifier field (P[22:21]). P
[23] indicates support for merge-write masking (e.g., when set to 0) or support for zeroing and merge-write masking (e.g., when set to 1).
[0110] The following table details an exemplary embodiment of the encoding of registers in instructions using the third prefix 601(C).
[0111]
[0112] Table 1: 32-register support in 64-bit mode
[0113]
[0114] Table 2: Encoding register specifiers in 32-bit mode
[0115]
[0116]
[0117] Table 3: Operation Mask Register Specifier Encoding
[0118] Program code can be applied to the input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0119] The program code can be implemented in a high-level procedural or object-oriented programming language to communicate with the processing system. If desired, the program code can also be implemented in assembly or machine language. In fact, the mechanisms described herein are not limited in scope to any particular programming language. In any case, the language can be a compiled language or an interpreted language.
[0120] Embodiments of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or a combination of these implementation approaches. Embodiments of the present invention can be implemented as a computer program or program code, executed on a programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0121] One or more aspects of at least one embodiment can be implemented by representative instructions stored on a machine-readable medium, which represent various logics within a processor. When these instructions are read by the machine, they cause the machine to fabricate the logic for performing the techniques described herein. These representations are referred to as "IP cores" and can be stored on a tangible machine-readable medium and provided to various customers or manufacturing facilities to be loaded into the fabrication machines that actually make the logic or processor.
[0122] These machine-readable storage media can include - but are not limited to - non-transitory tangible arrangements of articles manufactured or formed by a machine or device, including storage media such as: hard disks, any other type of disk (including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW), and magneto-optical disks), semiconductor devices (such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM)), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.
[0123] Accordingly, embodiments of the present invention also include non-transitory tangible machine-readable media that contain instructions or contain design data defining the structures, circuits, devices, processors, and / or system features described herein, such as a Hardware Description Language (HDL). Such embodiments may also be referred to as program products.
[0124] Emulation (including binary translation, code morphing, etc.)
[0125] In some cases, an instruction converter may be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter may translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation), morph, emulate, or otherwise convert the instructions to one or more other instructions to be processed by the core. The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be on the processor, outside the processor, or part on the processor and part outside the processor.
[0126] Figure 12The figure is a block diagram of converting binary instructions in a source instruction set into binary instructions in a target instruction set using a control software instruction converter according to some implementations. In the illustrated embodiment, the instruction converter is a software instruction converter, but alternatively, the instruction converter may be implemented using software, firmware, hardware, or various combinations thereof. Figure 12 It shows that a program in a high-level language 1202 can be compiled using a first ISA compiler 1204 to generate first ISA binary code 1206, which can be natively executed by a processor 1216 having at least one first ISA instruction set core. The processor 1216 having at least one first ISA instruction set core represents any such processor that can perform substantially the same functions as an Intel processor having at least one first ISA instruction set core by compatibly executing or otherwise processing (1) a substantial portion of the instruction set of the first ISA instruction set core or (2) a target code version of an application or other software targeted to run on a processor having at least one first ISA instruction set core, so as to achieve substantially the same results as a processor having at least one first ISA instruction set core. The first ISA compiler 1204 represents a compiler operable to generate first ISA binary code 1206 (e.g., target code), which can be executed on the processor 1216 having at least one first ISA instruction set core with or without additional linking processing. Similarly,
[0127] It shows that a program in a high-level language 1202 can be compiled using an alternative instruction set compiler 1208 to generate alternative instruction set binary code 1210, which can be natively executed by a processor 1214 without a first ISA core. An instruction converter 1212 is used to convert the first ISA binary code 1206 into code that can be natively executed by a processor 1214 without a first ISA instruction set core. This converted code may not be the same as the alternative instruction set binary code 1210 because it is difficult to make an instruction converter that can do this; however, the converted code will implement the overall operation and consist of instructions from the alternative instruction set. Thus, the instruction converter 1212 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device without a first ISA instruction set processor or core to execute the first ISA binary code 1206 through emulation, simulation, or any other process. Figure 12 Virtual extension of the global address space, and system security
[0128]
[0129] The use of a system and software libraries addresses the need for a global address space on large-scale multi-core systems. The application user perception of the system memory space provides hooks for exploiting it and circumventing security solutions.
[0130] A Global Address Space (GAS) is a memory architecture for parallel and distributed computing systems, in which all memory locations across multiple processing nodes can be directly accessed by any processor in the system without explicit message passing or data replication. In other words, the memory of each processing node is combined into a single, unified address space, allowing processors to read and write data from any location (as if that location were local memory).
[0131] This shared memory model simplifies the programming of parallel and distributed systems by providing a more intuitive way to manage data access and communication between processing nodes. However, implementing a GAS architecture presents challenges (such as ensuring data consistency, managing memory latency, and maintaining security) because all processors may potentially access any memory location in the system.
[0132] Security in a GAS architecture is typically handled by software that partitions the address mapping between various user jobs, system, and kernel spaces. Hardware approaches are limited and exist more as support methods for implementing software security. An example is by designing regions in the address specifically for users, kernels, and the system to guide software usage. Additionally, some previous hardware solutions included minor protection checks but did not provide extensive protection to isolate memory regions between different user jobs.
[0133] Software security solutions impose a significant overhead on programmers and will affect the overall performance of the application. Additionally, software-only security exposes the hardware to various physical inspection methods. Limited hardware guidance leaves significant security vulnerabilities because it does not protect user code and data regions from the influence of other users (or kernel / system software) by itself.
[0134] Some embodiments of the present invention relate to a Virtual Extension for Global Address Space (VEGAS), and systems, methods, and devices for system security. In one or more embodiments, the VEGAS system can provide a mechanism for performing isolated job execution without disturbing standard software tools and models. The VEGAS system provides memory regions within the global address space to "jobs", and these memory regions are only visible / accessible to that job. The compute blocks associated with the job can access unencrypted data within the assigned memory region, while the unencrypted data remains encrypted to all other resources in the system. Although a portion of the global address space is visible to the compute logic, the security aspects are managed externally using VEGAS logic. The resources associated with the job do not have visibility to the physically mapped memory, which prevents side-channel attacks. The VEGAS system method provides a scalable security solution for systems implementing a Global Address Space (GAS). Previous hardware solutions for implementing memory region protection within the address mapping, or basic supervisor mode, only provide limited security and can be easily bypassed. The VEGAS system provides an end-to-end system solution that fully protects the memory range for the duration of the existence of the job.
[0135] The VEGAS system is a general-purpose computing architecture that can be applied to a wide range of applications from small devices to supercomputing systems. This disclosure focuses on a single-node office server setup that includes compute resources, memory (DRAM), and non-volatile memory (NVM) connected via industry-standard interfaces such as NVMe. However, this can be applied to other servers and computing arrangements.
[0136] Some of the components of the VEGAS system may include:
[0137] Compute resources: Represented as compute slices, these compute cores can be based on various architectures such as x86, ARM, or RISC-V.
[0138] Memory: This includes DRAM and NVM, which store and manage data within the system.
[0139] Network: Multiple network levels connect the resources and memory within the system, although the VEGAS architecture can also be applicable to only one level of the network.
[0140] Accelerators: Referred to as Memory Processing Units (MPUs), these specialized components provide specific acceleration for certain functions.
[0141] The VEGAS system allows users to access all resources, where security provisions are typically implemented in software. However, the VEGAS system introduces dedicated hardware support for improved security and isolation. When a user runs a job on the VEGAS system, resources such as compute slices, memory regions, and external connectivity are allocated by a centralized scheduler. The VEGAS system ensures that data is only visible to specific jobs and is protected from unauthorized access, even as network transactions pass through other resources.
[0142] The VEGAS system provides regionalization and isolation for jobs running on the platform, thus ensuring data security and efficient resource allocation.
[0143] In one or more embodiments, the VEGAS system may use metadata attached to the virtual address space to provide process isolation from all other processes, where address translation is handled outside of the compute block. It should be understood that a compute block is a term used to describe any computing unit or structure within the VEGAS system. It may refer to a compute slice, a Compute Complex Tile (CCT), or any other computing element within the system.
[0144] In one or more embodiments, the VEGAS system may facilitate that a compute block can access its own allocated virtual memory space without being exposed to the physical address mapping, thus preventing side-channel attacks.
[0145] In one or more embodiments, the VEGAS system may facilitate the VEGAS logic placed in a near-memory compute block to perform data decryption and encryption for atomic computing operations, as well as security checks for data handling.
[0146] In one or more embodiments, the VEGAS system may facilitate the VEGAS block at the network boundary to perform security checks, metadata decoding, address interleaving, metadata encryption and decryption, metadata compression and decompression, and resource access pattern detection.
[0147] In one or more embodiments, the VEGAS system may provide a scalable directory structure connected to a socket network for coherence in memory regions of interest, where the VEGAS logic within the directory structure is used to prevent unrelated jobs from accessing cache information.
[0148] The above description is for illustrative purposes and is not intended to be restrictive. There may be many other examples, configurations, processes, algorithms, etc., some of which are described in more detail below. Example embodiments will now be described with reference to the accompanying drawings.
[0149] Figure 13Depicts an illustrative schematic diagram for a VEGAS system in accordance with one or more example embodiments of the present disclosure.
[0150] Referring Figure 13 , a conceptual diagram of a system with VEGAS components is shown. A compute slice (1308) represents a programmable compute engine with the ability to boot an operating system and run processes. A memory processing unit (MPU, 1312) consists of a memory controller, atomicity compute support, and data manipulation operations. A compute complex slice (CCT, 1306) consists of multiple compute slices and an MPU connected by a local area network. A socket (1302) includes multiple CCTs (1306) connected using a socket network (1324), and these can communicate with other system components using CXL (1320 and / or 1322), PCIe, or similar interfaces. Multiple sockets can be connected using a global network (1326), which also has a processing unit connected to a core similar to Xeon or with security features. The VEGAS blocks (1330, 1334, 1336, and 1338) are responsible for performing: (a) virtual address translation; (b) job isolation by allowing jobs to map only to specific resources; (c) metadata field management (decryption, encryption, compression); (d) data decryption and encryption; and (e) access rule checking and security management.
[0151] A process initiated on a security system connected to the proposed VEGAS system or launched from a compute slice within the VEGAS system will generate a security task ID linked to the process. This task ID is securely delivered via an encrypted channel to the VEGAS blocks (1330, 1334, 1336, or 1338) to mark the address space and rules associated with the task ID protection.
[0152] In the VEGAS system, process isolation can be achieved by attaching metadata to the virtual address space of each process. This metadata is invisible to the compute slices, thus ensuring that each process remains separate from other processes. For example, consider two processes A and B running on the system. The VEGAS system assigns unique metadata to each process, thus ensuring that they do not inadvertently access each other's resources, thereby maintaining isolation and security.
[0153] Process isolation is an important aspect of the VEGAS system as it ensures that different processes running on the system do not interfere with each other or access each other's resources. To achieve this, the VEGAS system employs metadata attached to the virtual address space of each process. For example, for each process running on the system, the VEGAS system associates metadata with its virtual address space. This metadata contains information about the process, such as its permissions, privileges, and access restrictions. By attaching the metadata to the virtual address space, the VEGAS system can enforce isolation between processes and control access to resources.
[0154] The metadata associated with a process's virtual address space is invisible to the compute slices. Instead, the address translation that maps virtual addresses to physical addresses is handled outside of the compute blocks by the VEGAS logic. This ensures that the compute blocks cannot access or manipulate the metadata, access or manipulation of which could introduce potential security vulnerabilities.
[0155] The VEGAS system enforces process isolation by leveraging the metadata associated with each process's virtual address space. When a compute block requests access to a resource, the VEGAS logic checks the metadata to determine if the block is authorized to access the resource. If the compute block is not authorized, the VEGAS logic denies access, thereby preventing unauthorized interactions between processes. For example, assume there are two processes A and B running on the VEGAS system. Each process has its own metadata associated with its virtual address space, which specifies the resources it can access. When process A attempts to access a resource reserved for process B, the VEGAS logic checks the metadata and identifies that process A is not authorized to access the resource. Thus, the VEGAS logic denies the access request, effectively isolating the process and ensuring the integrity and security of the system.
[0156] By using the metadata attached to the virtual address space, the VEGAS system can effectively isolate the processes running on the system, thereby protecting sensitive data and resources from unauthorized access and potential security threats.
[0157] In the VEGAS system, the compute blocks can access their own allocated virtual memory space without being exposed to the physical address mapping. This design choice prevents side-channel attacks by minimizing the compute blocks' knowledge of the physical memory. For example, a compute block may be responsible for processing sensitive data. By ensuring that the block cannot access the physical memory mapping, the VEGAS system minimizes the risk of an attacker exploiting a side-channel vulnerability to access this sensitive information.
[0158] Each computational block in the VEGAS system is assigned its own virtual memory space. This virtual memory space is separated from other computational blocks, ensuring that each block operates independently and securely within its own allocated memory.
[0159] The VEGAS logic lies outside the boundaries of the computational blocks and is responsible for translating virtual addresses into physical addresses. By separating the VEGAS logic from the computational blocks, the system ensures that the computational blocks do not have direct access to the physical address mapping. This separation minimizes the risk of security vulnerabilities that could arise if the computational blocks were aware of the physical memory layout.
[0160] Before a computational block publishes data on the system network, the VEGAS logic handles address translation and virtual expansion. This process ensures that the computational block only deals with virtual addresses and remains unaware of the actual physical memory locations.
[0161] The lack of knowledge about the physical memory in the computational blocks helps prevent side-channel attacks. In a side-channel attack, an attacker can use knowledge of the physical memory layout to infer sensitive information about other processes running on the system. By ensuring that the computational blocks only have access to their own virtual memory spaces, the VEGAS system effectively mitigates the risk of such attacks.
[0162] For example, assume there are two computational blocks X and Y, each assigned its own virtual memory space. When computational block X needs to access a resource, it uses the virtual address associated with its own memory space. The VEGAS logic translates this virtual address into the corresponding physical address, allowing computational block X to access the resource without being exposed to the physical memory layout. If an attacker tries to use knowledge of computational block X to obtain information about the memory of computational block Y, they will not succeed because computational block X only has access to its own virtual memory space and is unaware of the physical memory layout.
[0163] By maintaining the separation between the computational blocks and the physical address mapping, the VEGAS system effectively protects the memory access process, which significantly reduces the risk of side-channel attacks and ensures the overall security of the system.
[0164] In one or more embodiments, the VEGAS block / logic (e.g., 1330) is strategically placed near the MPU in a near-memory computational block. Some of the functions of the VEGAS block may include:
[0165] a) Data decryption and encryption for atomic computing operations: Atomic computing operations are indivisible and non-interruptible tasks that must be executed in their entirety to ensure data consistency and accuracy. The VEGAS logic plays a crucial role in protecting these operations by handling the decryption and encryption of data. When a computing block receives encrypted data, the VEGAS logic decrypts the data before the atomic computing operation is executed. Once the operation is complete, the VEGAS logic re-encrypts the data before it is sent back to memory or shared with other computing blocks. This process ensures that sensitive data remains secure while being processed, as only authorized computing blocks can access and decrypt the data. For example, assume that the task of a computing block is to perform an atomic operation on encrypted financial data. The VEGAS logic decrypts the data so that the computing block can perform the operation. After the operation is complete, the VEGAS logic encrypts the data again, thus ensuring that the processed data remains secure.
[0166] b) Security checks for authenticated data handling: In addition to data encryption and decryption, the VEGAS logic also performs security checks to verify the legitimacy of data handling operations. These checks help ensure that only authorized and authenticated operations can access or manipulate data within the system.
[0167] By combining data decryption / encryption with security checks, the VEGAS logic adds a robust security layer to near-memory computing blocks, thus protecting sensitive data and ensuring that only authorized operations can access and manipulate the data.
[0168] The VEGAS block located at the network boundary acts as a security gatekeeper by performing several key functions. Its main purpose is to ensure that the data and metadata transmitted through the network are secure, reliable, and compliant with established rules. Some of the functions of the VEGAS block may include:
[0169] a) Security checks for network packets: The VEGAS block examines incoming and outgoing network packets to verify their authenticity. Packets that fail authentication are discarded (thrashed), and an acknowledgment is sent back to the sender. This process helps maintain the integrity of the system and prevents unauthorized access or data tampering.
[0170] b) Metadata decoding and rule checking: The VEGAS block decodes the metadata associated with network packets and checks whether the packets comply with predefined rules. This step ensures that only legitimate packets are allowed to pass through the system, thus minimizing the risk of malicious activities or data leaks.
[0171] c) Address interleaving programmed by the process: The VEGAS block is responsible for address interleaving, which rearranges the memory addresses according to a specific pattern programmed by the process. This function helps to optimize memory access, reduce latency, and improve overall system performance.
[0172] d) Metadata encryption and decryption for secure transmission over the network: The VEGAS block encrypts the metadata before it is transmitted over the network, thus ensuring that sensitive information is protected from eavesdropping or interception. Similarly, when receiving the metadata, the VEGAS block decrypts it so that the system can process and interpret the information.
[0173] e) Metadata compression and decompression: To reduce the amount of data transmitted over the network and improve efficiency, the VEGAS block compresses the metadata before sending it. After receiving the metadata, the VEGAS block decompresses the metadata so that the system can interpret and utilize the information.
[0174] f) Detecting resource access patterns and blocking transactions when the rules do not allow them: The VEGAS block monitors resource access patterns to identify and block any transactions that violate the established rules. This function helps to prevent side-channel attacks that may occur due to repeated access to system resources (hammering).
[0175] For example, assume that a computing block wants to send data to another computing block within the system. Before the data is transmitted, the VEGAS block checks the security of the packet, decodes and validates the metadata, and interleaves the addresses as programmed. It also encrypts and compresses the metadata to ensure a secure and efficient transmission. At the receiving end, the VEGAS block decrypts and decompresses the metadata, verifies that the packet complies with the rules, and monitors the resource access patterns to prevent potential side-channel attacks.
[0176] By performing these critical functions at the network boundary, the VEGAS block plays a key role in maintaining the security and integrity of the system, thus ensuring that data and metadata are securely transmitted and processed according to the established rules and protocols.
[0177] The VEGAS system employs a scalable directory structure connected to a socket network to enforce consistency for memory regions of interest. This structure incorporates VEGAS logic to prevent unrelated jobs identified by unique job IDs from accessing cache information. For example, if two jobs with different job IDs are running on the system, the VEGAS logic ensures that they cannot access each other's cache information, thus maintaining separation and security between unrelated jobs.
[0178] Memory consistency is important for maintaining data consistency across different memory locations when multiple computational blocks access and modify data. The directory structure is responsible for keeping track of the memory regions of interest and ensuring that the most up-to-date data is available to the computational blocks.
[0179] The VEGAS logic integrated within the directory structure adds an extra layer of security to the system by preventing unrelated jobs from accessing cache information. Each job running on the system is assigned a unique job ID, which is used as an identifier for that particular job. The VEGAS logic uses these job IDs to distinguish between jobs and enforce access control rules.
[0180] By monitoring and controlling access to cache information, the VEGAS logic helps prevent unauthorized access and potential security breaches. This security measure is particularly important in shared computing environments where multiple users or applications may be running simultaneously and sensitive data must be protected from unauthorized access. For example, assume there are two jobs running on the VEGAS system: Job-A and Job-B, each with its unique job ID. Job-A is accessing and modifying data in memory region X, while Job-B is working on memory region Y. The scalable directory structure keeps track of the memory regions of interest for each job and ensures data consistency across the system. The VEGAS logic within the directory structure checks the job ID associated with each request for cache information. If a request from Job-A attempts to access cache information related to memory region Y, the VEGAS logic rejects the request, thus preventing unauthorized access and maintaining data security.
[0181] It should be understood that the above description is for illustrative purposes and is not intended to be restrictive.
[0182] Reference Figure 14 shows an example use of VEGAS metadata fields for isolating and protecting data within a user job.
[0183] In one or more embodiments, the VEGAS system can utilize a global address space to manage memory resources, allowing different jobs to run on a system with their own unique physical address spaces. This ensures that jobs are isolated from each other and cannot access each other's memory. The global address space encompasses a wide range of memory, from zero to tens of petabytes or zettabytes, depending on the requirements of the system. Within this global address space, multiple jobs can run concurrently, each with its own allocated memory region.
[0184] For example, Job 1 may request a specific memory region that can be linearly mapped within the global address space. Similarly, Job 2 may request a different amount of memory, which can also be allocated within the global address space but may not be linearly distributed.
[0185] To ensure isolation between jobs, the VEGAS system uses VEGAS blocks placed at different granularities, such as interfaces for computing resources, network, and memory. These blocks authenticate and authorize transactions based on their associated memory spaces.
[0186] If Job 1 runs on a compute slice and issues a memory request within its allocated memory region, the VEGAS block will authorize the transaction, allowing it to pass through the network and access the required memory. However, if Job 1 issues a request outside its allocated memory region, the VEGAS block will reject the transaction, ensuring that jobs cannot access each other's memory spaces.
[0187] In the context of the VEGAS system, the translation of virtual addresses to system physical addresses occurs for the so-called job physical addresses. Although the VEGAS system itself is not directly involved in the translation, it plays an important role in handling address spaces. Jobs are only aware of their own physical address spaces, and the VEGAS blocks are responsible for translating job physical addresses to system physical addresses. The two address spaces defined in the system are physical addresses and system physical addresses. The compute resources running the jobs do not have visibility of the system physical address space, which makes the system secure. Jobs only have access to their respective physical addresses, which ensures that they remain unaware of the overall resources of the system.
[0188] The VEGAS system ensures job isolation by mapping each job to specific resources. This means that if a compute slice can access a certain region of memory, the system will provide isolation to prevent two jobs from accessing each other's resources.
[0189] An important aspect of the VEGAS system is the transmission of metadata that carries the necessary information for authentication. This metadata contains various attributes such as access rules, security levels, job IDs, partition IDs, and access modes, etc. Metadata is crucial for encoding transactions and ensuring secure data transfer.
[0190] Some of the functions in the VEGAS block's functionality are used to authenticate transactions and send them to their destinations (such as DRAM). Authentication can occur at the source or the destination. In either case, once the transaction is authenticated, the VEGAS block forms a packet with the necessary information (indicating which job the transaction belongs to), along with any applicable rules.
[0191] The data being transmitted is encrypted, which ensures that no unencrypted information is transmitted over the network. Metadata encryption is also possible, involving the encryption of job IDs, partition IDs, and other relevant information. Then, this encrypted metadata is transmitted over the network to the destination. The encryption process is necessary because the destination cannot retrieve and decrypt the data without knowing which specific job the data belongs to. This is particularly important for accelerators sitting next to the memory, which need to read the decrypted data, perform operations, and write the data back. Thus, the additional information provided by the metadata must be encapsulated and sent over the network.
[0192] For a system that includes multiple compute slices and memory components using scalable network connections, the VEGAS system provides a mechanism for performing isolated job execution without disturbing standard software tools and models. In essence, assuming a 256-bit true GAS implementation, each "job" in the machine is exposed to a 64-bit address space of all the data belonging to that job, regardless of how it is physically machine-owned. The other address bits [255:65] encode a larger address space, where the higher bits are metadata. It should be understood that this is only an example for illustrative purposes and is not intended to be restrictive. Other bits can be used. This limits each individual job to a maximum of approximately 4 exabytes (EB) of data for direct access via load / store / atomic operations. An example of how to encode this data at the full level is shown in Figure 2 is shown.
[0193] The metadata fields (all fields excluding the Figure 2 address field in ) encode information such as user identity and / or access control list (ACL) properties, media type, access interleaving granularity, security rules, encryption requirements, and isolation such as job keys. Through a secondary retranslation of the extended address space, this extended address can be local due to properties such as the interleaving or elasticity properties expressed in the access pattern. The metadata fields can be selected according to the requirements of the application. For example, if the data is not interleaved, the interleaving bits can be removed. The flexibility in including metadata fields helps to reduce the packet size and thus helps system performance.
[0194] A description of each of the proposed metadata examples is as follows:
[0195] The rule set represents the security level or rules for concurrent access to common data. This can be data shared between user jobs, debuggers, or profilers.
[0196] The global job ID represents the generated ID of the operating system associated with the job that issued the request. This is unique and shared among the resources dedicated to that job. This job ID becomes the "index" for a transparent global key table used for automatic encryption. When implemented in the system, data that has not yet been encrypted will only exist in the core slice (and is observed as plain text). While in the core slice, only resources assigned the same job ID can understand the content. Once the data leaves the core slice (and passes through the VEGAS encryption block), it will only be seen as ciphertext.
[0197] The partition ID is a representation of the nature of the potential composite memory types used to represent regions of memory (NVM, scratchpad, DRAM, etc.).
[0198] Interleaving represents the granularity of access, which can be dynamically adjusted based on the memory type. For example, NVM can be interleaved at a higher granularity compared to a scratchpad closer to the computing unit.
[0199] The decryption key index or decryption key can also be embedded as ciphertext, as part of the metadata.
[0200] It should be understood that the above description is for illustrative purposes and is not intended to be restrictive.
[0201] Figure 15 The figure is a flowchart illustrating an illustrative VEGAS system in accordance with one or more example embodiments of the present disclosure.
[0202] In one or more embodiments, the computing system discussed incorporates a VEGAS system that facilitates process isolation between at least two processes running on their respective computing blocks. This process isolation ensures that each job only has access to the resources specifically allocated to it, thus providing a secure computing environment.
[0203] At block 1502, the device may execute at least two processes within the device in a computing environment, each process running on a respective one of at least two computing blocks.
[0204] At block 1504, the device may employ a virtual extension of the global address space (VEGAS) system that includes VEGAS logic to manage the allocation of virtual memory spaces for at least two computing blocks, where the VEGAS logic is located outside the boundaries of the computing blocks and disposes of address translation and virtual extension before being published on the system network.
[0205] At block 1506, the device may isolate the virtual memory spaces of at least two processes by allowing each computing block to access only its own allocated virtual memory space.
[0206] In addition to process isolation, the computing system also includes computer-executable instructions that generate metadata for each process. This metadata is attached to the virtual address space and remains invisible to the computing blocks. Additionally, the processing circuitry is configured to protect the system network from side-channel attacks by ensuring that each computing block remains unaware of physical memory information.
[0207] The VEGAS system also manages a global address space (GAS) consisting of a number of address bits. Each job in the computing environment is exposed to the first address bits for all data belonging to that job. The other address bits within the GAS encode an extended address space that includes metadata fields such as user identification, access control list properties, media type, access interleaving granularity, security rules, and encryption requirements.
[0208] The VEGAS system includes components for performing various tasks, including virtual address translation, job isolation, metadata field management, data encryption and decryption, and access rule checking and security management. Additionally, the processing circuitry is configured to allow for flexible selection of metadata fields to help reduce packet size and improve overall system performance.
[0209] It should be understood that the foregoing description is for illustrative purposes and is not intended to be limiting.
[0210] Figure 16 An embodiment of an exemplary system 1600 in accordance with one or more example embodiments of the present disclosure is illustrated.
[0211] In various embodiments, the computing system 1600 may include an electronic device or may be implemented as part of an electronic device. The embodiments are not limited to this context. More generally, the computing system 1600 is configured to implement all of the logic, systems, processes, logical flows, methods, equations, apparatuses, and functions described herein.
[0212] System 1600 can be a computer system with multiple processor cores, such as a distributed computing system, a supercomputer, a high-performance computing system, a computing cluster, a mainframe computer, a minicomputer, a client-server system, a personal computer (PC), a workstation, a server, a portable computer, a laptop computer, a tablet computer, a handheld device such as a personal digital assistant (PDA), or other devices for processing, displaying, or transmitting information. Similar embodiments can include, for example, entertainment devices such as portable music players or portable video players, smartphones or other cellular phones, telephones, digital video cameras, digital still cameras, external storage devices, etc. Further embodiments implement larger-scale server configurations. In other embodiments, system 1600 can have a single processor with one core or more than one processor. Note that the term "processor" refers to a processor with a single core or a processor package with multiple processor cores.
[0213] Computing system 1600 is configured to implement all of the logic, systems, processes, logical flows, methods, apparatuses, and functions described herein with reference to the above figures.
[0214] As used in this application, the terms "system" and "component" and "module" are intended to refer to a computer-related entity, whether it is hardware, a combination of hardware and software, software, or software in execution, examples of which are provided by the exemplary system 1600. For example, a component can be, but is not limited to: a process running on a processor, a processor, a hard disk drive, multiple storage drives (of optical and / or magnetic storage media), an object, an executable, an execution thread, a program, and / or a computer.
[0215] By way of illustration, both an application running on a server and the server can be components. One or more components can reside within a process and / or an execution thread, and a component can be localized on one computer and / or distributed between two or more computers. Further, components can be communicatively coupled to each other via various types of communication media to coordinate operations. Coordination can involve one-way or two-way information exchange. For example, a component can transmit information in the form of signals transmitted via a communication medium. The information can be implemented as signals assigned to respective signal lines. In such an assignment, each message is a signal. However, further embodiments can alternatively employ data messages. Such data messages can be sent across various connections. Exemplary connections include parallel interfaces, serial interfaces, and bus interfaces.
[0216] As shown in the figure, system 1600 includes a motherboard 1605 for mounting platform components. The motherboard 1605 is a point-to-point interconnect platform that includes a processor 1610, a processor 1630 coupled via a point-to-point interconnect as an Ultra Path Interconnect (UPI), and a VEGAS device 1619. In other embodiments, system 1600 may be another bus architecture, such as a multi-point bus. Additionally, each of processors 1610 and 1630 may be a processor package having multiple processor cores. By way of example, processors 1610 and 1630 are shown as including (one or more) processor cores 1620 and 1640, respectively. Although system 1600 is an example of a two-socket (2S) platform, other embodiments may include more than two sockets or one socket. For example, some embodiments may include a four-socket (4S) platform or an eight-socket (8S) platform. Each socket is a mount for a processor and may have a socket identifier. Note that the term platform refers to a motherboard that has certain components, such as processors 1610 and chipset 1660, mounted thereon. Some platforms may include additional components, and some platforms may include only sockets for mounting processors and / or chipset.
[0217] Processors 1610 and 1630 may be any of a variety of commercially available processors, including but not limited to: Core (2) and processors; and processors; application, embedded, and security processors; and and processors; IBM and Cell processors; and similar processors. Dual microprocessors, multi-core processors, and other multi-processor architectures may also be employed as processors 1610 and 1630.
[0218] Processor 1610 includes an integrated memory controller (IMC) 1614, registers 1616, and point-to-point (P-P) interfaces 1618 and 1652. Similarly, processor 1630 includes IMC 1634, registers 1636, and P-P interfaces 1638 and 1654. IMCs 1614 and 1634 couple processors 1610 and 1630 to respective memories: memory 1612 and memory 1632. Memories 1612 and 1632 can be part of the main memory (e.g., dynamic random-access memory (DRAM)) of a platform for a synchronous DRAM (SDRAM) such as double data rate type 3 (DDR3) or type 4 (DDR4). In this embodiment, memories 1612 and 1632 are locally attached to respective processors 1610 and 1630.
[0219] In addition to processors 1610 and 1630, system 1600 may further include a VEGAS device 1619. The VEGAS device 1619 can be connected to chipset 1660 via P-P interfaces 1629 and 1669. The VEGAS device 1619 can also be connected to memory 1639. In some embodiments, the VEGAS device 1619 can be connected to at least one of processors 1610 and 1630. In other embodiments, memories 1612, 1632, and 1639 can be coupled to processors 1610 and 1630 and the VEGAS device 1619 via a bus and a shared memory hub.
[0220] System 1600 includes a chipset 1660 coupled to processors 1610 and 1630. Additionally, the chipset 1660 can be coupled to a storage medium 1603, for example, via an interface (I / F) 1666. The I / F 1666 can be, for example, Peripheral Component Interconnect-enhanced (PCI-e). Processors 1610, 1630, and the VEGAS device 1619 can access the storage medium 1603 through the chipset 1660.
[0221] The storage medium 1603 may include any non-transitory computer-readable storage medium or machine-readable storage medium, such as an optical storage medium, a magnetic storage medium, or a semiconductor storage medium. In various embodiments, the storage medium 1603 may include an article of manufacture. In some embodiments, the storage medium 1603 may store computer-executable instructions, such as those for implementing one or more of the processes or operations described herein (e.g., Figure 15 ) of the computer-executable instructions 1602. The storage medium 1603 may store computer-executable instructions for any of the above equations. The storage medium 1603 may also store computer-executable instructions for the models and / or networks described herein (such as neural networks, etc.). Examples of computer-readable storage media or machine-readable storage media may include any tangible medium capable of storing electronic data, including volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and so on. Examples of computer-executable instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, object-oriented code, visual code, and so on. It should be understood that the embodiments are not limited to this context.
[0222] The processor 1610 is coupled to the chipset 1660 via P-P interfaces 1652 and 1662, and the processor 1630 is coupled to the chipset 1660 via P-P interfaces 1654 and 1664. A Direct Media Interface (DMI) may couple the P-P interfaces 1652 and 1662, and the P-P interfaces 1654 and 1664, respectively. The DMI may be a high-speed interconnect that facilitates, for example, eight gigatransfers per second (GT / s), such as DMI 3.0. In other embodiments, the processors 1610 and 1630 may be interconnected via a bus.
[0223] The chipset 1660 may include a controller hub, such as a platform controller hub (PCH). The chipset 1660 may include a system clock for performing clock functions and include an interface for an I / O bus, such as a universal serial bus (USB), a peripheral component interconnect (PCI), a serial peripheral interconnect (SPI), an integrated interconnect (I2C), etc., to facilitate the connection of peripheral devices on the platform. In other embodiments, the chipset 1660 may include more than one controller hub, such as a chipset having a memory controller hub, a graphics controller hub, and an input / output (I / O) controller hub.
[0224] In this embodiment, the chipset 1660 is coupled to a trusted platform module (TPM) 1672 and a UEFI, BIOS, flash memory component 1674 via an interface (I / F) 1670. The TPM 1672 is a dedicated microcontroller designed to secure the hardware by integrating cryptographic keys into the device. The UEFI, BIOS, flash memory component 1674 may provide pre-boot code.
[0225] In addition, the chipset 1660 includes an I / F 1666 for coupling the chipset 1660 to a high-performance graphics engine, a graphics card 1665. In other embodiments, the system 1600 may include a flexible display interface (FDI) between the processors 1610 and 1630 and the chipset 1660. The FDI interconnects the graphics processor cores in the processors with the chipset 1660.
[0226] Various I / O devices 1692 are coupled to a bus 1681 together with a bus bridge 1680 that couples the bus 1681 to a second bus 1691 and an I / F 1668 that connects the bus 1681 to the chipset 1660. In one embodiment, the second bus 1691 may be a low pin count (LPC) bus. Various devices may be coupled to the second bus 1691, including, for example, a keyboard 1682, a mouse 1684, a communication device 1686, a storage medium 1601, and an audio I / O 1690.
[0227] An artificial intelligence (AI) accelerator 1667 can be a circuit arranged to perform AI-related computations. The AI accelerator 1667 can be connected to a storage medium 1603 and a chipset 1660. The AI accelerator 1667 can deliver the processing power and energy efficiency required to implement large-scale data computations. The AI accelerator 1667 is a type of specialized hardware accelerator or computer system designed to accelerate artificial intelligence and machine learning applications, including artificial neural networks and machine vision. The AI accelerator 1667 can be applicable to algorithms for robots, the Internet of Things, and other data-intensive and / or sensor-driven tasks.
[0228] Many of the I / O devices 1692, communication devices 1686, and storage media 1601 can reside on the motherboard 1605, while the keyboard 1682 and mouse 1684 can be attached peripheral devices. In other embodiments, some or all of the I / O devices 1692, communication devices 1686, and storage media 1601 are attached peripheral devices and do not reside on the motherboard 1605.
[0229] Certain examples can be described using the expressions "in one example" or "example" and their derivatives. These terms mean that the specific features, structures, or characteristics described with respect to the example are included in at least one example. The appearances of the phrase "in one example" in various places in the specification do not necessarily all refer to the same example.
[0230] Some examples can be described using the expressions "coupled" and "connected" and their derivatives. These terms are not necessarily intended as synonyms for each other. For example, a description using the terms "connected" and / or "coupled" can indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" can also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other.
[0231] Apparatus and method for fine-grained per-user secure access control
[0232] For both high availability and utility purposes, large-scale datasets (up to petabytes in size) are typically required to remain resident in memory. Such datasets must be available to a large number of potential concurrent users while enforcing that these users are only permitted to access or modify specific subsets of the entire dataset to enforce permission or privacy issues.
[0233] In current implementations, per-user access is controlled through software abstractions that define a formalism for representing data and APIs for accessing and manipulating that data. The software then enforces the permissions for accessing the data. Although these methods are effective, they are inherently vulnerable to software bugs and security vulnerabilities that abuse the permission checking mechanism.
[0234] Embodiments of the present invention include a fine-grained access control list (ACL) definition supported by hardware extensions and enforcement of memory access on a per-user basis. Some embodiments rely on a data structure that includes a per-job security index linked to user data blocks (e.g., data words, data bytes, etc.), which point to an array with several user fields, where each field indicates a per-user security control value.
[0235] As used herein, "resource" refers to any accessible computing or memory in the system. "Job" refers to a process or other program code entity identified by a unique job ID, which may be associated with a set of reserved resources. "User" refers to any entity that has an account on the system and can access job resources. "Data set" refers to data that resides in memory or storage and can be shared by multiple users with different access permissions.
[0236] Figure 17 An embodiment of the per-user authentication circuit / logic 1701 is illustrated, which may be implemented in hardware (e.g., integrated in the memory controller 1740 or integrated within the memory access circuit and directly coupled to the memory controller 1740), or implemented via a combination of hardware and software (e.g., firmware executed by a dedicated microcontroller or processor).
[0237] Figure 17 The memory access circuit 1700 in may be coupled to other system components and may operate at least in part as described above with respect to Figure 3B the memory access circuit 364 and / or the memory cell 370 in. Additionally, multiple cores (such as Figure 3B the core illustrated in or Figure 2 the cores 202A-N illustrated in) may execute software (e.g., a "job" as described below) on behalf of different users. Execution of the software may trigger various memory access requests, which are then processed by the per-user authentication circuit / logic 1701, as described in further detail below.
[0238] In one embodiment, the memory allocation details 1705 and security rules 1730 are indicated in the configuration write operation 1751, which is a privileged or supervisory operation available only to the system software or other trusted components. In one embodiment, only the root (trusted source) entity is permitted to write configuration data to the per-user authentication circuit / logic 1701, which may include a set of configuration registers (e.g., MSRs) for storing the memory allocation details 1705, security rules 1730, and other security metadata and security control data.
[0239] The memory requests 1752 of these embodiments include a user ID combined with a memory address. The address generation circuit 1707 generates a user data address (e.g., a physical address in the system memory 1750), a rule index address, and an ACL address based on the user ID and the provided memory address. In one embodiment, the access control list (ACL) and the corresponding rule index information are stored in the system memory 1750 and cached in the rule index cache 1735 to provide efficient access to the ACL data 1720 by the per-user access check circuit 1710. As described herein, the system memory 1750 may be partitioned into a region 1760 for user data, a region 1761 for rule index data, and a region 1762 for ACL data. The data memory region includes data sets protected with a given data granularity ranging from 1 to N bytes. The ACL region contains a corresponding entry per data granularity. The ACL granularity depends on the number of rules encoded for the use case and may range from 1 to M bytes. As described below, this field is referred to as the "security index".
[0240] In one embodiment, the rule index address provided by the address generation circuit 1707 is used to perform a lookup in the rule index cache 1735. In response to a cache miss, a request is sent to the memory controller 1740 to fetch that portion of the rule index and the corresponding ACL data (e.g., ACL data 1720), which is then stored in the cache 1735 and provided to the per-user access check circuit 1710. If available, the ACL data 1720 is provided directly from the rule index cache 1735.
[0241] The memory controller 1740 accesses the system memory 1750 based on the user data address provided by the address generation circuit 1707 and according to the metadata specified by the security rules 1730 (e.g., based on user / application permissions specified in the metadata). If the metadata indicates that the memory request 1752 can be satisfied, the data is returned to the per-user access check circuit 1710, which verifies the request using the ACL data 1720 before returning the (authenticated) data 1753 to the requester (e.g., an application executed by the user).
[0242] Figure 18 FIG. illustrates an embodiment of an ACL data structure 1801 that is used within the framework described herein to define access control lists and enforce a memory access permission scheme in hardware. As mentioned, the system memory 1750 can be partitioned into a user data region 1760 and an ACL region 1761 (containing rule indices and ACL data). The user data memory region 1760 includes data sets protected at a given "data granularity" ranging from 1 to N bytes. The ACL region 1761 includes a corresponding entry per data granularity. The ACL granularity depends on the number of rules that are required to be encoded for the use case and can range from 1 to M bytes.
[0243] For example, in Figure 18 , three 1-byte fields 1802A - 1802C are shown within the ACL data structure 1801, which specify the security indices to be applied in the security index table 1810. Each entry of the ACL list is a mask bit that encodes the access type for a particular rule and user. For example, Figure 18 each column of the security index table 1810 shown in is associated with a particular user (e.g., user-0 (User-0) to user-255 (User-255) in this example), and each row indicates the mask bits for a particular rule. Depending on the use case for which the embodiment of the present invention is implemented, the data and ACL granularities can be adjusted to achieve an acceptable trade-off in terms of the protection provided and the subsequent space overhead.
[0244] In the illustrated example, a data granularity of 8 bytes (words 1803A - 1803C) is used in combination with a 1-byte ACL granularity. This means that memory accesses can be protected at an 8-byte granularity and enforced by encoding 256 rules (via 1-byte ACLs). Thus, a given setup can encode 256 unique users, each with exclusive access to their respective data. Alternatively or additionally, 256 rules can be specified, where each rule is encoded on K bits, thus providing 2 KCombinations are provided for use. For example, a use case can define a "group" or "color" with access rights to data in certain memories, and a user executing the code is granted permission to access such data. In Figure 19 Memory overheads are shown for some of the different arrangements and user data access granularities.
[0245] Different embodiments of the present invention can provide storage of metadata inline with memory words across memory regions, as Figure 18 illustrated (e.g., where ACL data bytes 1802A - 1802C are interleaved with data words 1803A - 1803C), or provide storage of metadata as a contiguous storage block within the system memory range, as Figure 19 shown (e.g., where user data 1901 is stored in a first memory range and security metadata 1911 is stored in a second non - overlapping memory range).
[0246] In the above embodiments, each ACL mask bit is associated with a specific user, and the rule set encodes per - user access control, where 0 - access is not allowed and 1 - access is allowed. For example, in the Figure 18 security index table 1810 shown, each column of bits is associated with a specific user, and each individual bit within the column indicates per - user access control for the corresponding rule. To provide more fine - grained access permissions, some embodiments extend the mask encoding to a greater number of bits. For example, 2 - bit encoding can be used to provide four options: 00 - not accessible, 01 - read - only, 10 - write - only, 11 - read - write permitted. Different embodiments of the present invention can use any number of access control bits to provide a greater number of access permission options. Regardless of the number, the mask bits can be arranged in an array, where the number of entries in the array is equal to the number of users (e.g., the Figure 18 columns in).
[0247] Figure 20An example of a per - user system security model is shown, where system memory 2001 refers to any memory / storage device in the system, and network 2040 is an on - chip / off - chip network that provides access to system memory 2001 via authentication circuitry / logic 2010. In this embodiment, authentication circuitry / logic 2010 and ACL cache 2020 operate as security blocks to process per - user security checks as described herein in response to local memory access request 2050A and remote memory access request 2050B, respectively. As used herein, a "local" memory access request 2050A may be a request generated from an entity (e.g., a core, an accelerator, etc.) that is on the same chip or within the same package as the authentication logic / circuit, while a "remote" memory access request 2050B may be a request generated from outside the chip or package (e.g., from a different socket or a different system coupled via a network). In one embodiment, authentication circuitry / logic 2010 isolates memory 2001 from memory access requests 2050A - 2050B and performs a security check before allowing a memory access operation (e.g., a read / write operation). As mentioned, ACL cache 2020 coupled to authentication circuitry / logic 2010 accelerates security index table access.
[0248] In some embodiments, authentication circuitry / logic 2010 performs an initial security check in response to memory access requests 2050A - 2050B and allows data access at a finer granularity (i.e., using ACL data) only after the initial authentication has been performed. In these embodiments, authentication circuitry / logic 2010 and the corresponding ACL may be accessible only by trusted sources (e.g., a "root" source). Thus, access to the security index is initiated by authentication circuitry / logic 2010 and operates transparently to users and user applications.
[0249] The virtual extension of the global address space (VEGAS) described above provides a technique for isolating system resources at a per - process or per - job granularity. In a given process / job, multiple users may access data resources (e.g., data stores in a database or memory), and the described embodiments may provide per - user secure access to data resources at variable granularity levels. Figure 17-Figure 20 described embodiments may provide per - user secure access to data resources at variable granularity levels.
[0250] As described herein, fine-grained access control lists (ACLs) defined by hardware extensions and enforcement of memory access on a per-user basis allow data set management software to delegate permission and access control responsibilities to the hardware. This provides an added layer of protection against accessing data by exploiting software bugs and executing malicious code. These embodiments allow an authenticated entity that may be resident in the data set access system to bypass the software layer (similar to a query language) by directly accessing the machine memory within a job. These embodiments are also beneficial in governing access to data structures in memory that have a simple functional API or data structures in memory that can be directly addressed via load and store instructions. Thus, embodiments of the present invention can help improve the security of existing legacy code without requiring an additional security software layer, or allow programmers to develop new protected software, and can potentially improve the software development cycle. Additionally, these embodiments provide a data security solution in which multiple users can access a shared data set.
[0251] In Figure 21 FIG. illustrates a method for performing per-user security checks in accordance with some embodiments of the present invention. The method may be implemented within the context of the various architectures described herein, but is not limited to any particular system or processor architecture.
[0252] At 2101, a memory access request associated with a particular user and process is received, the memory access request including a unique user ID for identifying the user and a unique job ID for uniquely identifying the job / process. At 2102, the user ID, job ID, and destination address in the system memory are read from the memory access request.
[0253] At 2103, the memory access request is sent to the system memory based on the destination address. This may include a single request for embodiments in which the ACL metadata is stored inline with the user data (e.g., as shown in Figure 18 ), or if the ACL metadata is stored in a separate region of the system memory (e.g., as shown within the system memory 1750 in Figure 17 ), then this may include multiple requests.
[0254] At 2104, a per-data block (e.g., per-data word, per-data double word, per-data byte, etc.) portion of the security index table is obtained from the memory or cache according to the security metadata (e.g., user ID, job ID, destination address) tagged to the memory access request. For example, all security index metadata associated with a particular user can be obtained from the system memory and stored in the ACL cache.
[0255] At block 2105, per-data-block authentication is performed. For example, based on the user ID, job ID, and destination address, a relevant portion of the security index table 1810 can be indexed and read to determine whether a memory access request should be verified and passed. If it passes (i.e., if the relevant bit of the security index table indicates that the memory access is permitted), then at 2106, the memory access is allowed to complete (e.g., the requested data is sent to the source for a read operation, and the data is written to the memory for a write operation). If it does not pass, then an exception is raised at 2107 (indicating an invalid access), and no response (or ERROR message) is sent.
[0256] In existing implementations, a Page Table Entry (PTE) includes various access protection bits / fields that are marked to indicate the access level for the corresponding memory region (e.g., read-only, read-write, no-execute, etc. for each memory page). Some embodiments of the present invention store security metadata in one or more of these PTE fields to reduce the memory tax described above with respect to Figure 19 the memory tax described above. In some embodiments, additional fields are included in the PTE to store portions of the security metadata 1911.
[0257] As mentioned, in some embodiments, only the root source (trusted source) can read and write access control lists and / or other configuration data to the configuration registers and security index of the per-user authentication circuit / logic 1701 and has visibility to the metadata. Embodiments of the present invention provide a mode switch for selectively enabling or disabling per-user and / or per-word security controls. This mode switch can be specified to turn on or off some or all of the per-user authentication techniques described herein.
[0258] Example
[0259] The following are example implementations of different embodiments of the present invention.
[0260] Example 1. A processor includes: a plurality of cores for executing instructions associated with a plurality of jobs on behalf of a plurality of users to generate memory access requests; a memory access circuit for coupling at least one of the plurality of cores to a memory, the memory access circuit including: a per-user authentication circuit operable to perform an access check for a request to access a data block in the memory at a sub-page granularity, the request including a security index associated with the data block; the per-user authentication circuit for using the security index to identify a corresponding bit within an access control data structure to determine whether to provide access to the data block in response to the request.
[0261] Example 2. A processor as in Example 1, wherein the security index comprises or is based on a combination of a user ID code and a job ID code, the user ID code uniquely identifying the user among a plurality of users associated with the request, and the job ID code uniquely identifying the job among a plurality of jobs.
[0262] Example 3. A processor as in Example 1 or 2, wherein the memory access circuit further comprises: an address generation circuit for generating a user data address and an index address based on the request, the user data address for identifying the location of a data block in the memory, and the index address for identifying a corresponding portion of an access control data structure based on the security index.
[0263] Example 4. A processor as in any one of Examples 1-3, further comprising: an index cache for storing a corresponding portion of the access control data structure according to a cache policy.
[0264] Example 5. A processor as in any one of Examples 1-4, wherein the data block comprises a granularity of at least one of the following: byte, word, double word, or quad word.
[0265] Example 6. A processor as in any one of Examples 1-5, wherein the data block comprises one of a plurality of data blocks stored in the memory, wherein the plurality of data blocks are for association with corresponding plurality of security index values.
[0266] Example 7. A processor as in any one of Examples 1-6, wherein the plurality of security index values are for storage in a first region of the memory, the first region being separate from a second region of the memory in which the plurality of data blocks are stored.
[0267] Example 8. A processor as in any one of Examples 1-7, wherein each of the plurality of security index values is for storage in a region of the memory in which the data block is stored.
[0268] Example 9. A processor as in any one of Examples 1-8, wherein the access control data structure comprises an access control table, the access control table comprising a plurality of entries, each entry for storing a plurality of access control bits associated with a user among a plurality of users.
[0269] Example 10. A processor as in any one of Examples 1-9, wherein the per-user authentication circuit is for reading a plurality of access control bits associated with a first user among a plurality of users associated with a request for a data block to determine whether to provide the first user with access to the data block.
[0270] Example 11. A method includes: performing multiple jobs on behalf of multiple users; receiving a request to access a data block in a memory at a sub-page granularity, the request being associated with a job among the multiple jobs and the request including a security index; obtaining a portion of an access control data structure from the memory based on the security index; and identifying corresponding bits within the portion of the access control data structure to determine whether to provide access to the data block in response to the request.
[0271] Example 12. The method of Example 11, wherein the security index includes or is based on a combination of a user ID code and a job ID code, the user ID code uniquely identifying the user among the multiple users associated with the request, and the job ID code uniquely identifying the job among the multiple jobs.
[0272] Example 13. The method of Example 11 or 12, wherein the memory access circuit further includes: an address generation circuit configured to generate a user data address and an index address based on the request, the user data address being used to identify the location of the data block in the memory, and the index address being used to identify the corresponding portion of the access control data structure based on the security index.
[0273] Example 14. The method of any one of Examples 11-13, further comprising: caching the corresponding portion of the access control data structure in a cache according to a caching policy.
[0274] Example 15. The method of any one of Examples 11-14, wherein the data block includes at least one of the following granularities: byte, word, double word, or quad word.
[0275] Example 16. The method of any one of Examples 11-15, wherein the data block includes one data block among a plurality of data blocks stored in the memory, wherein the plurality of data blocks are used to be associated with corresponding plurality of security index values.
[0276] Example 17. The method of any one of Examples 11-16, wherein the plurality of security index values are used to be stored in a first region of the memory, the first region being separated from a second region of the memory in which the plurality of data blocks are stored.
[0277] Example 18. The method of any one of Examples 11-17, wherein each of the plurality of security index values is used to be stored in the region of the memory in which the data block is stored.
[0278] Example 19. The method of any one of Examples 11-18, wherein the access control data structure includes an access control table, the access control table including a plurality of entries, each entry being used to store a plurality of access control bits associated with a user among the multiple users.
[0279] Example 20. The method of any one of Examples 11-19, wherein a plurality of access control bits associated with a first user associated with a request among a plurality of users are used for a data block to be read to determine whether to provide the first user with access to the data block.
[0280] Example 21. A machine-readable medium having program code stored thereon that, when executed by a machine, causes the machine to perform operations including: performing a plurality of jobs on behalf of a plurality of users; receiving a request to access a data block in a memory at a sub-page granularity, the request being associated with a job among the plurality of jobs and the request including a security index; obtaining a portion of an access control data structure from the memory based on the security index; and identifying corresponding bits within the portion of the access control data structure to determine whether to provide access to the data block in response to the request.
[0281] Example 22. The machine-readable medium of Example 21, wherein the security index comprises or is based on a combination of a user ID code and a job ID code, the user ID code uniquely identifying the user among the plurality of users associated with the request, and the job ID code uniquely identifying the job among the plurality of jobs.
[0282] Example 23. The machine-readable medium of any one of Examples 21-22, wherein the memory access circuit further comprises: an address generation circuit for generating a user data address and an index address based on the request, the user data address for identifying the location of the data block in the memory, and the index address for identifying a corresponding portion of the access control data structure based on the security index.
[0283] Example 24. The machine-readable medium of any one of Examples 21-23, further comprising: caching a corresponding portion of the access control data structure in a cache according to a cache policy.
[0284] Example 25. The machine-readable medium of any one of Examples 21-24, wherein the data block comprises at least one of the following granularities: byte, word, double word, or quad word.
[0285] Example 26. The machine-readable medium of any one of Examples 21-25, wherein the data block comprises one of a plurality of data blocks stored in the memory, wherein the plurality of data blocks are used to be associated with corresponding plurality of security index values.
[0286] Example 27. The machine-readable medium of any one of Examples 21-26, wherein the plurality of security index values are used to be stored in a first region of the memory, the first region being separated from a second region of the memory in which the plurality of data blocks are stored.
[0287] Example 28. A machine-readable medium as in any one of Examples 21-26, wherein each of the plurality of security index values is for storage in a region of a memory in which data blocks are stored.
[0288] Example 29. A machine-readable medium as in any one of Examples 21-29, wherein the access control data structure includes an access control table that includes a plurality of entries, each entry for storing a plurality of access control bits associated with a user among a plurality of users.
[0289] Example 30. A machine-readable medium as in any one of Examples 21-29, wherein the plurality of access control bits associated with a first user among the plurality of users associated with a request are for being read for a data block to determine whether to provide the first user with access to the data block.
[0290] Some embodiments of the present invention implement an optimization that does not rely on a fixed ratio of security metadata bytes per application payload byte. Instead, these embodiments use a more flexible scheme, such as encoding the ratio in a page table, for automated disposal at a greater granularity. The ability to identify most of the code and data as generally accessible or only accessible by trusted sources is an optimization that can greatly reduce any potential overhead of the proposed scheme.
[0291] Embodiments of the present invention may include the steps described above. These steps may be embodied as machine-executable instructions that may be used to cause a general-purpose or special-purpose processor to execute these steps. Alternatively, these steps may be performed by specific hardware components that include hardwired logic for performing these steps, or by any combination of programmed computer components and custom hardware components.
[0292] As described herein, an instruction may refer to a specific configuration of hardware such as an application specific integrated circuit (ASIC) that is configured to perform certain operations or has a predefined function or software instructions stored in a memory embodied in a non-transitory computer-readable medium. Thus, the techniques illustrated in the figures may be implemented using code and data stored on and executed on one or more electronic devices (e.g., a terminal station, a network element, etc.). Such electronic devices use computer machine-readable media (internally and / or via a network with other electronic devices) to store and transmit the code and data, the computer machine-readable media such as non-transitory computer machine-readable storage media (e.g., magnetic disks; optical disks; random access memory; read-only memory; flash devices; phase change memory) and transitory computer machine-readable communication media (e.g., electrical, optical, acoustic, or other forms of propagated signals - such as, carrier waves, infrared signals, digital signals, etc.). Additionally, such electronic devices typically include a collection of one or more processors coupled to one or more other components such as one or more storage devices (non-transitory machine-readable storage media), user input / output devices (e.g., a keyboard, a touch screen, and / or a display), and network connections. The coupling of the collection of processors to the other components is typically through one or more buses and bridges (also known as bus controllers). The storage devices and signals carrying network traffic represent one or more machine-readable storage media and machine-readable communication media, respectively. Thus, the storage devices of a given electronic device typically store code and / or data for execution on the collection of one or more processors of that electronic device. Of course, different combinations of software, firmware, and / or hardware may be used to implement one or more portions of the embodiments of the present invention. Throughout this detailed description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that the present invention may be practiced without some of these specific details. In some instances, well-known structures and functions are not described in detail to avoid obscuring the subject matter of the present invention. Accordingly, the scope and spirit of the present invention should be determined according to the appended claims.
Claims
1. A processor, comprising: a plurality of cores for executing instructions associated with a plurality of jobs on behalf of a plurality of users to generate memory access requests; A memory access circuit, the memory access circuit being used to couple at least one core of the plurality of cores to a memory, the memory access circuit comprising: a per-user authentication circuit operable to perform an access check on a request to access a block of data in the memory at a sub-page granularity, the request including a security index associated with the block of data; The per-user authentication circuit is to use the security index to identify a corresponding bit within an access control data structure to determine whether to provide access to the data block in response to the request.
2. The processor of claim 1, wherein: The security index includes or is based on a combination of a user ID code that uniquely identifies a user of the plurality of users associated with the request and a job ID code that uniquely identifies a job of the plurality of jobs.
3. The processor according to claim 1 or 2, wherein: The memory access circuit further comprises: An address generation circuit is configured to generate a user data address and an index address based on the request, wherein the user data address is used to identify a location of the data block in the memory, and the index address is used to identify a corresponding portion of the access control data structure based on the security index.
4. The processor of claim 3, further comprising: An index cache is used to store the corresponding part of the access control data structure according to a cache policy.
5. The processor according to any one of claims 1 to 4, wherein: The data block comprises a granularity of at least one of: byte, word, doubleword, or quadword.
6. The processor of claim 5, wherein: The data block includes one data block among a plurality of data blocks stored in the memory, wherein the plurality of data blocks are used to be associated with a corresponding plurality of security index values.
7. The processor of claim 6, wherein: The plurality of security index values are for being stored in a first region of the memory separate from a second region of the memory in which the plurality of data blocks are stored.
8. The processor of claim 6, wherein: Each security index value of the plurality of security index values is configured to be stored in a region of the memory in which a data block is stored.
9. A processor as claimed in any one of claims 1 to 8, wherein: The access control data structure includes an access control table including a plurality of entries, each entry for storing a plurality of access control bits associated with a user of the plurality of users.
10. The processor of claim 9, wherein: The per-user authentication circuit is configured to read a plurality of access control bits associated with a first user of the plurality of users associated with a request for the data block to determine whether to provide the first user with access to the data block.
11. A method comprising: Execute multiple jobs on behalf of multiple users; receiving a request to access a data block in a memory at a sub-page granularity, the request being associated with a job from the plurality of jobs and the request including a security index; retrieving a portion of an access control data structure from the memory based on the security index; as well as A corresponding bit within the portion of the access control data structure is identified to determine whether to provide access to the data block in response to the request.
12. The method of claim 11, wherein: The security index includes or is based on a combination of a user ID code that uniquely identifies a user among the plurality of users associated with the request and a job ID code that uniquely identifies a job among the plurality of jobs.
13. The method according to claim 11 or 12, wherein: The memory access circuit further comprises: An address generation circuit is configured to generate a user data address and an index address based on the request, wherein the user data address is used to identify a location of the data block in the memory, and the index address is used to identify a corresponding portion of the access control data structure based on the security index.
14. The method of claim 13, further comprising: The corresponding portion of the access control data structure is cached in a cache according to a cache policy.
15. The method according to any one of claims 11 to 14, wherein: The data block comprises a granularity of at least one of: byte, word, doubleword, or quadword.
16. The method of claim 15, wherein: The data block includes one data block among a plurality of data blocks stored in the memory, wherein the plurality of data blocks are used to be associated with a corresponding plurality of security index values.
17. The method of claim 16, wherein: The plurality of security index values are for being stored in a first region of the memory separate from a second region of the memory in which the plurality of data blocks are stored.
18. The method of claim 16, wherein: Each security index value of the plurality of security index values is configured to be stored in a region of the memory in which a data block is stored.
19. The method according to any one of claims 11 to 18, wherein: The access control data structure includes an access control table including a plurality of entries, each entry for storing a plurality of access control bits associated with a user of the plurality of users.
20. A machine-readable medium having program code stored thereon, the program code, when executed by a machine, causing the machine to perform operations, the operations comprising: Execute multiple jobs on behalf of multiple users; receiving a request to access a data block in a memory at a sub-page granularity, the request being associated with a job from the plurality of jobs and the request including a security index; retrieving a portion of an access control data structure from the memory based on the security index; as well as A corresponding bit within the portion of the access control data structure is identified to determine whether to provide access to the data block in response to the request.