Data encryption / compression based on memory address translation

By selectively routing data to encryption or compression modules based on memory address translation attributes, the method addresses the challenge of securely managing encrypted and compressed data in multi-threaded and multi-core processor systems, enhancing security and performance.

DE102013200161B4Active Publication Date: 2025-05-08INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102013200161
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-01-23
Filing Date
2013-01-09
Publication Date
2025-05-08
Estimated Expiration
2033-01-09

AI Technical Summary

Technical Problem

Existing data processing systems face challenges in securely managing encrypted and compressed data, especially in multi-threaded and multi-core processor environments, where unauthorized access to unencrypted data in caches poses a significant risk.

Method used

The method involves selectively routing data to encryption or compression modules based on attributes stored in a memory address translation data structure, such as ERAT or TLB, to control encryption/decryption and compression/decompression processes during memory access requests.

Benefits of technology

This approach enhances data security by ensuring that secure data is only accessed and processed by authorized threads, reducing the risk of unauthorized access and improving system performance by minimizing the performance costs associated with data encryption and compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for accessing data in a data processing system, wherein the method comprises: in response to a memory access request initiated by a thread in a processing kernel (450), accessing a memory address translation data structure (456) to perform a memory address translation for the memory access request, wherein the processing kernel has a secure L1 cache (460) for storing encrypted data, another L1 cache (458) for storing unencrypted data, and an integrated encryption module (452); Accessing an encryption-related page attribute in the memory address translation data structure to determine whether the memory page associated with the memory access request is encrypted; and Fulfilling the memory access requirement by Selective routing of secure data on the storage side through the integrated encryption module to a decryption operation, depending on whether the storage side associated with the storage access request has been identified as encrypted and is stored in the secure L1 cache; and Using data on the memory page depends on whether the memory page associated with the memory access request is stored in the other L1 cache.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the invention

[0001] The invention relates generally to data processing and, more particularly, to the encryption and / or compression of data stored and / or used by data processing systems and processors. Background of the invention

[0002] The protection of secure data stored or used by the processors of a data processing system is of utmost importance in many data processing applications. Encryption algorithms are typically applied to secure data so that it is unintelligible without the application of a decryption algorithm, and secure data is typically stored in mass storage and other non-volatile storage media in encrypted format, which requires decryption before the secure data can be read and / or processed by a processor in a data processing system. However, the decryption of encrypted secure data often results in the secure data being stored in unencrypted form on various types of volatile storage in a data processing system, e.g.in main memory or in various levels of cache memory used to speed up access to frequently used data. However, whenever data is stored in an unsecured form in a data processing system's memory, the data may be exposed to unauthorized access, potentially compromising data confidentiality.

[0003] However, encrypting and decrypting data usually requires a certain amount of processing effort to process secure data, and in itself it is even desirable for applications to keep other, insecure data in a data processing system so that processing this other data does not involve the same processing effort associated with decryption and encryption.

[0004] Furthermore, as semiconductor technology gradually reaches practical limits in terms of increased clock speeds, architects are increasingly focusing on parallelism in processor architectures to improve performance. At the chip level, multiple processing cores are often arranged on the same chip, operating in much the same way as separate processor chips or, to some extent, as completely separate computers. Furthermore, parallelism is even applied within cores by using multiple execution units specifically designed to handle certain types of operations. Pipelining ("parallel processing of instructions") is also used in many cases, so that certain operations that may require multiple clock cycles to execute are divided into stages, allowing other operations to begin before previous ones complete.Multithreading is also used to enable parallel processing of multiple instruction streams, allowing more work to be performed overall in each clock cycle.

[0005] Because of this increased parallelism, the challenges associated with maintaining secure data in a data processing system are more significant than in previous systems without parallel computing. For example, in a data processing system containing a single processor with a single thread, secure data can be stored in encrypted form outside the processor and decrypted as needed by that single thread after the data has been loaded into the processor. However, if additional threads and even additional processing cores are located on the same processor chip, it may be necessary to restrict access to secure data to only certain threads or processing cores on the chip.For example, if multiple threads or processing cores share a shared cache, storing any secure data in unencrypted form in that cache may pose a risk that an unauthorized party could gain access to that data via a thread or processing core other than the one authorized to access the secure data. Furthermore, as modern processor designs grow into system-on-chip (SOC) systems with hundreds of processing cores on a single processor chip, protecting unencrypted data even from other processes on the same processor chip is becoming increasingly important.

[0006] Traditionally, encryption and decryption are performed by software running on a processor. However, encryption and decryption require significant processing power, and thus, dedicated hardware-based encryption modules have been developed that can perform the encryption / decryption of secure data faster and more efficiently than is typically the case with software, thereby reducing the processing overhead associated with these operations. Conventional encryption modules are usually located outside of a processing core, for example, between the processing core and a memory controller or otherwise connected to a memory bus located outside the processing cores.In order to make it easier to determine which data is encrypted, secure data can also be stored in specific memory address ranges, so that an encryption module can be controlled by filtering so that only data stored in the identified ranges of memory addresses is encrypted or decrypted.

[0007] However, such an architecture can result in unencrypted data being cached and accessible by other threads and / or processing cores on a chip. Furthermore, a memory controller located outside a processor chip typically needs to set up and manage memory address ranges where secure data is stored, allowing encryption to be selectively enabled for memory transactions containing secure data, resulting in inefficient throughput for secure data.

[0008] The challenges associated with data compression are similar. Data compression can be used to reduce the amount of memory required to store data; however, compressed data must be decompressed before it can be used by a processing core or thread. Compressing and decompressing data involves processing overhead, and implementing such functions in software often results in a performance penalty. In addition, specific compression modules are also being developed to reduce the processing overhead associated with compressing and decompressing data; however, such modules are usually located outside of a processor core, e.g.within a memory controller, and consequently, compressed data may be required to be stored in decompressed format across different cache levels, reducing the storage space for other data in such caches and thus reducing memory system performance. Furthermore, as with encrypted data, a memory controller may need to establish and manage ranges of memory addresses in which to store compressed data so that a compression engine can be selectively enabled for memory transactions with compressed data, resulting in inefficient throughput for compressed data.

[0009] For this reason, there remains a pressing need in the field for a method to minimize the performance overhead associated with accessing and managing encrypted and / or compressed data in a computing system, as well as for providing additional protection for encrypted data within a multi-threaded and / or multi-core processor chip and a computing system incorporating it.

[0010] The document US 2006 / 0 059 553 A1 relates to a system for storing data in a memory, comprising: a CPU; an encryption circuit in communication with the CPU; a memory that communicates with the encryption circuit; and an address bus having a plurality of address lines, wherein a value of at least one address line determines a key selected from a plurality of keys to be used in the encryption circuit to encrypt data transferred from the CPU to the memory.

[0011] The document US 2008 / 0 077 922 A1 relates to methods for providing multi-level memory protection, the method comprising: defining a hierarchy of one or more predecessor processes and their respective successor processes; and establishing a data structure for defining access rights for each of the predecessor processes and their respective successor processes in the defined hierarchy. Summary of the invention

[0012] The invention relates to a method for accessing data in a data processing system, a circuit arrangement, and a program product, the features of which are specified in the corresponding patent claims. The embodiments are specified in the dependent patent claims.

[0013] The invention addresses these and other problems in the prior art by providing a method and circuitry that selectively routes data to an encryption or compression module based on encryption- and / or compression-related page attributes stored in a memory address translation data structure, such as an Effective To Real Translation (ERAT) or Translation Lookaside Buffer (TLB). For example, a memory address translation data structure may be accessed in conjunction with a memory access request for data on a memory page, so that attributes associated with the memory page in the data structure can be used to control whether data is encrypted / decrypted and / or compressed / decompressed as part of processing the memory access request.

[0014] Therefore, according to one aspect, in response to a memory access request initiated by a thread in a processing core, a memory address translation data structure is accessed to perform a memory address translation for the memory access request. Furthermore, an encryption-related page attribute in the memory address translation data structure is accessed to determine whether the memory page associated with the memory access request is encrypted. Thereafter, secure data on the memory page is selectively passed through an encryption module to subject the secure data to an encryption operation, depending on whether the memory page associated with the memory access request is determined to be encrypted.

[0015] According to another aspect, in response to a memory access request initiated by a thread in a processing core, a memory address translation data structure is accessed to perform a memory address translation for the memory access request. Furthermore, a compression-related page attribute in the memory address translation data structure is accessed to determine whether the memory page associated with the memory access request is compressed. Thereafter, data on the memory page is selectively passed through a compression module to subject the secure data to a compression operation, depending on whether the memory page associated with the memory access request is determined to be compressed.

[0016] These and other advantages and features which characterize the invention are set forth in the appended claims and form a further part hereof. To better understand the invention and the advantages and objects attendant upon its use, reference should be made to the drawings and the accompanying description, in which exemplary embodiments of the invention are described. Short description of the drawings Fig. 1 is a block diagram of an exemplary automated computing system including an exemplary computer useful in data processing in accordance with embodiments of the present invention. Fig. Figure 2 is a block diagram of an exemplary NOC implemented in the computer of Fig. 1 is implemented. Fig. Figure 3 is a block diagram showing an exemplary implementation of a node of the NOC of Fig. 2 shows in more detail. Fig. Figure 4 is a block diagram showing an example implementation of an IP block of the NOC of Fig. 2 shows. Fig. 5 is a block diagram of an exemplary data processing system illustrating data encryption / decryption based on memory address translation in accordance with the invention. Fig. Figure 6 is a block diagram of an example ERAT input format for the ERAT of Fig. 5. Fig. 7 is a block diagram illustrating exemplary memory access using a data processing system that supports data encryption / decryption based on memory address translation in accordance with the invention. Fig. 8 is a flow chart showing an example processing sequence for accessing data in the data processing system of Fig. 7 shows. Fig. 9 is a flowchart illustrating an exemplary processing sequence for performing a read bus transaction in the data processing system of Fig. 7 shows. Fig. 10 is a flowchart illustrating an exemplary processing sequence for performing a write bus transaction in the data processing system of Fig. 7 shows. Fig. 11 is a flowchart illustrating an alternative processing sequence for performing a read bus transaction in the data processing system of Fig. 7 in conjunction with layer-selective encryption / compression. Fig. 12 is a flowchart illustrating an alternative processing sequence for performing a write bus transaction in the data processing system of Fig. 7 in conjunction with layer-selective encryption / compression. Fig. 13 is a block diagram illustrating exemplary memory access using another data processing system including an integrated encryption module according to the invention. Fig. 14 is a block diagram showing another data processing system including an integrated encryption module according to the invention and further including a separate secure cache memory for use in storing secure data. Detailed description

[0017] Embodiments according to the invention selectively route data to an encryption or compression module based on encryption- and / or compression-related page attributes stored in a memory address translation data structure, such as an Effective To Real Translation (ERAT) or Translation To Read Buffer (TLB). For example, a memory address translation data structure may be accessed in conjunction with a memory access request for data on a memory page, so that attributes associated with the memory page in the data structure may be used to control whether data is encrypted / decrypted and / or compressed / decompressed as part of processing the memory access request.

[0018] Furthermore, in some embodiments according to the invention, an integrated encryption module within a processing core of a multi-core processor may be used to perform encryption operations—i.e., encryption and decryption of secure data—in conjunction with memory access requests accessing such data. Combined with a memory address translation data structure, such as an Effective To Real Translation (ERAT) or a Translation To Real Buffer (TLB) that has been augmented with encryption-related page attributes to indicate whether memory pages identified in the data structure are encrypted, secure data associated with a memory access request in the processing core may be selectively directed to the integrated encryption module based on the encryption-related page attribute for the memory page associated with the memory access request.Furthermore, in some embodiments, an integrated encryption module may be coupled to an L1 cache within the processing core, which is considered effectively secure given that it stores secure data. The L1 cache may be configured to store both secure and insecure data, or alternatively, the L1 cache may be dedicated to secure data, and a second, insecure L1 cache may be used to cache insecure data. In either case, secure data can only be decrypted within the processing core and encrypted whenever the secure data is outside the processing core, providing increased security compared to conventional designs.

[0019] Other variations and modifications will be apparent to those skilled in the art. Therefore, the invention is not limited to the specific embodiments discussed herein. Hardware and software environment

[0020] Referring to the drawings, in which like reference numerals designate like parts throughout the several views, Fig. 1 illustrates an exemplary automated computing system including an exemplary computer 10 useful in data processing according to embodiments of the present invention. The computer 10 of Fig. 1 includes at least one computer processor 12 or “CPU” and a random access memory 14 (RAM) connected to a processor 12 and other components of the computer 10 via a high-speed memory bus 16 and a bus adapter 18.

[0021] An application program 20, a module of user-level computer program instructions for performing specific data processing tasks, such as word processing, spreadsheets, database operations, video games, securities market simulations, simulations of atomic quantum processes, or other user-level applications, is stored in RAM 14. An operating system 22 is also stored in RAM 14. Operating systems useful in connection with embodiments of the invention include UNIX™, Linux™, Microsoft Windows XP™, AIX™, IBM i5 / OS™, and others, as will be apparent to those skilled in the art. The operating system 22 and the application 20 in the example of Fig. 1 are shown in RAM 14, but many components of the software are usually also stored in non-volatile memory, e.g., on a disk drive 24.

[0022] As will become more apparent below, embodiments according to the invention may be implemented within network-on-chip (NOC) integrated circuit units or chips, and as such, computer 10 is shown as including two exemplary NOCs: a video adapter 26 and a coprocessor 28. The NOC video adapter 26, which may alternatively be referred to as a graphics card, is an example of an I / O adapter specifically designed for providing graphics output to a display device 30, such as a display screen or a computer monitor. The NOC video adapter 26 is connected to processor 12 via a high-speed video bus 32, bus adapter 18, and front-side bus 34, which is also a high-speed bus. The NOC coprocessor 28 is connected to the processor 12 via the bus adapter 18 and the front-side buses 34 and 36, which is also a high-speed bus.The NOC processor from . Fig. 1 can be optimized, for example to accelerate certain data processing tasks at the request of the main processor 12.

[0023] The exemplary NOC video adapter 26 and the NOC coprocessor 28 of Fig. 1 each include a NOC comprising integrated processor (IP) blocks, routers, memory data transfer controllers and network interface controllers, with reference to the details in connection with the Fig. 2 to 3 will be discussed in more detail below. The NOC video adapter and the NOC coprocessor are each optimized for programs that utilize parallel processing and also require fast random access to shared memory. Those skilled in the art will understand that, in light of the present disclosure, the invention may be implemented in other devices and device architectures besides the NOC devices and NOC device architectures. Thus, the invention is not limited to implementation within a NOC device.

[0024] The Computer 10 of Fig. 1 includes a drive adapter 38 connected to the processor 12 and other components of the computer 10 via an expansion bus 40 and the bus adapter 18. The drive adapter 38 connects non-volatile data storage to the computer 10 in the form of a drive 24 and may be implemented, for example, using Integrated Drive Electronics (IDE) adapters, Small Computer System Interface (SCSI) adapters, and others, as will be apparent to one skilled in the art. Non-volatile computer memory may also be implemented as an optical drive, electronically erasable programmable read-only memory (so-called EPROM or flash memory), RAM drives, and the like, as will be apparent to one skilled in the art.

[0025] Computer 10 further includes one or more input / output (I / O) adapters 42 that implement user-oriented input / output, for example, via software drivers and computer hardware for controlling output to display devices, such as computer display screens, as well as user input through user input devices 44, such as a keyboard and mouse. Furthermore, computer 10 includes a communications adapter 46 for communicating with other computers 48 and for communicating with a communications network 50. Such communications may be performed serially via RS-232 connections, via external buses such as Universal Serial Bus (USB), via communications networks such as IP communications networks, and in other ways as will be apparent to those skilled in the art.Communication adapters implement the hardware layer of data transmission, through which one computer sends data transmissions to another computer directly or over a data transmission network. Examples of data transmission adapters suitable for use with computer 10 include modems for wired dial-up data transmission, Ethernet (IEEE 802.3) adapters for wired data transmission network data transmission, and 802.11 adapters for wireless data transmission network data transmission.

[0026] For further explanation, Fig. 2 is a functional block diagram of an exemplary NOC 102 according to embodiments of the present invention. The NOC of Fig. 2 is implemented on a "chip" 100, i.e., on an integrated circuit. The NOC 102 includes integrated processor (IP) blocks 104, routers 110, memory data transfer controllers 106, and network interface controllers 108, which are grouped into interconnected nodes. Each IP block 104 is connected to a router 110 via a memory data transfer controller 106 and a network interface controller 108. Each memory data transfer controller controls the data transfer between an IP block and the memory, and each network interface controller 108 controls the data transfer between the IP blocks via the routers 110.

[0027] In the NOC 102, each IP block represents a reusable unit of synchronous or asynchronous logical design used as a building block for data processing within the NOC. The term "IP block" is sometimes expanded to include "intellectual property block," where an IP block effectively refers to a design owned by a party that is the intellectual property of a party to be licensed to other users or developers of semiconductor circuits. However, within the scope of the present invention, there is no requirement that IP blocks be subject to any particular ownership, so the term throughout this specification should always be expanded to include "integrated processor block." As defined herein, IP blocks are reusable units of logic, cell, or chip layout design that may be subject to intellectual property. IP blocks are logical cores that may be formed as ASIC chip designs or FPGA logic designs.

[0028] One way to describe IP blocks using an analogy is that IP blocks are to a NOC design what a library is to computer programming or a standalone integrated circuit component is to a printed circuit board design. In NOCs according to embodiments of the present invention, IP blocks may be implemented as generic gate netlists, as complete, specific or general-purpose microprocessors, or in other ways that may be apparent to one of ordinary skill in the art. A netlist is a Boolean algebraic representation (gates, standard cells) of a logical function of an IP block, analogous to a list of assembly code for a high-level application. Furthermore, NOCs may be implemented, for example, in a synthesizable form described in a hardware description language such as Verilog or VHDL.In addition to netlist and synthesizable implementations, NOCs can also be provided in hardware-level physical descriptions. Analog IP block elements such as SERDES, PLL, DAC, ADC, etc., can be distributed in a transistor layout format such as GDSII. Digital elements of IP blocks are sometimes also offered in layout format. It should also be understood that IP blocks, as well as other logic circuits implemented according to the invention, can be distributed in the form of computer data files, e.g., logic definition program code, that define the functionality and / or layout of the circuitry implementing such logic at various levels of detail. Even if the invention is described in the context of circuitry,implemented in fully functional integrated circuit units, data processing systems employing such units, and other tangible physical hardware circuits, those skilled in the art will thus understand, in light of the present disclosure, that the invention may also be implemented within a program product, and that the invention applies equally regardless of the particular type of computer-readable storage medium used to distribute the program product. Examples of computer-readable storage media include, but are not limited to, physical writable media such as volatile and non-volatile memory devices, floppy disks, hard disk drives, CD-ROMs, and DVDs.

[0029] Each IP block 104 in the example of Fig. 2 is connected to a router 110 via a memory data transfer controller 106. Each memory data transfer controller is a combination of synchronous and asynchronous logic circuitry designed to provide data transfer between an IP block and memory. Examples of such data transfers between IP blocks and memory are load instructions for memory and store instructions for memory. The memory data transfer controllers 106 are described below with reference to Fig. 3. Each IP block 104 is also connected to a router 110 via a network interface controller 108, which controls the data transmission via the routers 110 between the IP blocks 104. Examples of data transmissions between IP blocks are messages containing data and instructions for processing the data among IP blocks in parallel applications and within the context of applications processed in a pipelined system. The network interface controllers 108 are also described below with reference to Fig. 3. The routers 110 and the corresponding links 118 between them implement the network operations of the NOC. The links 118 may be packet structures implemented on physical, parallel wired buses connecting all the routers. This means that each link may be implemented on a wired bus large enough to simultaneously accommodate an entire data switching packet, including all header information and payload. For example, if a packet structure contains 64 bytes, including an 8-byte header and 56 bytes of payload, the wired bus facing each link will be 64 bytes in size, 512 wires. Furthermore, each link may be bidirectional, so that if the link packet structure contains 64 bytes, the wired bus will actually contain 1024 wires between each router and each of its neighbors in the network.In such an implementation, a message could contain more than one packet, but each packet would fit exactly within the size of the wired bus. Alternatively, a connection may be implemented on a wired bus large enough to accommodate only part of a packet, so that a packet must be split into multiple clock cycles. For example, if a connection is implemented with a size of 16 bytes or 128 wires, a 64-byte packet can be split into four clock cycles. It should be understood that different implementations may use different bus sizes due to practical physical constraints as well as desired performance characteristics.If the connection between the router and each section of the wired bus is called a port, each router has five ports, one for each of the four data transmission directions in the network and a fifth port to connect the router to a specific IP block via a memory data transfer controller and a network interface controller.

[0030] Each memory data transfer controller 106 controls the data transfer between an IP block and the memory. The memory may include an off-chip main RAM 112, an off-chip memory 114 directly connected to an IP block via a memory data transfer controller 106, an on-chip memory enabled as an IP block 116, and on-chip cache memory. In the NOC 102, for example, each of the on-chip memories 114, 116 may be implemented as an on-chip cache memory. All of these memory forms may be located in the same address space and at the same physical or virtual addresses, even for the memory directly connected to an IP block. For this reason, memory-addressed messages can be fully bidirectional with respect to IP blocks, since such memory can be directly addressed by an IP block anywhere on the network.Memory 116 on an IP block can be addressed by that IP block or another IP block in the NOC. Memory 114 directly connected to a memory data transfer controller can be addressed by the IP block connected to the network through that memory data transfer controller—and can also be addressed by another IP block anywhere in the NOC.

[0031] The NOC 102 includes two memory management units (MMUs) 120, 122, illustrating two alternative memory architectures for NOCs according to embodiments of the present invention. The MMU 120 is implemented within an IP block, allowing a processor within the IP block to operate in virtual memory while allowing the rest of the NOC's architecture to operate in a physical memory address space. The MMU 122 is implemented off-chip and connected to the NOC via a data link port 124. Port 124 includes the contacts and other connections necessary to route signals between the NOC and the MMU, as well as sufficient intelligence to convert message packets from the NOC packet format to the bus format required by the external MMU 122.The external location of the MMU means that all processors in all IP blocks of the NOC can operate in a virtual memory address space, with all conversions to physical addresses of off-chip memory handled by the off-chip MMU 122.

[0032] In addition to the two memory architectures illustrated through the use of MMUs 120, 122, data transfer port 126 shows a third memory architecture useful in NOCs that may be used in embodiments of the present disclosure. Port 126 provides a direct connection between an IP block 104 of the NOC 102 and off-chip memory 112. This architecture does not have an MMU in the processing path and instead provides for the use of one physical address space by all IP blocks of the NOC. With bidirectional address space sharing, all IP blocks of the NOC can access memory in the address space using memory-directed messages, including loads and stores, routed through the IP block directly connected to port 126.The connector 126 includes the contacts and other connections necessary to conduct signals between the NOC and the off-chip memory 112, as well as sufficient intelligence to convert message packets from the NOC packet format to the bus format required by the off-chip memory 112.

[0033] In the example of Fig. 2, one of the IP blocks is designated as a host interface processor 128. A host interface processor 128 provides an interface between the NOC and a host computer 10 in which the NOC may be installed, and also provides data processing services for the other IP blocks in the NOC, such as receiving NOC data processing requests from the host computer and distributing them among the IP blocks. For example, an NOC may implement a video graphics card 26 or a coprocessor 28 on a larger computer 10, as described above with reference to Fig. 1. In the example of Fig. 2, the host interface processor 128 is connected to the larger host computer via a data link port 130. Port 130 includes the contacts and other connections necessary to conduct signals between the NOC and the host computer, as well as sufficient intelligence to convert message packets from the NOC format to the bus format required by the host computer 10. In the example of the NOC coprocessor in the computer of Fig. 1, such a connector provides data transmission format translation between the interconnect structure of the NOC coprocessor 28 and the protocol required for the front-side bus 36 between the NOC coprocessor 28 and the bus adapter 18.

[0034] Fig. 3 then shows a functional block diagram illustrating in more detail the components implemented within an IP block 104, a memory data transfer controller 106, a network interface controller 108, and a router 110 in the NOC 102, all shown at 132. The IP block 104 includes a computer processor 134 and I / O functionality 136. In this example, the computer memory is represented by a segment of random access memory (RAM) 138 in the IP block 104. As described with reference to Fig. 2, the memory may occupy segments of a physical address space whose contents on each IP block are addressable or accessible by each IP block in the NOC. The processors 134, the I / O functionality 136, and the memory 138 in each IP block effectively implement the IP blocks as general-purpose programmable microcomputers. However, as explained above, IP blocks within the scope of the present invention generally provide reusable units of synchronous or asynchronous logic used as building blocks for data processing within a NOC. Therefore, implementing IP blocks as general-purpose programmable microcomputers is not a limitation of the present invention, although a typical embodiment is useful for illustrative purposes.

[0035] At NOC 102 of Fig. 3, each memory transfer controller 106 includes a plurality of memory transfer execution modules 140. Each memory transfer execution module 140 is designed to execute memory transfer instructions from an IP block 104, for example, a bidirectional memory transfer instruction stream 141, 142, 144 between the network and the IP block 104. The memory transfer instructions executed by the memory transfer controller may originate not only from the IP block connected to a router via a particular memory transfer controller, but also from any IP block 104 located anywhere in the NOC 102.This means that each IP block in the NOC can generate a memory transfer instruction and transmit it via the NOC's routers to another memory transfer controller associated with another IP block to execute that memory transfer instruction. Such memory transfer instructions can include, for example, address translation buffer control instructions, cache control instructions, barrier instructions, and memory load and store instructions.

[0036] Each memory transfer execution module 140 is designed to execute a complete memory transfer instruction separately from and in parallel with other memory transfer execution modules. The memory transfer execution modules implement a scalable memory transaction processor optimized for concurrent throughput of memory transfer instructions. The memory transfer controller 106 supports multiple memory transfer execution modules 140, all executing concurrently, to execute multiple memory transfer instructions concurrently. A new memory transfer instruction is assigned to a memory transfer module 140 by the memory transfer controller 106, and the memory transfer execution modules 140 can accept multiple response events concurrently.In this example, all memory transfer execution modules 140 are identical. Scaling the number of memory transfer instructions that can be processed concurrently by a memory transfer controller 106 is therefore implemented by scaling the number of memory transfer execution modules 140.

[0037] At NOC 102 of Fig. 3, each network interface controller 108 is designed to convert data transfer instructions from command format to a network packet format for transmission between IP blocks 104 via routers 110. The data transfer instructions may be formulated in command format by IP block 104 or by memory data transfer controller 106 and provided to network interface controller 108 in command format. The command format may be a native format corresponding to the architectural register file of IP block 104 and memory data transfer controller 106. The network packet format is typically the format required for transmission across routers 110 of the network. Each of these messages is composed of one or more network packets.Examples of such data transfer instructions that are converted from command format to packet format in the network interface controller include memory load instructions and memory store instructions between IP blocks and memory. Such data transfer instructions may also include data transfer instructions that send messages between IP blocks containing data and instructions for processing the data between IP blocks in parallel applications and within pipelined applications.

[0038] At NOC 102 of Fig. 3, each IP block is designed to send memory address-based data transfers to and from memory via the IP block's memory data transfer controller, and then to the network via its network interface controller. A memory address-based data transfer is a memory access instruction, such as a load instruction or a store instruction, executed by a memory data transfer execution module of an IP block's memory data transfer controller. Such memory address-based data transfers typically originate in an IP block, are formulated in instruction format, and are passed to a memory data transfer controller for execution.

[0039] Many memory address-based data transfers are performed using message traffic because any memory to be accessed may be located anywhere in the physical memory address space, on or off-chip, directly connected to a memory data transfer controller in the NOC, or ultimately accessed through an IP block of the NOC—regardless of which IP block the particular memory address-based data transfer originated from. Thus, at the NOC 102, all memory address-based data transfers performed using message traffic are forwarded from the memory data transfer controller to the associated network interface controller in a message for conversion from instruction format to packet format and transmission over the network.When converting to packet format, the network interface controller further identifies a network address for the packet depending on the one or more memory addresses accessed by a memory address-based data transfer. Memory address-based messages are addressed with memory addresses. Each memory address is mapped by the network interface controllers to a network address, typically the network location of a memory data transfer controller responsible for a certain range of physical memory addresses. The network location of a memory data transfer controller 106 is naturally also the network location of that memory data transfer controller's associated router 110, the network interface controller 108, and the IP block 104.The instruction conversion logic 150 within each network interface controller is capable of converting memory addresses into network addresses for the purpose of transmitting memory address-based data transfers across routers of an NOC.

[0040] Upon receiving message traffic from the network's routers 110, each network interface controller 108 examines each packet for store instructions. Each packet containing a store instruction is forwarded to the store data transfer controller 106 associated with the receiving network interface controller, which executes the store instruction before sending the remaining packet payload to the IP block for further processing. In this way, memory contents are always prepared to support data processing by an IP block before the IP block begins executing a message's instructions, which depend on the specific memory contents.

[0041] At NOC 102 of Fig. 3, each IP block 104 is designed to bypass its memory data transfer storage unit 106 and send network-addressed data transfers 146 between IP blocks directly to the network via the IP block's network interface controller 108. Network-addressed data transfers are messages that are routed to another IP block by a network address. Such messages convey work data as part of pipelined applications, multiple data for a single program processing among IP blocks in a SIMD application, etc., as will be understood by one skilled in the art. Such messages differ from memory-address-based data transfers in that they are network-addressed from the outset by the originating IP block, which knows the network address to which the message is to be routed via routers of the NOC.Such network-addressed data transfers are forwarded from the IP block via I / O functions 136 in command format directly to the IP block's network interface controller, then converted into packet format by the network interface controller, and transmitted via routers of the NOC to another IP block. Such network-addressed data transfers 146 are bidirectional, potentially being forwarded to and from any IP block of the NOC, depending on their use in a particular application. However, each network interface controller is designed to both send and receive such data transfers to and from an associated router, and each network interface controller is designed to both send and receive such data transfers directly to and from an associated IP block, bypassing an associated memory data transfer controller 106.

[0042] Each network interface control unit 108 in the example of Fig. 3 is further designed to translate virtual channels in the network, characterizing network packets by type. Each network interface controller 108 includes virtual channel translation logic 148, which classifies each data transfer instruction by type and enters the instruction type into a field of the network packet format before forwarding the instruction in packet format to a router 110 for transmission in the NOC. Examples of data transfer instruction types include messages between IP blocks based on network addresses, request messages, responses to request messages, invalidation messages directed to caches, memory load and store messages, and responses to memory load messages, etc.

[0043] Each router 110 in the example of Fig. 3 includes forwarding logic 152, virtual channel control logic 154, and virtual channel buffers 156. The forwarding logic is typically implemented as a network of synchronous and asynchronous logic implementing a data transmission protocol stack for data transmission in the network formed by routers 110, interconnects 118, and bus lines between the routers. The forwarding logic 152 includes functionality that skilled readers might associate with forwarding tables in off-chip networking, where forwarding tables are considered too slow and cumbersome for use in an NOC in at least some embodiments. The forwarding logic implemented as a network of synchronous and asynchronous logic may be configured to make forwarding decisions as fast as a single clock cycle.In this example, the forwarding logic forwards packets by selecting a port to forward each packet received at a router. Each packet contains a network address to which the packet should be forwarded.

[0044] In the above description of memory address-based data transfers, each memory address was described as being mapped by network interface controllers to a network address, a network location of a memory data transfer controller. The network location of a memory data transfer controller 106 is naturally also the network location of the associated router 110 of that memory data transfer controller, the network interface controller 108, and the IP block 104. For this reason, when transferring data between IP blocks based on network addresses, it is also common for application-level computing to view network addresses as the location of an IP block within the network formed by the NOC's routers, links, and bus lines. Fig. Figure 2 shows that one design of such a network is a meshed network of rows and columns, where each network address can be implemented, for example, either as a unique identifier for each set of associated router(s), IP block, memory data transfer controller, and network interface controller of the meshed network, or x- and y-coordinates of each such set in the meshed network.

[0045] At NOC 102 of Fig. 3, each router 110 implements two or more virtual data transmission channels, each virtual data transmission channel being identified by a data transmission type. Data transmission instruction types and hence virtual channel types include the above-mentioned: messages between IP blocks based on network addresses, request messages, responses to request messages, invalidation messages directed to cache memories; load and store messages for memory; and responses to load messages for memory, etc. To support virtual channels, each router 110 in the example of Fig. 3 further includes virtual channel control logic 154 and virtual channel buffers 156. Virtual channel control logic 154 examines each received packet for its associated data transmission type and places each packet in an outgoing virtual channel buffer for that data transmission type for transmission over a port to a neighboring router in the NOC.

[0046] Each virtual channel buffer 156 has a finite amount of storage space. If many packets are received within a short period of time, the virtual channel buffer can fill up—to the point where no more packets can be stored in the buffer. With other protocols, packets arriving on a virtual channel whose buffer is full are lost. In this example, however, each virtual channel buffer 156 is designed with bus control signals so that surrounding routers are instructed via the virtual channel control logic to suspend transmission on a virtual channel—that is, to suspend the transmission of packets of a certain data transfer type. If one virtual channel is suspended in this way, all other virtual channels are unaffected—and can continue operating at full capacity.The control signals are wired through each router back to the router's associated network interface controller 108. Each network interface controller is configured to refuse to accept data transfer instructions for the suspended virtual channel upon receiving such a signal from the associated memory data transfer controller 106 or from the associated IP block 104. In this way, suspending a virtual channel affects all hardware implementing the virtual channel, all the way back to the source IP blocks.

[0047] One effect of suspending packet transmissions in a virtual channel is that packets are never dropped. When a router is exposed to a scenario where a packet might be dropped in an unreliable protocol, e.g., the Internet Protocol, the routers in the example of Fig. 3 suspend all transmissions of packets in a virtual channel through their virtual channel buffers 156 and their virtual channel control logic 154 until buffer space becomes available again, thereby avoiding the need to drop packets. The NOC of Fig. 3 can therefore implement highly reliable network data transmission protocols with an extremely thin hardware layer.

[0048] The exemplary NOC of Fig. 3 may further be configured to maintain cache coherence between both on-chip and off-chip memory caches. Each NOC may support multiple caches, each operating on the same underlying memory address space. For example, caches may be controlled by IP blocks, memory data transfer controllers, or cache controllers external to the NOC. Each of the on-chip memories 114, 116 may, in the example of Fig. 2 may also be implemented as a cache memory in the chip, and within the scope of the present invention, the cache memory may also be implemented off-chip.

[0049] Everyone in Fig. 3 includes five ports, four ports 158A to 158D connected to other routers via bus lines 118, and a fifth port 160 connecting each router to its associated IP block 104 via a network interface controller 108 and a memory data transfer controller 106. As can be seen from the illustrations of Fig. 2 and Fig. 3, the routers 110 and the connections 118 of the NOC 102 form a mesh network with vertical and horizontal connections connecting vertical and horizontal ports in each router. Fig. 3 For example, terminals 158A, 158D and 160 are referred to as vertical terminals and terminals 158B and 158D are referred to as horizontal terminals.

[0050] Fig. Figure 4 then shows, in another way, an exemplary implementation of an IP block 104 according to the invention, implemented as a processing element divided into an issue or instruction unit (IU) 162, an execution unit (XU) 164, and an auxiliary execution unit (AXU) 166. In the illustrated implementation, the IU 162 includes a plurality of instruction buffers 168 that receive instructions from an L1 instruction cache (iCACHE) 170. Each instruction buffer 168 is assigned to one of a plurality (e.g., four) of symmetric multithreaded (SMT) hardware threads. An Effective To Real Translation unit (iERAT) 172 is connected to the iCACHE 170 and is used to translate instruction fetch requests from a plurality of thread fetch sequencers 174 into absolute addresses for fetching instructions from low-level memory.Each thread fetch sequencer 174 is assigned to a specific hardware thread and is used to ensure that instructions to be executed by the associated thread are fetched into the iCACHE for dispatch to the appropriate execution unit. As also shown in . Fig. 4, instructions fetched into an instruction buffer 168 may also be monitored by branch prediction logic 176, which provides hints to each thread fetch sequencer 174 to minimize unsuccessful cache searches ("cache leaks") with respect to instructions caused by branches in executing threads.

[0051] The IU 162 further includes a dependency / issue logic block 178 assigned to each hardware thread and configured to resolve dependencies and control the issuance of instructions from the instruction buffer 168 to the XU 164. Furthermore, according to the illustrated embodiment, the separate dependency / issue logic 180 is provided in the AXU 166, allowing separate instructions to be issued concurrently from different threads to the XU 164 and the AXU 166. In an alternative embodiment, the logic 180 may be located in the IU 162 or omitted entirely, such that the logic 178 issues instructions to the AXU 166.

[0052] The XU 164, implemented as a fixed-point execution unit, includes a set of general purpose registers (GPRs) 182 connected to fixed-point logic 184, branch logic 186, and load / store logic 188. Load / store logic 188 is connected to an L1 data cache (dCACHE) 190, with effective-to-real translation provided by dERAT logic 192. The XU 164 can be configured to implement virtually any instruction set, such as all or part of a 32-bit or 64-bit PowerPC instruction set.

[0053] The AXU 166 operates as an auxiliary execution unit containing the dedicated dependency / issue logic 180 along with one or more execution blocks 194. The AXU 166 may contain any number of execution blocks and may implement virtually any type of execution unit, e.g., floating-point unit or one or more specific execution units such as encryption / decryption units, coprocessors, vector processing units, graphics processors, XML processing units, etc. In the illustrated embodiment, the AXU 166 includes a high-speed auxiliary interface to the XU 164, e.g., to support direct movement between AXU-architected state and XU-architected state.

[0054] The data transmission with the IP block 104 can be carried out as described above with respect to Fig. 2, via the network interface controller 108 connected to the NOC 102. Address-based data transfer, e.g., to access L2 caches, may be provided in conjunction with message-based data transfer. For example, each IP block 104 may include an associated inbox and / or outbox to handle cross-node data transfers between IP blocks.

[0055] Embodiments of the present invention may be practiced within the scope of the invention described above with respect to Fig. 1 to 4. However, those skilled in the art will understand, in light of the present disclosure, that the invention may be implemented in a variety of different environments, and that other changes may be made to the above-mentioned hardware and software embodiments without departing from the spirit and scope of the invention. As such, the invention is not limited to the particular hardware and software environment disclosed herein. Data encryption / compression based on memory address translation

[0056] Protecting unencrypted secure data from access outside a secure thread is of utmost importance in many computing applications. Generally, this data is only kept secure outside a chip, but as SOCs grow to hundreds of processors on a chip, it is increasingly important to protect unencrypted data even from other processors on the same chip. Typically, certain address ranges identified for encryption are filtered and encrypted / decrypted as the data is passed to and from a memory controller. However, this can result in unencrypted data being cached and accessible to other threads.

[0057] Encryption-related embodiments according to the invention, on the other hand, enable the protection of memory pages for exclusively authorized threads by adding one or more encryption-related page attributes to the memory address translation data structures used to perform memory address translation between virtual and absolute memory addresses. For example, in one embodiment, one or more encryption-related page attributes may be added to the page table entries (PTEs) of a processing core's Effective To Real Translation (ERAT) table, so that only trusted threads with access permission for a memory page are granted access to the page, as specified by the PTE for that page.If another thread attempts to access this page, a security interrupt is raised and handled as desired by a hypervisor or other supervisor-level software. The page attribute is also used to simplify hardware encryption / decryption by identifying memory accesses that need to be sent to an encryption module. For example, this page attribute can be sent in response to an L1 cache leak as a hint to an L2 cache or other hardware-level memory that the data to be reloaded must pass through the encryption module during load or store operations to encrypt / decrypt the data accordingly.

[0058] By integrating an encryption-related page attribute into a processor's memory address translation functionality, the overall security of secure data in a data processing system is improved while impacting performance to a lesser extent, particularly in a SOC with many SMT processing cores. Furthermore, PTEs are often not restricted to specific process identifiers, so different processes can be authorized to access the same encrypted data in some embodiments of the invention.

[0059] With regard to data being compressed in a data processing system, conventional data processing systems similarly control compressed memory regions using a series of registers to configure memory areas, often requiring additional hardware and complex software configuration, each of which can negatively impact system performance.

[0060] Compression-related embodiments according to the invention, on the other hand, add one or more compression-related page attributes to the memory address translation data structures used to perform memory address translation between virtual and absolute memory addresses in order to simplify the compression / decompression process and reduce latency associated with accessing compressed data.For example, in one embodiment, one or more compression-related page attributes may be added to the page table entries (PTEs) of a processing core's Effective To Real Translation (ERAT) table, so that when a PTE is initiated for a memory page containing compressed data, a corresponding compression-related page attribute may be passed to the memory subsystem when a load operation is performed, allowing the data from memory or a higher-level cache to pass directly through a hardware compression engine to be decrypted. The converse also applies; all memory data passes through the compression engine before being sent to the first level of compressed memory.This simplifies the process for managing compressed data and reduces the amount of supporting hardware and performance overhead associated with managing the compressed data.

[0061] Furthermore, in some embodiments according to the invention, page attributes may include level attributes that can be used to configure pages to be selectively encrypted and / or compressed in different levels of a memory system, e.g., in main memory or in higher-level caches such as L1, L2, or L3 caches. Thus, for example, some pages may be encrypted / compressed in L2 or L3 caches, while other pages may be encrypted / compressed in memory but decrypted / decompressed in L2 or L3 caches. This provides greater flexibility and can improve performance, particularly when it is desirable to accelerate memory access performance for frequently accessed data maintained in higher levels of a memory system.Among other advantages, compression as implemented here can effectively increase the size of a cache without requiring a corresponding increase in memory bandwidth.

[0062] Fig. Figure 5, for example, shows an exemplary data processing system 200 suitable for implementing data encryption / decryption based on memory address translation according to the invention. The system 200 is shown with a memory bus 202 connecting a plurality of processing cores 204 to a memory management unit (MMU) 206. Although in Fig. 5 only two processing cores 204 are shown, it should be understood that any number of processing cores may be used in different embodiments of the invention.

[0063] Each processing core 204 is an SMT core that includes a plurality (N) of hardware threads 208, along with an Effective To Real Translation (ERAT) table 210 and an integrated L1 cache 212. According to common understanding in the art, the ERAT table 210 serves as a cache for memory address translation data, e.g., PTEs, and is typically associated with a low-level data structure, e.g., a translation address buffer (TLB) 214 located within the MMU 206 or otherwise accessible. The TLB 214 may also serve as a cache for a larger page table, typically stored in memory 216.

[0064] The memory system may include multiple levels of memory and caches, and as such, data processing system 200 is shown with an L2 cache 218 coupled to MMU 206 and shared by processing cores 204. However, it should be understood that various alternative memory architectures may be used in other embodiments of the invention. For example, additional levels of cache, e.g., L3 caches, may be used, and memory 216 may be partitioned in some embodiments, e.g., in non-uniform memory access (NUMA)-based data processing systems.Furthermore, additional cache levels can be assigned to specific processing cores, so that, for example, each processing core includes a dedicated L2 cache that can be integrated into the processing core or connected between the processing core and the memory bus. In some embodiments, an L2 or L3 cache can be connected directly to the memory bus rather than through a dedicated interface to an MMU.

[0065] Furthermore, it should be understood that the Fig. 5 can be integrated into the same integrated circuit unit or chip, or arranged in multiple such chips. For example, in one embodiment, each processing core is implemented as an IP block in an NOC array, and a bus 202, an MMU 206, and an L2 cache 218 are integrated into the same chip as the processing cores in an SOC array.

[0066] In other embodiments, a bus 202, an MMU 206, an L2 cache 218, and / or a memory 216 may each be integrated into the same chip or into different chips of the processing cores, and in some cases, the processing cores may be located on separate chips.

[0067] For this reason, given the wide variety of known processor and memory architectures with which the invention may be used, it should be understood that the invention is not limited to the particular memory architecture shown herein.

[0068] To implement data encryption / compression based on memory address translation according to the invention, a data processing system 200 includes an encryption module 220 and a compression module 222 connected to a bus 202 and thus accessible for encryption / decryption and compression / decompression of data transmitted over bus 202. Although modules 220 and 222 are referred to as encryption and compression modules, respectively, it should be understood that module 220 typically includes both encryption and decryption logic, and that module 222 typically includes both compression and decompression logic, regardless of the designations used herein.It should be understood that in some embodiments, the encryption and decryption logic may be located in separate "modules," as may the compression and decompression logic. However, for the purposes of the invention, an encryption module may be considered a collection of hardware logic capable of performing encryption and / or decryption of data, and a compression module may be considered a collection of hardware logic capable of performing compression and / or decompression of data.

[0069] To simplify discussion of the invention, data processing system 200 is described as including both encryption and compression functionality. However, in some embodiments, it may be desirable to support only encryption or only compression functionality, and as such, embodiments according to the invention need not support both data encryption and data compression. Thus, in some embodiments, the respective module 220, 222 may be omitted, and the page attributes used to indicate whether a memory page is encrypted or compressed may, in some embodiments, include only encryption-related attributes or compression-related attributes.In addition, although the modules 220, 222 are shown as being connected to a bus 202, one or each of the modules 220, 222 may be connected to or integrated with an MMU 206.

[0070] As noted above, data encryption / compression based on memory address translation can be implemented by adding one or more page attributes to a memory address translation data structure, e.g., a page table entry (PTE). Fig. For example, Figure 6 shows an exemplary PTE 230 that may be managed in an ERAT 210 and extended to include various page attributes 232 through 238 to support data encryption / compression based on memory address translation. For encryption, an encrypted attribute 232, e.g., a 1-bit flag, may be used to indicate whether the data on the page is encrypted. Similarly, for compression, a compressed attribute 234, e.g., a 1-bit flag, may be used to indicate whether the data on the page is compressed.

[0071] Furthermore, in some embodiments, it may be desirable to optionally specify a level at which the data on a page is encrypted and / or compressed, such that the data is decrypted / decompressed at a higher level in the memory architecture (or optionally at the specified level), and the data is encrypted / compressed at the specified level and lower levels in the memory architecture. For example, a 2-bit level attribute, e.g., a level attribute 236 for encryption and a level attribute 238 for compression, may be provided to encode up to four memory levels, e.g., L1 = "00", L2 = "01", L3 = "10", and Memory = "11". Alternatively, a separate 1-bit tag may be associated with each memory level. Furthermore, in some embodiments, the encrypted and / or compressed attributes 232, 234 may be merged with the associated level attributes.For example, if no L3 cache is supported, one of the four states encoded in a 2-bit level attribute (e.g., "00") may represent that the page is unencrypted or uncompressed.

[0072] The PTE 230 also stores additional data, similar to conventional PTEs. For example, additional page attributes 240, e.g., attributes indicating whether a page is cacheable, protected, or read-only, whether memory coherence or write-through is required, an endian mode bit, etc., may be included in a PTE, as may one or more bits associated with user mode data 242, e.g., for software coherence or control over cache locking options. An access control page attribute 244 is provided to control which processes are granted access to a memory page, e.g.,by specifying a process identifier (PID) associated with the process that has access to the page, or optionally, a combination of matching and / or masking data or other data suitable for specifying a set of processes with access to a memory page. For example, the access control attribute may mask one or more LSBs from a PID, so that any PID that matches the MSBs in the access control attribute is granted access to the corresponding memory page. The ERAT page attribute 246 stores the effective-to-real translation data for the PTE, which typically includes the absolute address corresponding to the effective / virtual address used to access the PTE, as well as the effective / virtual address also used to index the ERAT via a CAM function.

[0073] It should be understood that the format of PTE 230 may also be used in TLB 214 and any other page table resident in the memory architecture. Alternatively, PTEs stored at different levels of the memory architecture may contain different data or omit some data, depending on the requirements of that particular level of the memory architecture. Furthermore, it should be understood that although in the embodiments discussed herein, the terms ERAT and TLB are used to describe various hardware logic that stores or caches memory address translation information in a processor or processing core, such hardware logic may be referred to by other names, and thus the invention is not limited to use with ERATs and TLBs. Furthermore, other PTE formats may be used, and for this reason, the invention is not limited to the particular PTE format described in Fig. 6 shown PTE format.

[0074] Because encryption-related and compression-related attributes are stored in a PTE, the determination of whether a page is encrypted and / or compressed, and the actual decryption and decompression of such data, are limited to only those processes and hardware threads executing on its behalf and authorized by the functionality in the computing system that otherwise controls access to the pages themselves. Page-based access control, typically used to prevent processes executing in a computing system from accessing or corrupting the memory of other processes resident in a computing system, can therefore be extended to support the management of encrypted and / or compressed data stored therein.

[0075] As is well known in the art, a hypervisor or other supervisor-level software, executing, for example, in firmware, a kernel, a partition manager, or an operating system, is typically used to allocate memory pages to specific processes and to handle access violations that might otherwise occur when a process attempts to access a page to which it is not authorized. Each supervisor-level software may, for example, manage an entire page table for the computing system, using dedicated hardware in the computing system to cache PTEs from a page table in the TLB 214 and ERATs 210. For this reason, embodiments according to the invention are capable of leveraging existing supervisor-level access controls to restrict access to encrypted and / or compressed data.For example, in many cases, a process may not even be able to determine whether a memory page is encrypted or compressed, because access to the PTE is restricted by supervisor-level software. For example, if no process executing on one processing core is authorized to access a memory page allocated to a process on another processing core, a process in the former processing core may not even be allowed to retrieve data from that memory page, even in encrypted or compressed form.

[0076] Fig. For example, FIG. 7 illustrates an example data processing system 250, and in particular, an example processing core provided therein, to illustrate an example memory access utilizing data encryption / decryption based on memory address translation in accordance with the invention. Address generation logic 252, such as provided in a load / store unit of a processing core, may generate a memory access request to access data (e.g., a cache line) from a particular memory page, e.g., in response to an instruction executed by a hardware thread (not shown) executing in the processing core.The memory access request is issued in parallel to both an ERAT 253 and an L1 cache 254, the former performing an address translation operation and determining whether the memory access request is legitimate for the PID associated with the requesting hardware thread, and the latter determining whether the cache line specified by the memory access request is currently cached in the L1 cache. In the illustrated embodiment of FIG. Fig. 7, the ERAT 253 is labeled "dERAT" and the L1 cache 254 is labeled "dCache" to indicate that these components are associated with data accesses and that corresponding iERAT and iCache components may be provided to handle instruction accesses (not shown).

[0077] In response to the memory access request, the ERAT 253 accesses a PTE 256 for the memory page specified by the memory access request. Hypervisor protection exception handler logic 258 compares a PID for the memory access request with the access control bits in the PTE, and if an access violation occurs because the PID is not authorized to access that memory page, the logic 258 signals an interrupt by issuing a software exception to the supervisor-level software, as shown at 260. If a memory access request is authorized but an L1 cache leak occurs, the memory access request is forwarded to a load / leak queue 262, which forwards the request to low-level memory, e.g., an L2 cache 264.

[0078] Fig. 8 shows a more detailed flow of operations 270 that may be performed in response to memory access requests issued by a hardware thread on behalf of a process in a data processing system 250. Protection logic, e.g., handler logic 258, accesses an ERAT 253 to determine whether a PTE 256 indicates that the requesting thread has permission to access the page associated with the memory access request (block 272). If permission is present (block 274), a determination is made as to whether the request can be satisfied by an L1 cache 254 (block 276). If the memory access does not produce a leak in the L1 cache 254, the request is satisfied by the L1 cache 254 (block 278), and processing of the memory access request is complete.

[0079] However, if the request produces a leak in the L1 cache 254, the request is forwarded to the load / leak queue 262 to add an entry to the queue corresponding to the request. In addition, it may be desirable to set an indicator in the entry to indicate that the request is associated with encrypted and / or compressed data. Next, before issuing the request to low-level storage, e.g., via a memory bus to an L2 cache or low-level storage, a determination is made in block 282 as to whether the page is indicated as encrypted and / or compressed, as determined from the page attributes in the PTE 256. If not, a bus transaction for the memory access request is issued in block 284.On the other hand, if the page is encrypted and / or compressed, a bus transaction is issued in block 286 with additional encryption / compression-related sideband data from the PTE 256.

[0080] The encryption / compression-related sideband data can be transmitted over a memory bus in various ways according to the invention. For example, additional control lines can be provided in a bus architecture to determine whether a bus transaction is associated with encrypted and / or compressed data, so that the status of one or more control lines can be used to determine whether the data is encrypted or compressed. Alternatively, transaction types can be associated with encrypted and / or compressed data, so that a determination can be made simply based on the transaction type of the bus transaction. In particular, in the latter case, no encryption module or compression module would be required to search for specific memory areas; instead, it could simply search for certain transaction types.

[0081] Referring again to block 274, if the requesting thread does not have access permission to the requested page, control transfers to block 288, which handles the access violation. In contrast to conventional access violation handling, in some embodiments, it may be desirable to perform alternative or enhanced operations on access violations associated with encrypted data to address the additional concerns related to attempts to gain unauthorized access to secure data. Thus, when an access violation is detected, block 288 determines whether the page is encrypted by accessing the encryption-related page attribute for the page. If the page is not encrypted, a software exception is raised (block 290) and handled in the conventional manner.If, on the other hand, the page is encrypted, control transfers to block 292 to determine whether to shut down the data processing system. In some high-security embodiments, e.g., where highly confidential information is potentially stored in a data processing system, it may be desirable to shut down a system immediately in the event of a potential attack. Thus, it may be desirable to provide a configurable mode in which an access violation results in an immediate system shutdown. If the configurable mode is set, block 292 thus returns control to block 294 to shut down the system. Otherwise, block 292 returns control to block 296 to throw a software exception similar to a conventional exception, except for an indicator that the page for which the exception is signaled was encrypted. A software, e.g.,a kernel or other supervisor-level program may then perform advanced exception handling, such as logging an attempt to access an encrypted page, notifying a central processor, sending messages over a network to an external device, cleaning up memory and state contents, or other desired operations, depending on the security requirements of the data and / or the computing system.

[0082] The Fig. 9 and Fig. 10 each represent sequences of operations related to the execution of read and write bus transactions, which are executed in response to the steps described above in connection with Fig. 8 described memory access request can be issued. Fig. 9, for example, shows a sequence of operations 300 for processing a read bus transaction. The read bus transaction first results in the request being fulfilled by the bus (block 302), e.g., by an L2 or L3 cache if the data has already been cached, or by memory if not. Determinations are then made as to whether the transaction indicates that the data is encrypted (block 304) or compressed (block 306), e.g., by encryption and compression modules, each of which scans the bus for a transaction type associated with encrypted or compressed transactions, or by checking for asserted control lines on the bus.

[0083] If the data is not encrypted or compressed, the data is returned in the conventional manner (block 308). However, if the data is encrypted, the returned data is encrypted by an encryption module (e.g., encryption module 220 of Fig. 5) to decrypt the data before returning the data to the requesting processing core (block 310). If the data is compressed, the returned data is similarly passed through a compression module (e.g., compression module 222 of Fig. 5) before the data is returned to the requesting processing core (block 312).

[0084] As in Fig. 10, a write bus transaction is processed similarly to a read bus transaction via a sequence of operations 320. The write bus transaction includes a cache line to be written to low-level memory; however, before forwarding the data to an appropriate destination, e.g., from an L2 or L3 cache or main memory, determinations are made as to whether the transaction indicates that the data is encrypted (block 322) or compressed (block 324), e.g., by encryption and compression modules, each of which scans the bus for a transaction type associated with encrypted or compressed transactions, or looks for asserted control lines on the bus.

[0085] If the data is not encrypted or compressed, the data is written to the corresponding destination in the conventional manner (block 326). However, if the data is encrypted, the write data is first encrypted by an encryption module (e.g., encryption module 220 of Fig. 5) to encrypt the data before writing (block 328). Similarly, if the data is compressed, the write data is first passed through a compression module (e.g., compression module 222 of Fig. 5) before the data is written (block 330).

[0086] The Fig. 11 and Fig. 12 each represent sequences of operations related to the execution of read and write bus transactions, which are carried out in response to the above-mentioned Fig. 8 discussed memory access request, but for an implementation that supports level attributes in a PTE to control at which levels a memory page is encrypted / compressed in a multi-level memory architecture.

[0087] Fig. 11, for example, shows a sequence of operations 350 for processing a read bus transaction. The read bus transaction first results in the request being fulfilled by the bus (block 352), e.g., by an L2 or L3 cache if the data has already been cached, or by memory if not. Next, at the level of the memory architecture providing the data, determinations are made as to whether the transaction indicates that the data is encrypted (block 354) or compressed (block 356), e.g., by encryption and compression modules, each of which scans the bus for a transaction type associated with encrypted or compressed transactions, or by checking for asserted control lines on the bus.In the illustrated embodiment, the requested data is encrypted or compressed if the associated level attribute is equal to or higher than the memory level from which the data is provided (e.g., if the level attribute indicates the L2 cache, the data is encrypted / compressed in the L2 cache, in an L3 or low-level cache, and in main memory).

[0088] If the data is not encrypted or compressed at the level at which the data is provided, the data is returned in the conventional manner (block 358). However, if the data is encrypted at the source level, the returned data is encrypted by an encryption module (e.g., encryption module 220 of Fig. 5) to decrypt the data before returning the data to the requesting processing core (block 360). If the data is compressed at the source level, the returned data is similarly compressed by a compression module (e.g., compression module 222 of Fig. 5) before the data is returned to the requesting processing core (block 362).

[0089] As in Fig. 12, a write bus transaction is processed similarly to a read bus transaction via a sequence of operations 370. The write bus transaction includes a cache line to be written to low-level memory; however, before forwarding the data to an appropriate destination, e.g., from an L2 or L3 cache or main memory, determinations are made as to whether the transaction indicates that the data is encrypted (block 372) or compressed (block 374) at the transaction's destination level, e.g., by encryption and compression modules, each scanning the bus for a transaction type associated with encrypted or compressed transactions, or checking for asserted control lines on the bus.

[0090] If the data is not encrypted or compressed at the target level, the data is written to the corresponding target in the conventional manner (block 376). However, if the data is encrypted at the target level, the write data is first encrypted by an encryption module (e.g., encryption module 220 of Fig. 5) to encrypt the data before writing (block 378). Similarly, if the data is compressed at the target level, the write data is first passed through a compression module (e.g., compression module 222 of Fig. 5) before the data is written (block 380).

[0091] Thus, the data in the Fig. 11 to 12, the memory is encrypted and / or compressed only at selected levels in the memory hierarchy, with different selected levels potentially being defined for different memory pages. As such, a significant degree of flexibility can be provided for different applications. However, it should be understood that level attributes may not be implemented in all implementations, and furthermore, in implementations that support both encryption and compression, level attributes may not be supported with respect to both functionalities.

[0092] In some embodiments according to the invention, additional memory transactions may be supported. For example, read-modify-write transactions may be used to encrypt and / or recompress an entire cache line when updating data in an encrypted or compressed cache line. Integrated encryption module

[0093] As noted above, in some embodiments, the encryption module may be integrated into a processing core to effectively provide a secure cache within the processing core, thereby providing further protection for secure data throughout the memory system. The encryption-related page attributes mentioned above may be used to access this secure cache to prevent secure data from ever leaving the processing core in an unencrypted form, so that the secure data is encrypted whenever it is outside the processing core. The integrated encryption module is typically configured to perform encryption operations such as encryption and decryption.

[0094] For example, in one embodiment, a separate secure L1 cache may be provided along with a standard (insecure) L1 cache so that the secure L1 cache can be accessed in parallel with the standard cache and, once accessed, multiplexed into the pipeline. Upon a loss in the secure cache, a load may be dispatched to the next level of the memory system, e.g., an L2 cache. An encryption-related page attribute may be maintained in the load / lose queue so that when the encrypted data is returned, it can be passed through the onboard encryption engine for decryption and into the secure cache, ensuring that only an authorized secure thread ever gains access to the unencrypted data.However, in another embodiment, a separate secure L1 cache may not be used, and the integrated encryption module may be used to selectively encrypt and decrypt data stored in the L1 cache based on the encryption-related page attributes associated with that data.

[0095] Fig. For example, FIG. 13 shows an example data processing system 400, and in particular, an example processing core provided therein to illustrate an example memory access utilizing an integrated encryption module according to the invention. Address generation logic 402, such as provided in a load / store unit of a processing core, may generate a memory access request to access data (e.g., a cache line) from a particular memory page, e.g., in response to an instruction executed by a hardware thread (not shown) executing in the processing core.The memory access request is issued in parallel to both an ERAT 406 and an L1 cache 404, where the former determines whether the memory access request is authorized for the PID associated with the requesting hardware thread, and the latter determines whether the cache line specified by the memory access request is currently cached in the L1 cache. In the illustrated embodiment of FIG. Fig. 13, the ERAT 406 is referred to as "dERAT" and the L1 cache 404 is referred to as "dCache" to indicate that these components are associated with data accesses and that corresponding iERAT and iCache components may be provided to handle instruction accesses (not shown).

[0096] In response to the memory access request, the ERAT 406 accesses a PTE 408 for the memory page specified by the memory access request. Hypervisor protection exception handler logic 410 compares a PID for the memory access request with the access control bits in the PTE, and if an access violation occurs because the PID is not authorized to access that memory page, the logic 410 signals an interrupt by issuing a software exception to the supervisor-level software, as shown at 412. If a memory access request is authorized but an L1 cache leak occurs, the memory access request is forwarded to a load / leak queue 414, which forwards the request to low-level memory, e.g., an L2 cache 416.

[0097] In addition, the L1 cache 404 includes an encryption module 418 connected thereto and integrated into the processing core. The encryption module 418 may be used, for example, to encrypt data written from the processing core, e.g., from the L1 cache 404, or to decrypt data used by the processing core and received from low-level memory or the L1 cache. For example, the encryption module 418 may be configured to forward decrypted data to a register file 420 or to forward decrypted data back to the L1 cache 404 via a bus 422. Furthermore, the encryption module 418 may be connected to a bypass network 424 to bypass the register file 420 and provide decrypted data directly to an execution unit (not shown).

[0098] At the Fig. A variety of scenarios can be implemented in the configuration shown in Figure 13. For example, an L1 cache 404 can be used to store both encrypted and unencrypted data. Alternatively, an L1 cache 404 can be a secure cache, with a separate L1 cache (not shown) used to store insecure data. Furthermore, in some embodiments, the L1 cache 404 can store decrypted data, so the encryption module 418 is used to decrypt data received from and encrypt data written from low-level memories.Alternatively, for particularly security-sensitive applications, it may be desirable to store all secure data in encrypted form in the L1 cache 404, where the encryption module 418 may be used to decrypt data retrieved from the L1 cache 404 into a register file 420 and to decrypt data written from the register file 420 back to the L1 cache 404.

[0099] In connection with some embodiments, it may be desirable to place an indicator or attribute in an entry in the load / lose queue associated with a memory access request for a secure memory page. In this way, when the requested data is returned, the indicator can be accessed to determine whether the data is encrypted. For example, if two separate secure and insecure L1 caches are provided in a processing core, the indicator can thus be used to direct the returned data to the appropriate L1 cache. Alternatively, if secure data is stored in an L1 cache in an unencrypted format, the indicator can be used to cause the returned data to be passed through the encryption module before being stored in the L1 cache.

[0100] Referring to Fig. 14, in some cases, as mentioned above, it may be desirable to use separate secure and insecure L1 caches in a processing core, for example, so that there is no performance overhead associated with encrypting and decrypting data for insecure data. For example, a processing core 450 includes a decryption module 452 connected to a load / store unit 454. An ERAT 456 in the load / store unit 454 provides encryption-related and other page attributes for both an insecure L1 cache 458 and a secure L1 cache 460. Each L1 cache 458, 460 may have a separate load / leak queue 462, 464 and bus connections 466, 468 to a memory bus 470, or alternatively, may share the same load / leak queue and bus connection.

[0101] Conventional designs. For example, the use of memory attributes often bypasses the processing overhead otherwise incurred by a memory management unit in setting up and managing regions or ranges of encrypted or compressed data. Furthermore, when multiple processing cores are connected to the same bus, other processing cores are typically unable to access encrypted or compressed data for another processing core, because the page attributes for these are managed within each processing core and only authorized processes are allowed to access them.

[0102] Furthermore, various modifications may be made in accordance with the invention. For example, as with an integrated encryption module, it may be desirable to provide an integrated compression module within a processing core or other types of acceleration modules capable of utilizing page attributes to selectively process data based on them.

[0103] Other changes may be made to the disclosed embodiments without departing from the spirit and scope of the invention. For this reason, the invention lies in the claims. Implementation of an integrated encryption module may vary in different embodiments. For example, in some embodiments, an integrated encryption module may be implemented as an AXU, as described above in connection with Fig. 4. However, the invention is not limited to such an embodiment.

Claims

[1] A method for accessing data in a data processing system, the method comprising: in response to a memory access request initiated by a thread in a processing core (450), accessing a memory address translation data structure (456) to perform a memory address translation for the memory access request, the processing core comprising a secure L1 cache (460) for storing encrypted data, another L1 cache (458) for storing unencrypted data, and an integrated encryption module (452); Accessing an encryption-related page attribute in the memory address translation data structure to determine whether the memory page associated with the memory access request is encrypted; and Fulfilling the memory access requirement by selectively passing secure data on the memory page through the integrated encryption module to undergo a decryption operation, depending on whether the memory page associated with the memory access request has been determined to be encrypted and is stored in the secure L1 cache; and Use data on the memory page, depending on whether the memory page associated with the memory access request is stored in the other L1 cache. [2] The method of claim 1, wherein the memory address translation data structure includes a plurality of page table entries, each page table entry including an absolute address associated with the memory page associated with the page table entry and the encryption-related page attribute associated with that memory page. [3] The method of claim 1, further comprising, in response to a second memory access request initiated by a second thread, restricting access to the memory page by the second thread by comparing a process ID associated with the second thread to access control data in the memory address translation data structure, and denying access to the memory page by the second thread in response to the comparison. [4] The method of claim 3, wherein the memory access request comprises a first memory access request associated with a first process executed by the processing core, the method further comprising throwing a software exception in response to a second memory access request associated with a second process that does not have access permission to the memory page. [5] The method of claim 4, further comprising, in response to the second memory access request, determining whether the memory page is encrypted and, if so, indicating that the page is encrypted when the software exception is thrown. [6] The method of claim 3, wherein the memory access request comprises a first memory access request associated with a first process executed by the processing core, the method further comprising, in response to a second memory access request associated with a second process that is not authorized to access the memory page, determining whether the memory page is encrypted and, if so, shutting down the system. [7] The method of claim 1, wherein the encryption-related page attribute comprises an encrypted attribute indicating whether the memory page is encrypted and a level attribute identifying at least one level of a multi-level memory architecture at which the memory page is encrypted. [8] The method of claim 7, wherein the selective routing of secure data on the memory page by the encryption module is performed depending on whether the level attribute indicates that the memory page is encrypted at a source or destination level for the memory access request. [9] Circuit arrangement, comprising: a multi-core processor containing a plurality of processing cores; and a memory address translation data structure (456) disposed in a first processing core (450) of the plurality of processing cores, the memory address translation data structure configured to store address translation data for a memory page, the memory address translation data structure further configured to store an encryption-related page attribute for the memory page, the first processing core comprising a secure L1 cache (460) for storing encrypted data, another L1 cache (458) for storing unencrypted data, and an integrated decryption module (452); wherein the multi-core processor is configured to perform the method according to any one of the preceding claims. [10] A program product comprising a computer-readable medium and logic definition program code stored in the computer-readable medium and defining the circuit arrangement of claim 9.

Citation Information

Patent Citations

  • System and method for using address bits to affect encryption

    US20060059553A1

  • Multi-level memory architecture

    US20080077922A1