QUICKLY CLEAR SYSTEM MEMORY
Dynamic compression of cache lines at the data link layer addresses latency issues in processor connections, enhancing data processing efficiency and reducing power consumption.
Patent Information
- Application Number
- DE112023004089
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-07-12
- Publication Date
- 2025-09-04
AI Technical Summary
Data transmission between processors in multiprocessor systems is constrained by latency issues due to speed differentials in processor connections, leading to performance degradation and increased power consumption.
Dynamically compressing cache lines at the data link layer of the network stack, using a framer and data link layer components, to optimize bandwidth and latency without software intervention.
Reduces latency and improves data processing efficiency by enabling real-time adaptive compression, maintaining performance without compromising security or increasing complexity.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] The present invention relates generally to data processing systems and, more particularly, to various embodiments for compressing data in latency-critical processor connections of a data processing system in a data processing environment using a data processor. SUMMARY
[0002] According to one embodiment of the present invention, a method for compressing data in latency-critical processor links of a data processing system in a data processing environment using a data processing processor is presented. One or more cache lines may be dynamically compressed at a lowest level of a network stack based on one or more of a plurality of parameters prior to transferring a single cache line, the network stack including a framer and a data link layer.
[0003] One embodiment includes a computer-usable program product. The computer-usable program product includes a computer-readable storage device and program instructions stored in the storage device.
[0004] One embodiment includes a computer system. The computer system includes a processor, a computer-readable memory, and a computer-readable storage device, as well as program instructions stored in the storage device for execution by the processor via the memory.
[0005] Therefore, in addition to the preceding exemplary embodiment of the method, further exemplary embodiments of systems and computer program products are provided. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is a block diagram of an exemplary cloud computing node according to an embodiment of the present invention. Fig. 2 illustrates a cloud computing environment according to an embodiment of the present invention. Fig. 3 illustrates abstraction model layers according to an embodiment of the present invention. Fig. 4 is another block diagram illustrating an exemplary functional relationship between various aspects of the present invention. Fig. 5 is a block diagram illustrating operations for compressing data in latency-critical processor interconnects of a data processing system in a data processing environment according to an embodiment of the present invention. Fig. 6 is another block diagram illustrating operations for compressing data in latency-critical processor interconnects of a data processing system in a data processing environment according to an embodiment of the present invention. Fig. 7 is another block diagram illustrating operations for compressing data in latency-critical processor interconnects of a data processing system in a data processing environment according to an embodiment of the present invention. Fig. 8 is another block diagram illustrating operations for compressing data in latency-critical processor interconnects of a data processing system including compression in a framer and decompression in a parser in a data processing environment according to an embodiment of the present invention. Fig. 9 is a flowchart illustrating another exemplary method for compressing data in latency-critical processor interconnects of a data processing system in a data processing environment according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE DRAWINGS
[0006] In modern computer systems, a multiprocessor system comprises central processing units (CPUs) with multiple cores in a single module. Typically, data transfer between processors in the multiprocessor system occurs via an interprocessor bus (also called a processor link). The processors connected by the processor link (i.e., a driver processor and a target processor) are typically associated with I / O parameters that govern analog properties of a signal transmitted by the driver processor and the corresponding signal received by the target processor. Characterizing the processor link during a test / validation phase can help determine the best I / O parameters to reliably achieve the desired performance levels.
[0007] In addition, computer hardware caches are temporary storage locations for quickly accessing frequently used memory data. In other words, to reduce or eliminate the time delay (or "latency") when accessing data stored in a computer's main memory, modern computer processors include a cache that stores recently accessed data so the processor can quickly access it again. Data stored in a cache can be quickly accessed by a processor without having to access main memory (or "memory"), thereby increasing the performance of the processor and the computer overall. A cache has a shorter access time than the computer system's memory (e.g., often referred to as "DRAM" (dynamic random-access memory)). Caches are typically implemented with static random-access memory ("SRAM"), which is faster than DRAM.However, cache capacities are lower than those of DRAM. Access speed and cache / memory capacity are inversely proportional.
[0008] Several different cache layers can be deployed in a computer system. For example, a Level 1 (or primary cache) is used to store data for system memory (which includes random access memory, i.e., RAM) so that a processor can access it. A Level 1 (“L1”) cache can be integrated directly into the processor and run at the same speed as the processor, providing the fastest possible access time. A Level 2 (or secondary) (“L2”) cache is also used to store a portion of the system memory and may be integrated into a chip package but is separate from the processor. A Level 2 cache has a larger capacity than a Level 1 cache but is slower. Some systems may even include a Level 3 (“L3”) cache, which has an even larger capacity than a Level 2 cache.However, a Level 3 cache is typically slower than a Level 2 cache, but still faster than the primary storage unit, and can be located outside the chip package.
[0009] Data in a cache is stored in "lines," which are contiguous blocks of data (i.e., with a byte length equal to a power of two and aligned on boundaries equal to that size). This means that data is typically transferred and accessed in groupings called cache lines, which may contain more than one piece of data.
[0010] In a typical computing architecture, a processor core may be connected to a cache (e.g., an on-chip cache or a "nest"), which in turn is connected to the processor interconnect ("interconnect"). The interleaving and interconnect are designed to operate at a much slower speed than the processor core. The interleaving may be a computing infrastructure that handles the transfer of data between the core and other cores, main memory, and input / output units ("I / O units"). The interleaving is responsible for maintaining the coherence of data rows in the system. It is not an industry-standard term, so perhaps it should be further defined before being used.In a data processing system, there are many data connections—connections between processors and I / O devices (PCIe, Ethernet), connections between processors and memory, and connections between processors. Interleaving encompasses the entire infrastructure that manages these connections. Thus, an interconnect can be a connection that enables the transfer of data to / from a processor. In other words, in a computer system configuration where caches are shared between different processors connected by interprocessor interconnects, coherent transfer of cache data lines between processors can occur.
[0011] Typical data processing systems have a handshake protocol for transferring cache lines, and only a limited number of requests can be outstanding over the link. Requests to transfer cache lines are satisfied but are limited by the link speed. Since not all processor core computational operations require interleaving or interconnecting, computer system performance is provided at a reasonable power consumption. However, when the processor core requests a cache line over the processor interconnect, the speed difference negatively impacts computer system performance. For example, physical constraints limit the link speed, which is a combination of the number of serial channels per interconnect and the speed of the operation.
[0012] The present invention thus provides for improving and optimizing the effective bandwidth and latency of such a limited processor connection, while reducing the overhead of the data processor to improve and increase the efficiency of data processing.
[0013] In this way, the present invention provides increased performance of data processing systems by reducing the latency in transferring a single cache line through lossless compression.
[0014] Accordingly, various embodiments are provided herein for compressing data in latency-critical processor interconnects of a computing system in a computing environment using a computing processor. One or more cache lines may be dynamically compressed at the lowest level of a network stack based on one or more of a plurality of parameters prior to transmitting a single cache line, wherein the network stack includes a framer and a data link layer. In some examples, data packets prepared for transmission over the interconnect may include header information, cache line data, and some control information. The cache line data may be compressed, while the other components of the data packet (e.g., header, control flags) are not.In compressed form, more cache line data can actually fit into a single data packet. This means that the cache lines can be dynamically compressed at the lowest level of the network stack (framer / DLL), as opposed to the application layer.
[0015] It should be noted that a network stack, as used here, can contain the following layers: 1) an application layer (the application layer can be the core and any software running), 2) a transport layer (the layer that ensures the coherence of data transmissions—it creates packets from the cache line data for transmission and processes the received packets to send them to the application layer), 3) a data link layer (the layer responsible for ensuring error-free and orderly delivery of data packets), and 4) a physical layer (the actual electrical circuitry). Thus, a core can be the application layer, an interleaving can be the transport layer, and a link can be the link layer and the physical layer.
[0016] As such, low-level compression is a purely hardware-controlled mechanism (HW), without any intervention from the kernel or software application.
[0017] For example, the framer can be defined as a component / module that appends some additional information in a predefined organization to the data for transmission over the connection so that the data can be processed on the receiving side. The DLL can be the data link layer.
[0018] In some examples, various embodiments provide compression of data in the data link layer with framer feedforward. In some examples, various embodiments provide compression of data in a framer with bus feedback. In some examples, various embodiments provide dynamic compression switching based on cache line contention. For example, an attacker may attempt to generate frequent transfers of cache lines between processors to obtain information about their compressibility. If a high number of transfers is observed in a particular address space, meaning there is heavy contention for some cache lines, then it is safer to disable compression so that no information is lost.
[0019] In this way, various embodiments provide pipelined compression at the data link layer, enabling compression in a data processing system without compromising performance. That is, the present invention provides low-latency compression on the data links without compromising security via side channels, unlike when the data is constantly compressed only in the cache. By enabling dynamic compression operations on the data links, any impact (positive or negative) of compression on performance can be evaluated in real time, allowing corrective actions / decisions to be made to adjust compression and reduce latency.
[0020] In some examples, various embodiments provide for compressing data at the data link layer by using control signals at the link layer. That is, data can be dynamically compressed by dynamically adapting it to the conditions of the data link layer to increase the compression efficiency with respect to latency to at least the same speed as the data link speed without compression.
[0021] In some embodiments, various embodiments provide for implementing link-level compression to reduce link latency without increasing cache implementation complexity and to make the overall latency of cache data access equal to or better than a system without compression. The cache line is considered the smallest entity to work with and can therefore be implemented on link data.
[0022] It should be noted that one or more calculations may be performed using various mathematical operations or functions, which may involve one or more mathematical operations (e.g., analytical or computational solution of differential equations or partial differential equations using addition, subtraction, division, multiplication, standard deviations, means, averages, percentages, statistical modeling using statistical distributions, by determining minima, maxima or similar thresholds for combined variables, etc.).
[0023] In general, "optimize," as used here, can refer to and / or be defined as "maximizing," "minimizing," "best," or achieving one or more specific goals, objectives, or goals. Optimizing can also refer to maximizing a benefit for a user (e.g., maximizing the benefit of a trained machine learning planning agent). Optimizing can also refer to the most effective or expedient use of a situation, opportunity, or resource.
[0024] Furthermore, optimizing does not necessarily refer to a best solution or result, but can, for example, refer to a solution or result that is "good enough" for a particular application. In some implementations, the goal is to propose a "best" combination of operations, schedules, compression / decompression operations, and / or framer / DLL manager options, but there may be a variety of factors that can lead to an alternative proposal of a combination of operations, schedules, compression / decompression operations, and / or framer / DLL manager options that yields better results. Here, the term "optimize" can refer to such results achieved based on minima (or maxima, depending on which parameters are considered in the optimization problem).In another aspect, the terms "optimize" and / or "optimizing" may refer to an operation performed to achieve a better result, e.g., lower execution costs or higher resource utilization, regardless of whether the optimal result is actually achieved. Likewise, the term "optimize" may refer to a component for performing such an improvement operation, and the term "optimized" may be used to describe the result of such an improvement operation.
[0025] While this disclosure includes a detailed description of cloud computing, it is to be understood in advance that implementation of the teachings presented herein is not limited to a cloud computing environment. Rather, embodiments of the present invention may be practiced in conjunction with any type of computing environment now known or later developed.
[0026] Cloud computing is a service delivery model that enables seamless, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model can include at least five characteristics, at least three service models, and at least four implementation models.
[0027] The properties are as follows:
[0028] On-Demand Self-Service: A cloud user can unilaterally and automatically provision computing capabilities such as server time and network storage as needed, without requiring human interaction with the service provider.
[0029] Broad Network Access: Functions are available over a network and accessed through standard mechanisms that support use by heterogeneous lightweight or performance-intensive client platforms (e.g., mobile phones, laptops, and PDAs).
[0030] Resource Pooling: The provider's computing resources are pooled to serve multiple users using a multi-tenant model, with various physical and virtual resources dynamically allocated and reassigned as needed. There is a perceived location independence, as the user generally has no control or knowledge over the exact location of the provided resources, but may be able to specify a location at a higher level of abstraction (e.g., country, state, or data center).
[0031] Rapid Elasticity: Features can be deployed quickly and elastically for rapid scale-out, in some cases automatically, and released quickly for rapid scale-in. To the user, the features available for deployment often appear unlimited, and they can be purchased in any quantity at any time.
[0032] Measured Service: Cloud systems automatically control and optimize resource usage by leveraging a measurement function at a certain level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and logged, providing transparency for both the provider and the customer of the service used.
[0033] There are the following service models:
[0034] Software as a Service (SaaS): The functionality provided to the user consists of using the provider's applications running on a cloud infrastructure. The applications are accessible from various client devices via a lightweight client interface such as a web browser (e.g., web-based email). The user does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, with the possible exception of limited user-specific application configuration settings.
[0035] Platform as a Service (PaaS): The functionality provided to the user is to deploy applications created or obtained by the user, using programming languages and tools supported by the provider, on the cloud infrastructure. The user does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but has control over the deployed applications and possibly over configurations of the application hosting environment.
[0036] Infrastructure as a Service (LAA): The functionality provided to the user consists of providing processing, storage, networking, and other basic computing resources, enabling the user to deploy and run any software, including operating systems and applications. The user does not manage or control the underlying cloud infrastructure, but has control over operating systems, storage, deployed applications, and possibly limited control over selected network components (e.g., host firewalls).
[0037] There are the following deployment models:
[0038] Private Cloud: The cloud infrastructure is operated solely for an organization. It can be managed by the organization or a third party and can be located on its own premises or on a third-party site.
[0039] Community Cloud: The cloud infrastructure is shared by multiple organizations and supports a specific user community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by the organizations or a third party and can be located on their own premises or on a third-party premises.
[0040] Public Cloud: Cloud infrastructure is made available to the general public or a large industry group and is owned by an organization that sells cloud services.
[0041] Hybrid Cloud: The cloud infrastructure consists of two or more clouds (private, community, or public) that remain separate entities but are interconnected by a standardized or proprietary technology that enables data and application portability (e.g., cloud targeting for load balancing between clouds).
[0042] A cloud computing environment is service-oriented and focuses on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure comprising a network of interconnected nodes.
[0043] With reference now to Fig. Figure 1 illustrates a schematic of an example cloud computing node. Cloud computing node 10 is merely one example of a suitable cloud computing node and is not intended to suggest any limitation on the scope of use or functionality of embodiments of the invention described herein. Nevertheless, cloud computing node 10 may be implemented and / or perform any of the functionalities set forth above.
[0044] Within the cloud computing node 10, there is a computer system / server 12 that is operable with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of known computing systems, environments, and / or configurations that may be suitable for use with the computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, lightweight clients, high-performance clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, desktop computers, programmable consumer electronics, network PCs, minicomputer systems, mainframe systems, and distributed cloud computing environments that include any of the foregoing systems or devices, and the like.
[0045] The computer system / server 12 may be described in the general context of computer system-executable instructions, e.g., program modules executed by a computer system. In general, program modules may include routines, programs, objects, components, logic, data structures, etc., that perform specific tasks or implement specific abstract data types. The computer system / server 12 may be deployed in distributed cloud computing environments where the tasks are performed by remotely located processing units connected via a communications network. In a distributed cloud computing environment, program modules may be located on both local and remote computer system storage media, including storage units with random access memories.
[0046] As in Fig. 1, the computer system / server 12 in the cloud computing node 10 is represented as a general-purpose computing device. The components of the computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 connecting various system components, including the system memory 28, to the processor 16.
[0047] Bus 18 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port (AGP) interface, and a processor or local bus utilizing any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0048] The computer system / server 12 typically includes a variety of computer-readable media. These media may be any available media accessible to the computer system / server 12, including volatile and non-volatile media, removable and non-removable media.
[0049] System memory 28 may include computer system-readable media in the form of volatile memory, e.g., random access memory (RAM) 30 and / or cache 32. Computer system / server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. For example only, storage system 34 may be provided to read and write to a non-removable, non-volatile magnetic medium (not shown and commonly referred to as a "hard drive"). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic storage disk (e.g., "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical storage disk such as a CD-ROM, DVD-ROM, and other optical media may be provided.In such cases, each may be connected to bus 18 via one or more data media interfaces. As further illustrated and described below, system memory 28 may include at least one program product having a (e.g., at least one) set of program modules configured to perform the functions of embodiments of the invention.
[0050] The program / utility 40 with (at least) one set of program modules 42 may be stored, for example and not limited to, in the system memory 28, as well as an operating system, one or more application programs, other program modules, and program data. The operating system, one or more application programs, other program modules, and program data, or a combination thereof, may each include an implementation of a network environment. The program modules 42 generally perform the functions and / or methodologies of embodiments of the invention described herein.
[0051] The computer system / server 12 may also communicate with one or more external devices 14, e.g., a keyboard, a pointing device, a display 24, etc.; as well as with one or more devices that enable a user to interact with the computer system / server 12; and / or any devices (e.g., network card, modem, etc.) that enable the computer system / server 12 to communicate with one or more computing devices. Such data communication may occur via the input / output (I / O) interfaces 22. Furthermore, the computer system / server 12 may communicate with one or more networks, e.g., a local area network (LAN), a general purpose wide area network (WAN), and / or a public network (e.g., the Internet), via the network adapter 20.As shown, network adapter 20 communicates with the other components of computer system / server 12 via bus 18. It should be understood that other hardware and / or software components may be used in conjunction with computer system / server 12 even if not shown. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external hard disk drive arrays, RAID systems, tape drives, and data archiving storage systems, among others.
[0052] With reference now to Fig. 2, an illustrative cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud users, such as the personal digital assistant (PDA) or mobile phone 54A, the desktop computer 54B, the laptop computer 54C, and / or the automotive computer system 54N, can communicate. The nodes 10 can communicate with each other. They can be physically or virtually grouped in one or more networks, such as private, shared, public, or hybrid clouds as described above, or a combination thereof (not shown). This enables the cloud computing environment 50 to offer infrastructure, platforms, and / or software as services for which a cloud user does not need to maintain resources on a local computing device.It can be seen that the types of data processing units 54A to N shown in . Fig. 2 are intended to be illustrative only and the data processing nodes 10 and the cloud computing environment 50 can communicate with any type of computer-based device over any type of network and / or network-addressable connection (e.g., via a web browser).
[0053] With reference to Fig. 3 now shows a set of functional abstraction layers used by the cloud computing environment 50 ( Fig. 2) are provided. It is understood from the outset that the Fig. The components, layers, and functions illustrated in Figure 3 are intended to be illustrative only, and embodiments of the invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided:
[0054] The device layer 55 includes physical and / or virtual devices that are and / or are integrated with standalone electronics, sensors, actuators, and other objects to perform various tasks in a cloud computing environment 50. Each of the devices in the device layer 55 is networkable with other functional abstraction layers so that information retrieved from the devices can be provided to them and / or information from the other abstraction layers can be provided to the devices. In one embodiment, the various devices, including the device layer 55, may comprise a network of entities commonly known as the "Internet of Things" (IoT). Such a network of entities enables data transmission among them, as well as collection and dissemination of data for a wide variety of purposes, as will be appreciated by those skilled in the art.
[0055] The illustrated unit layer 55 includes, as shown, the sensor 52, the actuator 53, the "learning" thermostat 56 with integrated processing, sensing and networking electronics, the camera 57, the controllable household outlet 58 and the controllable electrical switch 59. Other possible units include, but are not limited to, various additional sensor units, network units, electronic units (e.g., a remote control unit), additional actuator units, so-called "smart" appliances such as a refrigerator or a washer / dryer, and a wide variety of other possible interconnected objects.
[0056] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: mainframe computers 61; Reduced Instruction Set Computer (RISC)-based servers 62; servers 63; blade servers 64; storage units 65; and networks and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0057] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71; virtual storage 72; virtual networks 73; including virtual private networks; virtual applications and operating systems 74; and virtual clients 75.
[0058] In one example, the management layer 80 may provide the functions described below. Resource provisioning 81 enables dynamic provisioning of computing resources and other resources used to perform tasks in the cloud computing environment. Metering and pricing 82 provides cost tracking when using resources in the cloud computing environment, as well as billing or invoicing for the use of these resources. In one example, these resources may include licenses for application software. The security function provides identity verification for cloud users and tasks, as well as protection for data and other resources. A user portal 83 provides users and system administrators with access to the cloud computing environment.Service Level Management 84 provides allocation and management of cloud computing resources so that the required service level is achieved. Service Level Agreement (SLA) Planning and Fulfillment 85 provides pre-allocation and procurement of cloud computing resources whose future needs are anticipated based on a service level agreement.
[0059] The workload layer 90 provides examples of functionality for which the cloud computing environment can be used. Examples of workloads and functions that can be provided by this layer include: mapping and navigation 91; software development and lifecycle management 92; providing virtual classroom education 93; data analytics processing 94; transaction processing 95; and, in the context of the illustrated embodiments of the present invention, various workloads and functions 96 for compressing data on latency-critical processor interconnects of a data processing system in a data processing environment.Furthermore, the workloads and functions 96 for compressing data in latency-critical processor interconnections of a data processing system in a data processing environment may include operations such as interleaving and, as further described, functions for managing users and devices. Those skilled in the art will appreciate that the workloads and functions 96 for compressing data in latency-critical processor interconnections of a data processing system in a data processing environment may also interact with other portions of the various abstraction layers, such as hardware and software 60, virtualization 70, management 80, and other workloads 90 (e.g., data analytics processing 94), to achieve the various purposes of the illustrated embodiments of the present invention.
[0060] As previously mentioned, the present invention provides novel solutions for improving and optimizing the effective bandwidth and latency of such a constrained processor interconnect, reducing the overhead of the data processing processor to improve data processing efficiency and provide increased data processing efficiency. Accordingly, various embodiments are provided herein for compressing data in latency-critical processor interconnects of a data processing system in a data processing environment using a data processing processor. One or more cache lines may be dynamically compressed at the lowest level of a network stack based on one or more of a plurality of parameters prior to transferring an individual cache line, wherein the network stack includes a framer and a data link layer.This means that the cache lines can be dynamically compressed at the lowest level of the network stack (Framer / DLL) as opposed to the application layer.
[0061] In some of the examples described herein, the various embodiments for dynamically compressing cache lines are provided at the lowest level of the network stack (framer / DLL), as opposed to the application layer, where pipelined compression is performed. In some examples, data is sent from a framer to the DLL, which uses this data to determine the level of compression to achieve. One or more queue levels may be monitored within a framer.
[0062] In some of the examples described herein, the various embodiments provide for dynamic compression by dynamically turning dynamic compression on or off, which may be based on rate matching. Various embodiments provide compression operations designed to account for timing differences between interconnect and interleaving using rate matching. In some of the examples described herein, the various embodiments provide for dynamic compression by preventing timing side-channel attacks by dynamically controlling compression on hot cache lines (e.g., a cache line that is frequently transferred, e.g.,more frequently than other cache lines, or with a transfer frequency above a predetermined threshold, or whose transfer frequency is at or exceeds a defined percentage, and / or with an nth transfer frequency in a defined period of time, etc., where n is a positive integer). That is, an attack that could exploit the duration of a given operation to gain additional information about the operation itself. Thus, if a certain amount of data is observed to be transferred quickly between processors, this indicates that the data has low entropy, and the attacker can attempt to gain further information based on this low entropy.
[0063] With reference now to Fig. 4 shows a block diagram of exemplary functional components of a cache system 400 for use in compressing data in latency-critical processor interconnects of a data processing system in a data processing environment according to various mechanisms of the illustrated embodiments. In one aspect, one or more of the Fig. 1 to 3 described components, modules, services, applications and / or functions in Fig. 4. As can be seen, many of the functional blocks can also be considered as “modules” or “components” of functionality, in the same descriptive sense as previously used in the Fig. 1 to 3.
[0064] The system 400 includes a memory 410, a cache 420 (e.g., cache data array), and a cache directory 430 that can communicate with a processor 450.
[0065] Data in the cache 420 is stored in "lines," for example, memory lines 422, which are contiguous blocks of data (i.e., with a byte length corresponding to a power of two and aligned on boundaries corresponding to this size). A cache line 422 of memory data, whose length corresponds to units of 64 to 256 bytes, is stored at a position in the data array, and the respective memory address is determined as in Fig. 4. Processor 450 supplies the memory address to the cache, e.g., cache data field 420. If the address is found in cache directory 430, processor 450 can access a corresponding data line in cache 420. If the address is not found in cache directory 430, processor 450 can access a corresponding data line in memory 410.
[0066] With reference now to Fig. 5 shows a block diagram of exemplary functional components of a system 500 for compressing data in latency-critical processor connections of a data processing system in a data processing environment according to various mechanisms of the illustrated embodiments. In one aspect, one or more of the Fig. 1 to 4 described components, modules, services, applications and / or functions in Fig. 5. As can be seen, many of the functional blocks can also be considered as “modules” or “components” of functionality, in the same descriptive sense as previously used in the Fig. 1 to 4.
[0067] Shown is a data compression service 510 having a processing unit 520 ("processor") for performing various calculation, data processing, and other functions according to various aspects of the present invention. In one aspect, the processor 520 and the memory 530 may be located in and / or outside of the interleaving service 510 and may be located in and / or outside of the data processing system / data processing server 12. The interleaving service 510 may be as shown in Fig. 1 may be located within the data processing system / data processing server 12 and / or external thereto. The processing unit 520 may communicate with the memory 530. The data compression service 510 may include a compression component 540, a framer / DLL component 550 (DLL - Data Link Layer), a framing component 560, and a decompression component 570.
[0068] In one aspect, system 500 may provide virtualized computing services (i.e., virtualized data processing, virtualized storage, virtualized networking, etc.). In particular, system 500 may provide virtualized data processing, virtualized storage, virtualized networking, and other virtualized services executing on a hardware substrate.
[0069] The data compression service 510, using the compression component 540, the framer / DLL component 550, the determination component 560, and the decompression component 570, may determine to compress one or more cache lines based on one or more of a plurality of parameters, wherein the plurality of parameters include a number of outstanding data packets, a quality of service (QoS), and a current packet type, wherein the current packet type is either a data packet or a control packet.
[0070] The data compression service 510 may, using the compression component 540, the framer / DLL component 550, the commit component 560, and the decompression component 570, dynamically compress one or more cache lines at the lowest level of a network stack based on one or more of a plurality of parameters prior to transmitting a single cache line, wherein the network stack includes a framer and a data link layer.
[0071] The data compression service 510 may, using the compression component 540, the framer / DLL component 550, the setting component 560, and the decompression component 570, enable compression of the one or more cache lines based on data in a queue above a defined threshold or disable compression of the one or more cache lines based on data in a queue below a defined threshold.
[0072] The data compression service 510 may send data from the framer to the data link layer using the compression component 540, the framer / DLL component 550, the determination component 560, and the decompression component 570 to determine an amount of data to be compressed based on one or more of the plurality of parameters.
[0073] The data compression service 510, using the compression component 540, the framer / DLL component 550, the setting component 560, and the decompression component 570, may dynamically switch the compression of the one or more cache lines using a Bloom filter on a set of previously transmitted cache line addresses.
[0074] The data compression service 510 may dynamically compress one or more cache lines to match a data input rate using the compression component 540, the framer / DLL component 550, the setting component 560, and the decompression component 570.
[0075] The data compression service 510 may dynamically control the compression of the one or more cache lines based on a number of hot cache lines using the compression component 540, the framer / DLL component 550, the setting component 560, and the decompression component 570.
[0076] For further explanation, Fig. 6 is a block diagram illustrating operations for compressing data in latency-critical processor connections of a data processing system in a data processing environment according to an embodiment of the present invention. In one aspect, one or more of the Fig. 1 to 5 described components, modules, services, applications and / or functions in Fig. 6. As shown, various functional blocks are represented by arrows that illustrate the relationships between the blocks of the system 600 and the process flow (e.g., steps or operations). In addition, descriptive information about the individual functional blocks of the system 600 is also shown. As can be seen, many of the functional blocks can also be considered "modules" of functionality, in the same descriptive sense as previously described in the Fig. 1 to 5. A repeated description of similar elements used in other embodiments described herein is omitted for the sake of brevity.
[0077] In view of the above, the module blocks of system 600 can also be integrated into various hardware and software components of a data processing system, including a cloud computing environment according to the present invention. Many of the functional blocks of system 600 can be executed as background processes on various components, either in distributed computing components or in other environments.
[0078] As shown, in a first example of system 600, a framer 610A is shown in the data link layer for receiving data packets (e.g., receiving data from a processor core not shown for clarity). A queue 612A may be present in framer 610A, and queue 612A may be monitored. A set operation may be performed on queue 612A to determine if the data (e.g., a number of data packets in queue 612A) is below a threshold to disable compression. If so, a decision is made to disable compression, and a compressor 614A bypasses compression of data in queue 610A and transmits data 616A to a parser 618A, which receives the uncompressed data, e.g., 620A and 622A.In one example, the threshold can be dynamically set as the average compressibility of previous packets multiplied by the number of clock cycles required to compress a single packet.
[0079] In a second example of system 600, a framer 610A is shown in the data link layer for receiving data packets (e.g., receiving data from a processor core not shown for clarity). A queue 612B may be present in framer 610B, and queue 612B may be monitored. A set operation may be performed on queue 612B to determine if the data (e.g., a number of data packets in queue 612B) is above a threshold to enable compression. If so, a decision is made to enable compression, and the data may be sent to a compressor 614B, and a cache line may be dynamically compressed. The compressed cache lines 616B may be transferred to a parser 618B to decompress one or more cache lines, e.g., 620B and 622B.
[0080] In this way, embodiments dynamically determine how much (if at all) the next data should be compressed, allowing the performance of a computing system to be scaled by transferring more data using the same link bandwidth without impacting latency. While this change may add certain side channels, it provides one or more protections against side-channel attacks.
[0081] Fig. Figure 7 is an additional block diagram illustrating operations for compressing data in latency-critical processor interconnects of a data processing system in a data processing environment according to an embodiment of the present invention. In one aspect, one or more of the Fig. 1 to 6 described components, modules, services, applications and / or functions in Fig. 7. A repeated description of similar elements used in other embodiments described herein is omitted for the sake of brevity.
[0082] A network stack 700 is shown, which includes a framer 710, a compressor 720, and a data link layer (e.g., a data link adapter (DLA) 730, also referred to as a DLL), which may be associated with a processor core. Note that "DLA" and "DLL" may be used interchangeably. The DLA 730A may exchange data with an additional DLA, e.g., the DLA 730B associated with a decompressor 740 and a parser 750. In some implementations, the framer 710 and the DLA 730A may exchange data with one or more compression control units (KMP units).
[0083] During operation, framer 710 may receive incoming data (e.g., from a processor core). That is, the incoming packet may be 30 bytes in size and include a packet identifier (ID), which may indicate the packet type.
[0084] The framer 710 may indicate a variety of parameters to the DLL (e.g., DLA 730A) for compressing or not compressing the data. In one example, the framer 710 indicates at least three parameters to the DLA 730A: 1) a number of outstanding data packets (e.g., outstanding data), 2) a quality of service (QoS), and 3) a current packet type, where the current packet types are either a data packet or a control packet. In some examples, the delay signal from the link layer to the framer may be used to allow the DLL to control the data flow rate. In one example, the link may not be ready to accept new packets due to certain error conditions on the link, and the DLL may delay the framer.In some implementations, a modified version of the delay signal may be used to indicate that the compression engine is busy processing packets and no new packets should be fed in. A modified delay signal is thus the original DLL delay, along with any reasons the compression engine might have.
[0085] Based on these parameters, the compressor 720 can perform at least two phases: 1) a compression phase or a non-compression phase. This means that in the non-compression phase, the compressor does not attempt to compress current packet types that appear as control packets. This means that cache lines that appear as control data are not compressed. Instead, the compressor 720 forwards the data (e.g., the cache line) to the DLA 730A, assuming that control packets are difficult to compress.
[0086] Alternatively, in the compression phase, the compressor 720 attempts to compress data depending on the amount of pending data (e.g., pending data in a queue in the framer 710). For example, if the amount of pending data exceeds a threshold, the compressor 720 attempts to compress more cache lines. If the amount of pending data is below a threshold, the compressor 720 may alternatively compress data and pad it with zeros to packet size. Suppose there are three packets P0, P1, and P2. Ideally, if all three packets can be compressed together, only one packet is transmitted externally. After compressing P0 and P1, it can be observed that P2 cannot be compressed into the same packet and overflows. At this point, the compressed form of P0 and P1 (e.g.,P0+P1) can be sent and transmitted externally, along with zeros added to achieve the fixed packet size. Then P2 can be sent separately.
[0087] Compressor 720 also considers the quality of service (QoS) required for compression. Quality of service indicates how much additional latency can be added to this particular data packet (e.g., cache line) without negatively impacting the system. Higher latency indicates that compressor 720 attempts to cram more data (e.g., more cache lines) into a packet; otherwise, compressor 720 compresses, zero-pads to packet size, and discards earlier.
[0088] As described above in the P0, P1, P2 scenario, another option for the compression engine is to combine P0, P1, and P2 and attempt to increase compression rather than processing them separately. Compressing the packets individually is ideal for latency because it eliminates the need to buffer multiple packets. However, compressing the combined packets increases efficiency and can result in higher compression at the expense of latency. The quality of service factor indicates whether data can be delayed, even slightly, to identify a more efficient compression format. If latency is critical, a decision can be made to send only the packet without processing.
[0089] As an example of operations of Fig. 7, a transmission core bus (e.g. Tx side) can be enabled faster than the normal number of processor clock signals (“pclks”), e.g. 28 processor clock signals. In addition, the operations of Fig. 7 Complete a cache line transfer (256 bytes) an nth number of processor clocks earlier / faster (e.g., completing the cache line transfer 24 processor clocks earlier). Bandwidth also increases on average, e.g., by 20%. On average, each octet word ("OW") (32 bytes) reaches a number of processor clocks faster, e.g., each OW reaches 10 processor clocks faster. In one aspect, the first OW is delayed by 2 processor clocks, while all other OWs reach the processor clocks quickly.
[0090] Fig. Figure 8 is an additional block diagram illustrating operations for compressing data in latency-critical processor connections of a data processing system with compression in a framer with bus feedback (compression in a framer and decompression in a parser) in a data processing environment according to an embodiment of the present invention. In one aspect, one or more of the Fig. 1 to 7 described components, modules, services, applications and / or functions in Fig. 8. A repeated description of similar elements used in other embodiments described herein is omitted for the sake of brevity.
[0091] As shown, a compressor in a framer 810 can receive an input data stream with an identifier (n + 1) and data (n). The compressor can identify a buffer and an 8-way data buffer. Assume that the interleaving clock equals the processor core clock divided by two for the input stream speed (e.g., interleaving clock = processor core clock / 2 input stream speed) and the interconnect clock equals the processor core clock divided by three for the output stream speed (e.g., interconnect clock = processor core clock / 3 output stream speed). This means that the cache (e.g., interleaving) operates at half the speed of a processor core and the interconnect operates at one-third the speed of a processor core.
[0092] It should be noted that the calculation of the threshold can be similar to the previous case and the attempted maximum compression (e.g. 2x, 4x, 8x) can be set as the padding level divided by the output flow rate.
[0093] For a packet that can be received with 37 bytes, only twenty (20) bytes can be erased. This means that the DLL framer's padding rate is 37 bytes ("B") per interleaving cycle (37B / interleaving cycle). The DLL framer's erasure rate is twenty bytes per interleaving cycle (e.g., 20B / interleaving cycle).
[0094] A compressor of the framer (e.g., DLL compression) attempts to compress based on a rate comparison. That is, in some examples, the compressor of the framer 810 attempts to dynamically compress the cache line at a compression rate (e.g., eight times "8x") when the padding level is equal to or greater than 180 bytes. In some examples, the compressor of the framer 810 attempts to dynamically compress the cache line at a compression rate (e.g., four times "4x") when the padding level is equal to or greater than 100 bytes. In some examples, the compressor of the framer 810 attempts to dynamically compress the cache line at a compression rate (e.g., two times "2x") when the padding level is equal to or greater than 60 bytes.
[0095] In some examples, the compressor of framer 810 does not compress the data and allows the data to pass through. In this way, the various embodiments provide dynamic compression without increasing latency in cache line transfer.
[0096] The framer 810 may send the compressed or uncompressed data to the parser 820, and the decompressor of the parser 820 may decompress the data.
[0097] Further by way of example, in various embodiments for which dynamic compression and transfer ("Tx") is provided, a release per cache line 9 processor clocks earlier is achieved. When receiving ("Rx"), the transfer of cache lines completes 4 processor clocks earlier (and the nominal latency for 1 cache line, assuming everything else is idle). Furthermore, when transferring zeros to two OWs, no additional latency is encountered in the output from the parser 820. A third OW may only be delayed by 2 processor clocks. The transfer of OWs 4 through 7 may occur 1, 2, 3, or 4 processor clocks earlier, respectively. Thus, on average across an isolated cache line, each OW arrives 1 processor clock earlier. On average across multiple transactions, each OW arrives at least 7 processor clocks earlier.
[0098] Additionally, due to compressed cache lines, an attacker can launch an attack in which sensitive lines are constantly flushed and reloaded. Based on the time required for flushing and reloading, the attacker can determine line compressibility, thereby losing entropy. To overcome this vulnerability, the present invention can implement a Bloom filter over a set of previously transferred cache line addresses. If the same cache line moves back and forth too often (programmable threshold), compression can be disabled (e.g., dynamic compression is turned off). Dynamic compression can be turned back on if a computing system does not have a hot cache line that is flushed / reloaded too often. Compression can also be turned back on after a defined period of time (e.g., dynamic compression is turned back on).
[0099] With reference now to Fig. 9 illustrates a method 900 for compressing data in latency-critical processor interconnections of a data processing system in a data processing environment using a processor in which various aspects of the illustrated embodiments may be implemented. Functionality 900 may be implemented as a method (e.g., a computer-implemented method) executing on a machine as instructions, wherein the instructions are embodied on at least one computer-readable medium or a non-transitory machine-readable storage medium. Functionality 900 may begin in block 902.
[0100] One or more cache lines may be dynamically compressed at a lowest level of a network stack based on one or more of a plurality of parameters prior to transmitting a single cache line, wherein the network stack includes a framer and a data link layer, as in block 904. Functionality 900 may terminate in block 906.
[0101] In one aspect, the operations of method 900 may be performed in conjunction with and / or as part of at least one of the blocks of Fig. 9 The operations of method 900 may include the following: enabling compression of the one or more cache lines based on data in a queue above a defined threshold or disabling compression of the one or more cache lines based on data in a queue below a defined threshold.
[0102] The operations of method 900 may determine to compress the one or more cache lines based on one or more of the plurality of parameters, wherein the plurality of parameters includes a number of outstanding data packets, a quality of service ("QoS"), and a current packet type, wherein the current packet type is either a data packet or a control packet.
[0103] The operations of method 900 may send data from the framer to the data link layer to determine an amount of data to be compressed based on the one or more of the plurality of parameters. The operations of method 900 may dynamically switch compression of the one or more cache lines using a Bloom filter on a set of previously transmitted cache line addresses. The operations of method 900 may dynamically compress the one or more cache lines to match a data input rate. The operations of method 900 may dynamically control compression of the one or more cache lines based on a number of hot cache lines.
[0104] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions stored thereon for causing a processor to perform aspects of the present invention.
[0105] The computer-readable storage medium may be any physical device that can retain and store instructions for use by an instruction-executing device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM).Flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded device such as punched cards or raised structures in a groove on which instructions are stored, and any suitable combination thereof. A computer-readable storage medium, as used herein, shall not be construed as carrying transient signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., pulses of light traveling through an optical fiber), or electrical signals transmitted through a wire.
[0106] Computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing units or to an external computer or external storage unit via a network such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmission lines, wireless transmission, routers, firewalls, switching units, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing unit receives computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing unit.
[0107] Computer-readable program instructions for performing operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, or the like.as well as conventional procedural programming languages such as the C programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a standalone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer by any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, over the Internet using an Internet service provider).In some embodiments, electronic circuits, including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuits to perform aspects of the present invention.
[0108] Aspects of the present invention are described herein with reference to flowchart and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart and / or block diagrams, as well as combinations of blocks in the flowchart and / or block diagrams, may be implemented by computer-readable program instructions.
[0109] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, produce a means for implementing the functions / steps defined in the flowchart and / or block diagram block(s).These computer-readable program instructions may also be stored on a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored thereon comprises an article of manufacture, including instructions that implement aspects of the function / step specified in the block(s) of the flowchart and / or block diagrams.
[0110] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of process steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-executable process such that the instructions executing on the computer, other programmable apparatus, or other device implement the functions / steps defined in the block(s) of flowchart and / or block diagrams.
[0111] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions comprising one or more executable instructions for performing the particular logical function(s). In some alternative implementations, the functions specified in the block may occur in a different order than shown in the figures. For example, two blocks shown in succession may actually execute substantially concurrently, or the blocks may sometimes execute in reverse order depending on the corresponding functionality.It is further to be understood that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by special purpose hardware-based systems that perform the specified functions or steps, or by combinations of special purpose hardware and computer instructions.
[0112] The descriptions of the various embodiments of the present invention have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments. It will be apparent to those skilled in the art that many changes and modifications are possible without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, practical application, or technical improvement over technologies on the market, or to enable those skilled in the art to understand the embodiments described herein.
Claims
[1] A method for compressing data in latency-critical processor connections of a data processing system in a data processing environment by one or more processors, the method comprising: dynamically compressing one or more cache lines at a lowest level of a network stack based on one or more of a plurality of parameters prior to transmitting a single cache line, the network stack comprising a framer and a data link layer. [2] The method of claim 1, further comprising: Enabling compression of the one or more cache lines based on data in a queue above a defined threshold; or Disable compression of one or more cache lines based on data in a queue below a defined threshold. [3] The method of claim 1, further comprising determining to compress the one or more cache lines based on the one or more of the plurality of parameters, the plurality of parameters comprising a number of outstanding data packets, a quality of service ("QoS"), and a current packet type, the current packet type being one of a data packet and a control packet. [4] The method of claim 3, further comprising sending data from the framer to the data link layer to determine an amount of data to be compressed based on the one or more of the plurality of parameters. [5] The method of claim 1, further comprising dynamically switching the compression of the one or more cache lines using a Bloom filter on a set of previously transmitted cache line addresses. [6] The method of claim 1, further comprising dynamically compressing the one or more cache lines to adapt to a data input rate. [7] The method of claim 1, further comprising dynamically controlling compression of the one or more cache lines based on a number of hot cache lines. [8] A system for compressing data in latency-critical processor connections of a data processing system in a data processing environment, the system comprising: one or more computers with executable instructions that, when executed, cause the system to: dynamically compress one or more cache lines at a lowest level of a network stack based on one or more of a plurality of parameters prior to transmitting a single cache line, the network stack comprising a framer and a data link layer. [9] The system of claim 8, wherein the executable instructions, when executed, cause the system to: to enable compression of one or more cache lines based on data in a queue above a defined threshold; or to disable compression of one or more cache lines based on data in a queue below a defined threshold. [10] The system of claim 8, wherein the executable instructions, when executed, cause the system to determine to compress the one or more cache lines based on the one or more of the plurality of parameters, the plurality of parameters comprising a number of outstanding data packets, a quality of service ("QoS"), and a current packet type, the current packet type being one of a data packet and a control packet. [11] The system of claim 10, wherein the executable instructions, when executed, cause the system to send data from the framer to the data link layer to determine an amount of data to be compressed based on the one or more of the plurality of parameters. [12] The system of claim 8, wherein the executable instructions, when executed, cause the system to dynamically switch compression of the one or more cache lines using a Bloom filter on a set of previously transmitted cache line addresses. [13] The system of claim 8, wherein the executable instructions, when executed, cause the system to dynamically compress the one or more cache lines to match a data input rate. [14] The system of claim 8, wherein the executable instructions, when executed, cause the system to dynamically control compression of the one or more cache lines based on a number of hot cache lines. [15] A computer program product for compressing data in latency-critical processor connections of a data processing system in a data processing environment, the computer program product comprising: one or more computer-readable storage media and program instructions stored together on the one or more computer-readable storage media, the program instruction comprising: Program instructions to dynamically compress one or more cache lines at a lowest level of a network stack based on one or more of a plurality of parameters prior to transferring a single cache line, the network stack comprising a framer and a data link layer. [16] The computer program product of claim 15, further comprising program instructions to: to enable compression of one or more cache lines based on data in a queue above a defined threshold; or to disable compression of one or more cache lines based on data in a queue below a defined threshold. [17] The computer program product of claim 15, further comprising program instructions to specify: compress the one or more cache lines based on the one or more of the plurality of parameters, the plurality of parameters comprising a number of outstanding data packets, a quality of service (“QoS”), and a current packet type, the current packet type being either a data packet or a control packet; and Send data from the framer to the data link layer to determine an amount of data to be compressed based on the one or more of the plurality of parameters. [18] The computer program product of claim 15, further comprising program instructions to dynamically switch compression of the one or more cache lines using a Bloom filter on a set of previously transmitted cache line addresses. [19] The computer program product of claim 15, further comprising program instructions to dynamically compress the one or more cache lines to adapt to a data input rate. [20] The computer program product of claim 15, further comprising program instructions to dynamically control compression of the one or more cache lines based on a number of hot cache lines.