Improving the Performance and Reliability of Processor Store Operation Data Transfer

By predicting available entries and adjusting request delays, the system enhances data transfer efficiency and reliability in processor systems by optimizing communication protocols and reducing rejected transfers.

JP2025522914APending Publication Date: 2025-07-17INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025500290
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-21
Filing Date
2023-07-10
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Current processor systems face challenges in improving the performance and reliability of store operation data transfer between the processor core and cache memory due to desynchronization of communication protocols and competition for cache entries among multiple threads, leading to rejected data transfers and reduced efficiency.

Method used

The system predicts available entries in the cache store queue, sends confirmation responses, and adjusts request delays based on queue occupancy to enhance data transfer efficiency and reliability.

Benefits of technology

This approach improves data transfer performance by minimizing rejected transfers and optimizing communication protocols, ensuring efficient and reliable data handling between the processor core and cache memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522914000001_ABST
    Figure 2025522914000001_ABST
Patent Text Reader

Abstract

Improving the Performance and Reliability of Processor Store Operation Data Transfer The processor includes a load and store unit (LSU) and a cache memory, and transfers data information from the store queue in the LSU to the cache memory. When the cache memory determines that there is an available entry in the store queue within the cache memory, the cache memory requests an information packet from the LSU. The LSU acknowledges the request and transfers the information packet to the cache memory. Before receiving an additional request from the cache memory, the LSU predicts that there is an additional available entry in the cache memory, sends an additional acknowledgment response to the cache memory, and transfers an additional information packet. If there is an additional available entry in the cache store queue of the cache memory, the cache memory stores the additional information packet. If there is no additional available entry in the cache store queue of the cache memory, the cache memory rejects the additional information packet. Subsequently, if the cache memory later requests the additional information packet, the LSU must retry the transfer of the additional information packet. The cache memory can set or reset the time delay when requesting subsequent information packets based on multiple factors within the cache memory and the processor, including the number of available entries in the cache store queue. Corresponding methods and computer program products are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to computer processor systems, and more specifically, to systems and methods for improving the performance and reliability of store operation data transfer between a processor and a processor cache memory.

Background Art

[0002] In a computer processing system, a processor functions as a central processing unit and performs many important functions such as arithmetic and logical operations, execution of program instructions, data transfer, and other processing system control and logical functions. The processor includes a cache that assists in the transfer of data between the processor core and the processing system memory. The cache typically has multiple levels or hierarchies, including a relatively small and faster level-one (L1) cache and a relatively larger and slower level-two (L2) cache. The processor includes a relatively small storage queue for temporarily storing data, and each level of the cache includes a queue for holding data before storing it in a larger cache storage.

[0003] There is a communication protocol between the processor cache (L2) store queue and the processor core store queue for controlling data transfer between the processor core and the cache. The processor cache indicates that an area in its storage is available by sending a request signal to the processor core for each available entry in the cache queue. The processor core responds to the cache with an acknowledgement response signal and transfers a data packet from the processor queue to the cache. The cache receives the data packet along with the acknowledgement response and stores the data packet in the cache store queue. The LSU waits until the cache sends an additional request before sending an additional data packet to the cache.

[0004] In view of the above, what is needed is a system and method for improving data transfer performance and reliability between a processor core and a processor cache. The processor system attempts to improve processor performance by predicting that there is an empty entry in the processor cache store queue. In addition, multiple threads within the processor core share the processor cache, competing for entries into the cache store queue, causing a desynchronization of the communication protocol between the processor core store queue and the processor cache store queue. The present invention prevents the processor cache from rejecting data transfers and minimizes desynchronization of the communication protocol between the processor core and the processor cache store queue. SUMMARY OF THE INVENTION

[0005] The present invention has been developed in response to current state-of-the-art technologies, and in particular, in response to problems and needs in the art that have not yet been fully solved by currently available systems and methods. The features and advantages of the present invention will become more fully apparent from the following description and the appended claims, or may be learned by the practice of the invention as set forth hereinafter.

[0006] According to one embodiment of the invention described in this specification, a processor is provided for improving the performance of store operation data transfer in a computer processing system. In one embodiment, the computer processing system includes a processor for executing computer processing system functions, a memory, and a plurality of components. In one embodiment, the processor has a load and store unit (LSU) and a cache memory for storing data information to be transferred to the memory and / or other components within the computer processing system. In one embodiment, the LSU has a store queue including a plurality of entries for storing a plurality of information packets. In one embodiment, the cache memory has a store queue including a plurality of entries for storing a plurality of information packets. In one embodiment, the cache memory determines that the cache store queue includes available entries. The cache memory transmits a request to transfer an information packet to the cache memory to the LSU. In one embodiment, the LSU transmits an acknowledgment response in response to the cache request and transfers the information packet from an entry in the LSU store queue to the cache memory. In one embodiment, the cache memory receives the information packet from the LSU and stores the information packet in an available entry in the cache store queue.

[0007] In one embodiment, the LSU predicts that there are additional available entries in the cache memory, sends an additional confirmation response signal to the cache memory, and transfers an additional information packet from the additional entry in the LSU store queue to the cache memory. The cache memory determines that there are additional available entries in the cache store queue, receives the information packet from the LSU, and stores the information packet in the additional available entry in the cache store queue. In one embodiment, the cache memory delays subsequent requests for subsequent information packets, where the subsequent requests also function as a confirmation response to the LSU that the cache has successfully stored the additional information packet in the cache store queue. In one embodiment, the cache memory determines that there are no additional available entries in the cache store queue and rejects the additional information packet transferred from the LSU. The LSU must wait for the cache memory to send a new request before retrying the transfer of additional information from the LSU store queue to the cache memory.

[0008] In one embodiment, the cache memory calculates a time delay (subsequent request delay) for sending subsequent requests for subsequent information packets to the LSU based on the number of available entries in the cache store queue. In one embodiment, the cache memory calculates the subsequent request delay based on the average recent time for transferring information packets from the LSU store queue to the cache store queue. In one embodiment, the cache memory sets and resets the subsequent request delay based on a threshold value, where the threshold value is based on the number of available entries in the cache store queue. The cache memory sets the subsequent request delay to a determined time interval when the number of available entries in the cache store queue is less than the threshold value. The cache memory resets the subsequent request delay to no time delay when the number of available entries in the cache store queue is greater than or equal to the threshold value.

[0009] According to another embodiment of the invention described herein, a method for improving the performance of store operation data transfer within a computer processing system is provided, where the computer processing system comprises a processor and a cache memory, the processor has a load and store unit (LSU) including a store queue, and the cache has a store queue. In one embodiment, the method comprises the step of storing an information packet in an entry within the LSU store queue. In one embodiment, the method comprises the step of the cache memory determining that there is an available entry within the cache store queue and requesting the information packet from the LSU. In one embodiment, the method comprises the step of the LSU acknowledging the request from the cache memory and transferring the information packet from the entry within the LSU store queue to the cache memory. In one embodiment, the method comprises the step of the cache memory receiving the information packet from the LSU and storing the information packet in an available entry within the cache store queue.

[0010] In one embodiment, the method comprises the step of the LSU predicting that the cache memory has an additional available entry in the cache store queue. In one embodiment, the method comprises the step of the LSU sending an additional confirmation response to the cache memory and transferring an additional information packet to the cache memory before the cache memory requests the additional information packet. In one embodiment, the method comprises the step of the cache memory determining that there is an additional available entry in the cache store queue, receiving the additional information packet from the LSU, and storing the additional information packet in the additional available entry in the cache store queue. In one embodiment, the method comprises the step of the cache memory delaying a subsequent request for a subsequent information packet to the LSU, where the subsequent request functions as a confirmation response indicating that the additional information packet is stored in the cache memory. In one embodiment, the method alternatively comprises the step of the cache memory determining that there is no additional available entry in the cache store queue, rejecting the transfer of the additional information packet from the LSU, and thereby requesting the LSU to retry the transfer of the additional information packet if the LSU receives another request from the cache memory.

[0011] In one embodiment, the method comprises the step of a cache memory calculating a time delay (subsequent request delay) for sending a subsequent request for a subsequent information packet to an LSU based on the number of available entries in a cache store queue. In one embodiment, the method comprises the step of a cache memory calculating the subsequent request delay based on the most recent average time for transferring an information packet from an LSU store queue to a cache store queue. In one embodiment, the method comprises the step of a cache memory setting and resetting the subsequent request delay based on a threshold value, where the threshold value is based on the number of available entries in a cache store queue. The method comprises the step of a cache memory setting the subsequent request delay to a determined time interval if the number of available entries in a cache store queue is less than the threshold value. The method comprises the step of a cache memory resetting the subsequent request delay to no time delay if the number of available entries in a cache store queue is greater than or equal to the threshold value.

[0012] According to another embodiment of the invention described herein, a computer program product is provided for improving the performance of store operation data transfer within a computer processing system, where the computer processing system comprises a processor and a cache memory, the processor has a load and store unit (LSU) including a store queue, and the cache has a store queue. In one embodiment, the computer program product comprises a non-transitory computer-readable storage medium having computer-usable program code embodied therein. In one embodiment, the computer-usable program code is configured to perform operations when executed by the processor. In one embodiment, the operations of the computer program product comprise procedures for storing an information packet in an entry of the LSU store queue. In one embodiment, the operations of the computer program product comprise procedures for the cache memory to determine that there is an available entry in the cache store queue and request an information packet from the LSU. In one embodiment, the operations of the computer program product comprise procedures for the LSU to acknowledge a request from the cache memory and transfer an information packet from an entry in the LSU store queue to the cache memory. In one embodiment, the operations of the computer program product comprise procedures for the cache memory to receive an information packet from the LSU and store the information packet in an available entry in the cache store queue.

[0013] In one embodiment, the operation of the computer program product comprises a procedure in which the LSU predicts that the cache memory has an additional available entry in the cache store queue. In one embodiment, the operation of the computer program product comprises a procedure in which the LSU sends an additional confirmation response to the cache memory and transfers an additional information packet to the cache memory before the cache memory requests the additional information packet. In one embodiment, the operation of the computer program product comprises a procedure in which the cache memory determines that there is an additional available entry in the cache store queue, receives an additional information packet from the LSU, and stores the additional information packet in the additional available entry in the cache store queue. In one embodiment, the operation of the computer program product comprises a procedure in which the cache memory delays a subsequent request for a subsequent information packet to the LSU, where the subsequent request functions as a confirmation response indicating that the additional information packet is stored in the cache memory. In one embodiment, the operation of the computer program product alternatively comprises a procedure in which the cache memory determines that there is no additional available entry in the cache store queue, rejects the transfer of the additional information packet from the LSU, and thereby requests the LSU to retry the transfer of the additional information packet if the LSU receives another request from the cache memory.

[0014] In one embodiment, the operation of the computer program product includes a procedure in which the cache memory calculates a time delay (subsequent request delay) for sending subsequent requests for subsequent information packets to the LSU based on the number of available entries in the cache store queue. In one embodiment, the operation of the computer program product includes a procedure in which the cache memory calculates the subsequent request delay based on the most recent average time for transferring information packets from the LSU store queue to the cache store queue. In one embodiment, the operation of the computer program product includes a procedure in which the cache memory sets and resets the subsequent request delay based on a threshold value, where the threshold value is based on the number of available entries in the cache store queue. The operation of the computer program product includes a procedure in which the cache memory sets the subsequent request delay to a determined time interval when the number of available entries in the cache store queue is less than the threshold value. The operation of the computer program product includes a procedure in which the cache memory resets the subsequent request delay without a time delay when the number of available entries in the cache store queue is greater than or equal to the threshold value.

Brief Description of the Drawings

[0015] To more easily understand the advantages of the present invention, a more specific description of the present invention briefly described above is presented with reference to specific embodiments shown in the accompanying drawings. It is understood that these drawings show only typical embodiments of the present invention and are not considered to limit its scope. Embodiments of the present invention are described and explained with additional specificity and detail through the use of the accompanying drawings.

[0016]

Figure 1

[0017]

Figure 2

[0018]

Figure 3

[0019]

Figure 4

[0020] As generally described and illustrated in the figures of this specification, it will be readily understood that the components of the present invention can be arranged and designed in a wide variety of different configurations. Accordingly, as shown in the figures, the following more detailed description of embodiments of the present invention is not intended to limit the scope of the claimed invention, but is merely representative of specific examples of the embodiments currently contemplated in accordance with the present invention. The presently described embodiments will be best understood by reference to the drawings, in which like parts are designated by like numerals throughout.

[0021] Exemplary embodiments are described herein for improving the performance and reliability of store operation data transfer within a computer processing system. The computer processing system includes one or more processors, memory, and a plurality of components for performing computer processing functions and control. The processor has a load and store unit (LSU) and cache memory. The LSU and cache memory have a store queue that includes entries for storing information packets. The LSU transfers an information packet to the cache memory, and the information packet is stored in the cache memory until the information needs to be transferred from the cache memory to the main memory and / or other components within the computer processing system. When the cache memory determines that there is an available entry in the cache store queue, the cache memory requests the information packet from the LSU. The LSU acknowledges the request and transfers the information packet from the LSU store queue to the cache memory. The LSU speeds up data transfer by predicting that the cache memory has an available entry in the cache store queue and sends an additional acknowledgment and additional information packet to the cache memory. The cache memory accepts the additional information packet if it has an available entry in the cache store queue. Alternatively, the cache memory rejects the transfer of the additional information packet if it does not have an additional entry in the cache store queue. The LSU must then retry the transfer of the additional information packet. The cache memory delays sending subsequent requests to the LSU for subsequent information packets in order to avoid asking the cache to reject the additional information packet transferred from the LSU and asking the LSU to retry the transfer of the additional information packet.

[0022] Referring to FIG. 1, a computer processing system 100 according to one embodiment is generally shown. The computer processing system 100 can be an electronic computer framework that includes and / or employs any number and combination of computing devices and networks that utilize various communication technologies, as described herein. In certain embodiments, the computer processing system 100 can be easily scalable, extensible, and modular, and has the ability to change for different services or to reconfigure some functions independently of other functions. In certain embodiments, the computer processing system 100 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smartphone. Additionally, the computer processing system 100 can be a cloud computing node. In certain embodiments, the computer processing system 100 can be described in the general context of computer system executable instructions, such as program modules, being executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like that perform specific tasks or implement specific abstract data types. In certain embodiments, the computer processing system 100 can be practiced in a distributed cloud computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed cloud computing environment, program modules can be located in both local computer system storage media, including memory storage devices, and remote computer system storage media.

[0023] As shown in FIG. 1, the computer processing system 100 has one or more central processing units (CPUs) 101 (collectively or generally referred to as the processor 101). In certain embodiments, the processor 101 can be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The processor 101, also referred to as the processing circuit, can also include one or more accelerators (e.g., a graphics processing unit, i.e., GPU). In one embodiment, the processor has a cache 109 and a controller 108 to assist in performing processor-related functions such as arithmetic and logical operations, execution of program instructions, data transfer, and other processing system control and logical functions. In certain embodiments, the processor 101 is coupled to the system memory 103 and various other components via a system bus 102. The system memory 103 can include a read-only memory (ROM) 104 and a random access memory (RAM) 105. The ROM 104 is coupled to the system bus 102 and can include a basic input / output system (BIOS) that controls certain basic functions of the computer system 100. The RAM is a read-write memory coupled to the system bus 102 for use by the processor 101. In certain embodiments, the system memory 103 provides a temporary memory area for the operation of the above instructions during operation. The system memory 103 can include a random access memory (RAM), a read-only memory, a flash memory, or any other suitable memory system.

[0024] In certain embodiments, computer processing system 100 has an input / output (I / O) adapter 106 and a communication adapter 107 coupled to system bus 102. The I / O adapter 106 can be a small computer system interface (SCSI) adapter that communicates with a hard disk 108 and / or any other similar components. The I / O adapter 106 and flash memory (DRAM) 118 and / or hard disk drive 118 are collectively referred to herein as mass storage 110. In certain embodiments, software 111 for execution on computer processing system 100 can be stored in mass storage 110. Mass storage 110 is an example of a tangible storage medium readable by processor 101, where software 111 is stored as instructions for execution by processor 101 to operate computer system 100 as described hereinafter in this specification with respect to various figures. Examples of computer program products and execution of such instructions are discussed in more detail herein.

[0025] In certain embodiments, communication adapter 107 interconnects system bus 102 with a network 112, which can be an external network, enabling computer processing system 100 to communicate with other systems. In one embodiment, a portion of system memory 103 and mass storage 110 stores collectively an operating system, which can be any suitable operating system, such as z / OS or AIX operating system by IBM Corporation, for coordinating the functions of the various components shown in FIG. 1.

[0026] In certain embodiments, additional input / output devices are connected to system bus 102 via display adapter 115 and interface adapter 116. In one embodiment, adapters 106, 107, 115, and 116 may be connected to one or more I / O buses connected to system bus 102 via an intermediate bus bridge (not shown). In one embodiment, a display 119 (e.g., a display screen or monitor) is connected to system bus 102 through a display adapter 115 that may include a graphics controller and a video controller to improve the performance of graphics-intensive applications. In one embodiment, a keyboard 121, a mouse 122, a speaker 123, and / or other devices may be interconnected to system bus 102 via an interface adapter 116 that may include, for example, a Super I / O chip that integrates adapters for multiple devices into a single integrated circuit. In certain embodiments, suitable I / O buses for connecting peripheral devices such as hard disk controllers, network adapters, and graphics adapters typically include a common protocol such as Peripheral Component Interconnect (PCI). Thus, as configured in FIG. 1, computer processing system 100 has processing capabilities in the form of a processor 101, storage capabilities including system memory 103 and mass storage 110, input means such as a keyboard 121 and a mouse 122, and output capabilities including a speaker 123 and a display 119.

[0027] In certain embodiments, communication adapter 107 can transmit data using any suitable interface or protocol, such as, among other things, an Internet small computer system interface. Network 112 can be, among other things, a cellular network, a wireless network, a wide area network (WAN), a local area network (LAN), or the Internet. An external computing device can be connected to computing system 100 through network 112. In some embodiments, the external computing device can be an external web server or a cloud computing node.

[0028] It should be understood that the block diagram of FIG. 1 is not intended to show that computer processing system 100 should include all of the components shown in FIG. 1. Rather, computer processing system 100 can include any suitable fewer or additional components not shown in FIG. 1 (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). Further, the embodiments described herein with respect to computer processing system 100 can be implemented with any suitable logic, which, when referred to herein, can include, in various embodiments, any suitable hardware (e.g., among other things, a processor, an embedded controller, or an application specific integrated circuit), software (e.g., among other things, an application), firmware, or any suitable combination of hardware, software, and firmware.

[0029] Figure 2 represents a block diagram of a processor, or a central processing unit 101, according to an embodiment of the present invention. In one embodiment, the processor includes one or more processor cores 200, 201. In one embodiment, the processor cores 200, 201 have load-store units (LSU) 210, 211, and multiple levels of cache memory 109. In one embodiment, the cache memory 109 for each processor core consists of level-one (L1) caches 220, 221 and level-two (L2) caches 230, 231. The level-one (L1) caches 220, 221 typically include memories with the fastest access time and are smaller in size. The level-two (L2) caches 230, 231 typically include memories with a faster access time that is slower than that of the L1 caches 220, 221 and are much larger in size. As described above, the processor cores 200, 201 execute processor-related functions such as arithmetic and logical operations, execution of program instructions, data transfer, and other processing system control and logical functions. To execute these functions, the processor cores 200, 201 need to access information, instructions, and / or data from the processor memory 103 through the system bus 102. In one embodiment, the processor cores 200, 201 utilize the load-store units 210, 211 and the multi-level cache memories 220, 221, 230, 231 to transfer information more efficiently and quickly. Information including instructions and data can be prefetched from the system memory 103 based on an algorithm within the processor core 200 that predicts which information is likely to be accessed from the system memory 103 and stored in the L1 cache 220 or the L2 cache 230.

[0030] It should be understood that the block diagram of FIG. 2 is not intended to show all components included in the processor cores 200 and 201. Although FIG. 2 includes two processor cores 200 and 201, this is only intended to illustrate an exemplary embodiment. The central processing unit 101 may include a single processor core 200, or multiple (more than two) processor cores 200 and 201. As described above, the processor cores 200 and 201 may execute many processing-related functions and may include additional components and control logic not shown in FIG. 2 to assist in the execution of such functions. In addition, specific embodiments of the processor cores 200 and 201 may include more than two levels of cache memory 108. In one embodiment, the processor cores 200 and 201 include a level-three (L3) cache and a level-four (L4) cache to boost the performance and efficiency of accessing information from the system memory 103. Furthermore, the embodiments described herein with respect to the processor cores 200 and 201 may be implemented in any suitable logic, and when referred to herein, the logic may include any suitable combination of hardware (e.g., among others, a processor, an embedded controller, or an application-specific integrated circuit), software (e.g., among others, an application), firmware, or any suitable combination of hardware, software, and firmware in various embodiments.

[0031] FIG. 3 shows a block diagram of a processor core according to an embodiment of the present invention. In one embodiment, the processor core 200 includes a load store unit (LSU) 210, an L1 cache memory 220, and an L2 cache memory 230. In one embodiment, the LSU 210 includes a store queue 250 and a load queue 255 for transmitting and receiving information between the L1 cache memory 220 and the L2 cache memory 230. The information consists of data to be transferred, instructions to be executed, addresses to be accessed, or other information necessary for the processor 200 to perform its operations. In one embodiment, the L2 cache 230 includes a corresponding store queue 260 and a load queue 265. In one embodiment, the LSU 210 is coupled to the L1 cache 220 and the L2 cache 230 for transferring and receiving information. Similarly, the L2 cache 260 is coupled to the LSU 210 and the L1 cache 220 for transferring and receiving information. In one embodiment, the L2 cache 230 is coupled to the system bus 102 for transferring and receiving information from components within the processor system 100. In one embodiment, the LSU 210, the L1 cache 220, and the L2 cache 230 include logic for communicating between components and controlling the transfer of information between components.

[0032] The block diagram of FIG. 3 is not intended to show all components included in the processor core 200, but only to show an exemplary embodiment of components within the processor core 200 that facilitate the transfer of information within the processor core 200 and within the processor system 100. It should be understood that FIG. 3 includes a single processor core 200 having a single LSU 210, a single L1 cache 220, and a single L2 cache 230. In alternative embodiments, the L1 cache 220 and the L2 cache 230 may be divided into multiple cache segments. In such cases, separate LSUs 210 may be implemented to communicate with and control the transfer of information with the segments of the divided L1 cache 220 and L2 cache 230. Additionally, the embodiment described in FIG. 3 includes control signals for the communication and transfer of data among the LSU 210, the L1 cache 220, and the L2 cache 230. Such control signals may be implemented with any suitable logic, and when referred to herein, the logic may include any suitable combination of hardware (e.g., among other things, a processor, an embedded controller, or an application-specific integrated circuit), software (e.g., among other things, an application), firmware, or any suitable combination of hardware, software, and firmware in various embodiments.

[0033] FIG. 4 depicts a block diagram of an improved system and computer implemented method for transferring data information within computer processing system 100 using LSU 210 and L2 cache memory 230 within processor core 200. In one embodiment, the processor core includes LSU 210 coupled to L2 cache 230. In one embodiment, the LSU includes a store queue 250 for storing data information and a control module 270 for transmitting control and communication signals. In one embodiment, the L2 cache 230 includes a store queue for storing data information and a control module 280 for transmitting control and communication signals. In one embodiment, the LSU control module 270 is coupled to the L2 cache control module 280 through control bus 275 and data bus 285. In one embodiment, the LSU control module 270 and the L2 cache control module include logic, which may include hardware, software, firmware, or a combination of hardware, software, and firmware.

[0034] In one embodiment, when an entry in the L2 store queue 260 is available, the L2 cache control module 280 transmits a POP signal to the LSU control module 270 along the control bus 275. In one embodiment, when the LSU control module 270 is ready for data information to be transferred from the LSU store queue 250 to the L2 cache store queue 260, it transmits a PUSH signal along the control bus 275. The LSU 210 then starts transferring data information from the LSU store queue 250 to the L2 cache store queue 260 through the data bus 285 until the data transfer is complete. In one embodiment, the L2 cache control module 280 transmits successive POP signals along the control bus 275 to indicate that additional entries in the L2 cache store queue 260 are available. Next, the LSU control module 270 transmits a PUSH signal along the control bus 275 to indicate that it is ready for additional data information to be transferred from the LSU store queue 250 to the L2 cache store queue 260, and this process is repeated as the data information transfer from the LSU store queue 250 to the L2 cache store queue 260 is initiated through the data bus 285.

[0035] In a further embodiment, the LSU 210 speeds up the transfer of data information to the L2 cache 230 by predicting that the L2 cache store queue has additional available entries. In one embodiment, the LSU control module 270 sends a PUSH signal along the control bus 275 and starts the transfer of data information from the LSU store queue 250 to the L2 cache store queue 260 in response to a POP signal from the L2 cache control module 280. In one embodiment, the LSU control module 270 sends an additional PUSH signal to the L2 cache control module 280 along the control bus 275 when the previous data information transfer is complete and before receiving an additional POP signal from the L2 cache control module 280. As before, the LSU control module 270 starts a subsequent data information transfer from the LSU store queue 250 to the L2 cache store queue 260 along the data bus 285. Thus, the LSU 210 predicts that the L2 cache 230 contains additional available entries in the L2 cache store queue 260 and enhances the transfer speed and performance between the LSU store queue 250 and the L2 cache store queue 260. In one embodiment, the L2 cache control module 280 sends a POP signal indicating that the L2 cache control store 260 contains available entries along the control bus 275 to the LSU control module 270 and receives data information transfer from the LSU store queue 250. In one embodiment, the L2 cache control module 280 sends a BOUNCE signal indicating that the L2 cache control store 260 does not contain available entries along the control bus 275 to the LSU control module 270 and rejects the data information transfer from the LSU store queue 250. In this case, the LSU 210 has to wait for an available entry in the L2 cache 230 to retry the data information transfer, and thus the LSU control module 270 has to wait to receive a subsequent POP signal from the L2 cache control module 280 along the control bus 275.Therefore, the LSU incorrectly predicts that there are available entries in the L2 cache store queue 260, reducing the store queue transfer performance between the LSU 210 and the L2 cache 230 within the processor core 200. In addition, if the LSU 210 releases the data information transferred in a subsequent PUSH signal from the LSU store queue 250 before the LSU control module 270 receives a BOUNCE signal from the L2 cache control module 280, the risk of data information loss increases.

[0036] In a further embodiment, the L2 cache 230 may delay indicating to the LSU 210 that an entry in the L2 cache store queue 260 is available. As described above, the L2 cache control module 280 transmits a POP signal when an entry in the L2 cache store queue 260 becomes available. In one embodiment, the L2 delays transmitting the POP signal when an entry in the L2 cache store queue 260 becomes available. In one embodiment, the delay may be a time delay over an interval of fixed duration. In one embodiment, the delay may be a time delay whose length may vary based on certain factors and / or metrics measured within the processor core 200. In one embodiment, the delay may be calculated based on the time interval for transferring data information from the LSU store queue 250 to the L2 cache store queue 260. In one embodiment, the delay may be a time interval adjusted based on the number of available entries in the L2 cache store queue 260. In one embodiment, the delay may be calculated using an algorithm based on a combination of factors and metrics within the processor core, including but not limited to the number of available entries in the L2 store queue 260, the free space within the L2 cache 230 or the L1 cache 220, the average time for transferring data information between the LSU 230 and the L2 cache 250, and the frequency of store operations within the processor core. By delaying the transmission of the POP signal from the L2 cache control module 280 to the LSU control module 270, the L2 cache control module 280 ensures that the L2 cache store queue 260 can successfully complete additional data information transfers when the LSU 210 attempts to speed up data information transfer to the L2 cache 230 by predicting that the L2 store queue 260 contains additional available entries.

[0037] In a further embodiment, the delay can be a fixed time interval that turns on and off, or is set and reset, based on a threshold, or a variable time interval that is periodically calculated based on factors and metrics measured within the processor core 200. In an exemplary embodiment, the delay is set or reset based on the number of available entries in the L2 store queue 260. The L2 cache control module 280 turns on or sets the delay when the number of available entries in the L2 cache store queue is less than a threshold, and turns off or resets the delay when the number of available entries in the L2 cache store queue exceeds the threshold. As an example, the L2 cache 230 delays sending a POP signal to the LSU 210 when the number of available entries in the L2 store queue 260 is less than 4 entries, and does not delay sending a POP signal to initiate a data information transfer from the LSU 210 when the number of available entries in the L2 cache store queue 260 is greater than or equal to 4 entries. In another exemplary embodiment, the time delay is calculated to be longer or shorter based on specific performance metrics within the processor core, including but not limited to the number of available entries in the L2 cache store queue 260, the recent average value in the transfer time or transfer speed for transferring data information from the LSU store queue 250 to the L2 cache store queue 260, and / or the recent frequency of memory operations or store operations occurring within the processor core 200. The ability to strategically set or reset, or variably adjust, the time interval of the delay maximizes the opportunity for the L2 cache 230 to receive an accelerated data information transfer from the LSU 210, enhancing the performance of data information transfer between the LSU store queue 250 and the L2 cache store queue 260. When the L2 store queue 260 is near capacity and has few available entries, or when other factors indicate that the environment of the processor core 200 is under stress, the L2 cache control module 280 delays sending a POP signal to the LSU control module 270.Alternatively, if the L2 store queue 260 has multiple available entries, or if other factors indicate that the environment of the processor core 200 is not under stress, the L2 cache control module 280 does not delay the transmission of the POP signal to the LSU control module 270. By strategically implementing a time delay for the L2 cache 230 to initiate the transfer of data information from the LSU store queue 250 to the L2 cache store queue 260, the L2 cache 230 can successfully handle additional requests from the LSU 210 for additional data information transfers to be received, avoiding the need to reject such requests because the L2 cache store queue 260 has no available entries.

[0038] The present invention may be embodied as a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium (or multiple computer-readable storage media) having computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0039] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction-executing device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves in which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other electromagnetic wave propagating freely, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0040] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.

[0041] The computer-readable program instructions for carrying out operations of the present invention may be written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk®, C++, or the like, and conventional procedural programming languages such as the “C” programming language or similar programming languages, in either source code or object code.

[0042] Computer-readable program instructions can be executed entirely on a user's computer, can be executed partially on a user's computer as a stand-alone software package, can be executed partially on a user's computer and partially on a remote computer, or can be executed entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to implement aspects of the present invention.

[0043] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0044] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium containing the instructions comprises a manufacture including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0045] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer implemented process, whereby the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0046] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions described in the blocks may occur in a different order than that depicted in the figures. For example, two blocks shown in succession may actually be executed substantially simultaneously, or, depending on the functions involved, the blocks may alternatively be executed in reverse order. Other implementations may not require all of the disclosed steps to achieve the desired functionality. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified function or action, or by a combination of dedicated hardware and computer instructions.

Claims

1. A processor included in a central processing unit (CPU), wherein the CPU is included in a computer processor system, the CPU is coupled to a memory, and the processor: A load and store unit (LSU) having a control module and a store queue, the store queue including a plurality of entries for storing a plurality of information packets; A cache memory (L2 cache) coupled to the LSU, the L2 cache having a control module and a store queue, the store queue including a plurality of entries for storing a plurality of information packets Comprising The L2 cache determines that there is an available entry in the L2 cache store queue and sends a request to transfer an information packet to the LSU; The LSU sends an acknowledgement response from the L2 cache and transfers the information packet from the entry in the LSU store queue via a data bus coupled to the L2 cache; The L2 cache receives the information packet from the LSU store queue and stores the information packet in the available entry in the L2 cache store queue; The LSU predicts that the L2 cache store queue has additional available entries and sends an additional acknowledgement response to the L2 cache, and transfers an additional information packet from an additional entry in the LSU store queue before the L2 cache requests the additional information packet. Processor.

2. The L2 cache determines that there is an additional available entry in the L2 cache store queue, receives the additional information packet from the LSU store queue, and stores the additional information packet in the additional available entry in the L2 cache store queue; The L2 cache delays sending a subsequent request to transfer a subsequent information packet from the LSU to the LSU, and the subsequent request also functions as an acknowledgement response to the LSU regarding storing the additional information packet in the L2 cache. The processor according to claim 1.

3. The L2 cache requests the information packet by transmitting a POP signal from the L2 cache control module along a control bus coupled to the LSU control module, and the LSU confirms and responds to the request from the L2 cache by transmitting a PUSH signal from the LSU control module along the control bus coupled to the L2 cache control module, the processor according to claim 1.

4. The L2 cache control module determines to delay (subsequent request delay) the request for a subsequent information packet, and calculates the subsequent request delay based on the number of available entries in the L2 cache store queue, the processor according to claim 2.

5. The L2 cache control module calculates the subsequent request delay based on an average time for transferring the information packet from the LSU store queue to the L2 cache store queue, the processor according to claim 4.

6. The L2 cache control module sets and resets the subsequent request delay based on a threshold value, the threshold value being the number of available entries in the L2 cache store queue; The L2 cache control module sets the subsequent request delay at a fixed time interval when the number of available entries in the L2 cache store queue is less than the threshold value; The L2 cache control module resets the subsequent request delay without a time delay when the number of available entries in the L2 cache store queue is greater than or equal to the threshold value The processor according to claim 4.

7. The L2 cache determines that there are no additional available entries in the L2 cache store queue; The L2 cache control module transmits a BOUNCE signal indicating that the additional information packet has been rejected and not stored in the L2 cache store queue to the LSU control module along the control bus; The LSU must wait for a subsequent request from the L2 cache before retrying the transfer of the additional information packet The processor according to claim 1.

8. A method for improving the performance of store operation data information transfer within a computer processing system, the computer processing system comprising a processor and a memory, the processor having a cache memory and a load and store unit, the computer processing system further comprising a computer-readable storage medium in which computer-usable program code is embodied, the computer-usable program code being configured to perform operations when executed by the processor, the method comprising: Storing an information packet in the load and store unit (LSU), the information packet being destined for transfer to the cache memory (L2 cache), the LSU having a store queue including a plurality of entries for storing a plurality of information packets, the L2 cache having a store queue including a plurality of entries for storing a plurality of information packets; When the L2 cache determines that there is an available entry in the L2 cache store queue, sending a request from the L2 cache to the LSU to transfer the information packet to the L2 cache; Sending an acknowledgement response from the LSU to the L2 cache and transferring the information packet from an entry in the LSU store queue to the L2 cache; Receiving the information packet in the L2 cache and storing the information packet in the available entry in the L2 cache store queue; and Predicting that the L2 cache has additional available entries in the L2 cache store queue, sending an additional acknowledgement response from the LSU to the L2 cache, and transferring an additional information packet from an additional entry in the LSU store queue to the L2 cache before the L2 cache requests the additional information packet A method comprising the steps of: Claim 9 Determining that the L2 cache has additional available entries in the L2 cache store queue; Storing the additional information packet from the LSU store queue in the additional available entry in the L2 cache store queue; and delaying the transmission of a subsequent request to transfer a subsequent information packet from the LSU, the subsequent request also functioning as a confirmation response to the LSU regarding the storage of the additional information packet in the L2 cache; The method according to claim 8, further comprising.

10. The L2 cache has a control module, and the step in which the L2 cache transmits the request further includes the step of the L2 cache control module transmitting a POP signal along a control bus coupled to the LSU; The LSU has a control module, and the step in which the LSU transmits the confirmation response further includes the step of the LSU control module transmitting a PUSH signal to the control bus coupled to the L2 cache The method according to claim 8.

11. The step of delaying the request for the subsequent information packet (subsequent request delay) is determined by the L2 cache control module, and the L2 cache control module calculates the subsequent request delay based on the number of available entries in the L2 cache store queue. The method according to claim 10.

12. The method according to claim 11, wherein the L2 cache control module calculates the subsequent request delay based on an average time for transferring the information packet from the LSU store queue to the L2 cache store queue.

13. The L2 cache control module sets and resets the subsequent request delay based on a threshold value, where the threshold value is the number of available entries in the L2 cache store queue; The L2 cache control module sets the subsequent request delay at a fixed time interval when the number of available entries in the L2 cache store queue is less than the threshold value; The L2 cache control module resets the subsequent request delay without a time delay when the number of available entries in the L2 cache store queue is greater than or equal to the threshold value The method according to claim 11.

14. When the L2 cache determines that there is no additional available entry in the L2 cache store queue, rejecting the additional information packet, where the L2 cache control module transmits a BOUNCE signal indicating that the additional information packet has been rejected and not stored in the L2 cache store queue, along the control bus to the LSU control module; and When the L2 cache transmits a subsequent request to the LSU, the LSU retries the transfer of the additional information packet The method according to claim 10, further comprising.

15. A computer program product for improving the performance of store operation data information transfer in a computer processing system, the computer processing system comprising a processor and a memory, the processor having a cache memory and a load and store unit, the computer program product comprising a non-transitory computer-readable storage medium in which computer-usable program code is embodied, the computer-usable program code being configured to perform operations when executed by the at least one processor, the operations comprising: Storing an information packet in the load and store unit (LSU), the information packet being destined for transfer to the cache memory (L2 cache), the LSU having a store queue including a plurality of entries for storing a plurality of information packets, the L2 cache having a store queue including a plurality of entries for storing a plurality of information packets; When the L2 cache determines that there is an available entry in the L2 cache store queue, transmitting a request from the L2 cache to the LSU to transfer the information packet to the L2 cache; Transmitting an acknowledgement response from the LSU to the L2 cache and transferring the information packet from an entry in the LSU store queue to the L2 cache; Receiving the information packet in the L2 cache and storing the information packet in the available entry in the L2 cache store queue; and A procedure of predicting that the L2 cache has additional available entries in the L2 cache store queue, sending an additional confirmation response from the LSU to the L2 cache, and transferring an additional information packet from an additional entry in the LSU store queue to the L2 cache before the L2 cache requests the additional information packet A computer program product having the above. **Claim 16** A procedure for the L2 cache to determine that there are additional available entries in the L2 cache store queue; A procedure for storing the additional information packet from the LSU store queue in the additional available entry in the L2 cache store queue; and A procedure for delaying the transmission of a subsequent request from the L2 cache indicating that a subsequent information packet is to be transferred from the LSU, where the subsequent request also functions as a confirmation response to the LSU regarding the storage of the additional information packet in the L2 cache The computer program product according to claim 15, further having the above. **Claim 17** The L2 cache has a control module, and the procedure for the L2 cache to send the request further includes a procedure for the L2 cache control module to send a POP signal along a control bus coupled to the LSU; The LSU has a control module, and the procedure for the LSU to send the confirmation response further includes a procedure for the LSU control module to send a PUSH signal to the control bus coupled to the L2 cache The computer program product according to claim 15. **Claim 18** The procedure for delaying the request for a subsequent information packet (subsequent request delay) is determined by the L2 cache control module, and the L2 cache control module calculates the subsequent request delay based on the number of available entries in the L2 cache store queue. The computer program product according to claim 17. **Claim 19** The L2 cache control module sets and resets the subsequent request delay based on a threshold value, where the threshold value is the number of available entries in the L2 cache store queue; The L2 cache control module sets the subsequent request delay at a fixed time interval when the number of available entries in the L2 cache store queue is less than the threshold value; When the number of available entries in the L2 cache store queue is greater than or equal to the threshold, the L2 cache control module resets the subsequent request delay without time delay. The computer program product according to claim 17. **Claim 20** When the L2 cache determines that there is no additional available entry in the L2 cache store queue, a procedure for rejecting the additional information packet, where the L2 cache control module transmits a BOUNCE signal along the control bus to the LSU control module to indicate that the additional information packet has been rejected and not stored in the L2 cache store queue; and When the L2 cache transmits a subsequent request to the LSU, a procedure for the LSU to retry the transfer of the additional information packet The computer program product according to claim 17, further comprising.