MEMORY COHERENCE TESTING REGARDING PAGE TRANSFER DISABLED

By delaying the SRQ purge cycle in multi-core processors, the method ensures proper synchronization and coherence in TLB entry deactivation, addressing cache coherency issues and improving data consistency across processing elements.

DE112022003732B4Active Publication Date: 2025-11-13INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
DE112022003732
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-10-04
Filing Date
2022-09-29
Publication Date
2025-11-13
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

In multi-core processors, there is a conflict situation between the deactivation cycle of translation lookaside buffer (TLB) entries and the clean cycle of store reorder queues (SRQs), leading to potential cache coherency issues due to improper handling of page translation invalidation and SRQ logic corruption.

Method used

Implementing a delay in the purge cycle of SRQs to ensure proper synchronization and coherence by allowing additional time for checking SRQ logic functionality before sending acknowledgement signals, thereby reducing the likelihood of cache coherency failures.

Benefits of technology

Enhances cache coherency by ensuring that SRQs are properly cleared before deactivating TLB entries, minimizing conflicts and maintaining consistent data visibility across processing elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for disabling page replication entries in a data processing system with a plurality of processing elements, wherein the method comprises: Applying a delay (330) to a store reorder queue (SRQ) (229) cleanup cycle of a processing element; Clean up SRQ (229) during the delayed cleanup cycle; and Receiving a translation lookaside buffer invalidation (TLBI) instruction (302) from a connection linking the plurality of processing elements, wherein the TLBI instruction (302) is an instruction to invalidate an address lookaside buffer (TLB) entry (318, 328) corresponding to a virtual memory page and / or a physical memory section, wherein the TLBI instruction (302) is broadcast by another processing element, where by applying the delay (330) to the cleanup cycle of the SRQ (229) the amount of overlap between the cleanup cycle of the SRQ (229) and a deactivation cycle associated with the TLBI (302) is increased.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] The present invention relates to embodiments in a processor and in particular data processing and especially cache coherence and deactivation of page conversions in a multi-core processor, a microprocessor or in a multi-processor system.

[0002] For example, a data processing system can have virtual memory for accessing addresses in physical memory without knowing the exact location of the address in physical memory. A mapping of virtual memory addresses to physical memory addresses can be managed and stored as a page table. For example, when a program accesses a virtual memory address, an address translation can be performed using the page table to determine which physical memory address the accessed virtual memory address refers to. The data stored at the determined physical memory address can then be read from that physical memory address.

[0003] In a multiprocessor system with multiple processing elements (e.g., a system with multiple processors, a single processor with multiple cores), all cores can share the page table. To speed up access to the page table translations, each processing element (processor or core) can store its own translation lookaside buffer (TLB), where each TLB can be a cache representing a portion of the page table. A TLB can contain a number of page table entries, and each TLB entry can contain a mapping from a virtual address to a physical address. For example, the TLB entries can be managed so that the portion of the total available memory occupied by the TLB can contain the most recently accessed, the most frequently accessed, or the most likely to be accessed portion of the total available memory.As data is moved into and out of physical memory (e.g., when a new process is started or a context changes), the entries in the TLBs must be updated to reflect the new data, and the TLB entries associated with data swapped out of system memory must be disabled. Since each core manages its own TLB, the cores must exchange data with each other to maintain cache coherence.

[0004] Document US 2005 / 0114607A1 describes a device and method for synchronous page table updates with low overhead.

[0005] Document US 2017 / 0286300A1 describes a processor, a method, and a data processing system for translating effective addresses to real addresses. When a program is replaced by the operating system running on a microprocessor, only the entries associated with the replaced program and contained in the effective address translation units are replaced. The entries in the effective address translation units associated with the operating system and shared libraries, as well as all other software units operating on the microprocessor, are not invalidated.

[0006] Document US 6 338 128 B1 describes a detection method for hazards encountered when load and store instructions are executed out-of-order without using real addresses in a processor unit.

[0007] Document US 2019 / 0108034A1 describes a detection of hazards when executing load and store instructions out of order (out-of-order execution) without using real addresses in a processor unit. SUMMARY

[0008] The invention is described by the features of the independent claims. Embodiments are specified in the dependent claims.

[0009] The summary of the disclosure serves to facilitate understanding of the data processing systems and the methods for disabling page-turn entries and managing cache coherence, without limiting the disclosure or the invention. This disclosure is addressed to a person skilled in the art. It should be clear that various aspects and features of the disclosure may be advantageously used in certain cases alone or in other cases in combination with other aspects and features. Accordingly, variations and modifications can be made to the storage systems, the architectural structure, and the procedures to achieve different effects.

[0010] As described in some examples, a general procedure for disabling page reorder entries in a data processing system is outlined. The data processing system may contain multiple processing elements. The procedure may include delaying a store reorder queue (SRQ) cleanup cycle for a processing element. It may also include cleaning the SRQ during the delayed cleanup cycle. Furthermore, the procedure may include receiving a translation lookaside buffer invalidation (TLBI) instruction from a connection linking the multiple processing elements. The TLBI instruction may be an instruction to disable an entry in the address lookaside buffer (TLB), corresponding to a virtual memory page and / or a physical memory area.The TLBI application can be rounded up by another processing element. By adding a delay to the SRQ cleanup cycle, the difference between the SRQ cleanup cycle and a TLBI-related deactivation cycle can be reduced.

[0011] As described in some examples, a data processing system is generally configured to disable page-shift entries in a data processing system. The data processing system can contain a first processing element, a second processing element, and a connection associated with the first and second processing elements. The first processing element can be configured to broadcast an instruction to disable an address-shift buffer (TLBI) on the connection. The TLBI instruction can be an instruction to disable an entry of the address-shift buffer (TLB) corresponding to a virtual memory page and / or a physical memory area. The second processing element can be configured to delay a cleanup cycle of a memory sort queue (SRQ) of the second processing element.Furthermore, the second processing element can be configured to clear the SRQ during the delayed cleanup cycle. Additionally, the second processing element can be configured to receive the TLBI instruction from the connection. By adding the delay to the SRQ cleanup cycle, the difference between the SRQ cleanup cycle and a TLBI deactivation cycle can be reduced.

[0012] As described in some examples, a processing element is generally configured to disable page reorder entries in a data processing system. The processing element may contain a processor pipeline with one or more load store units (LSUs) configured to execute load and store instructions. The one or more LSUs may be configured to delay a cleanup cycle of the processing element's store reorder queue (SRQ). The one or more LSUs may further be configured to clean the SRQ during the delayed cleanup cycle. The one or more LSUs may be configured to receive a TLBI deactivation instruction for the address translation buffer (TLB) from a connection linking multiple processing elements.The TLBI instruction can be an instruction to disable a TLB entry corresponding to a virtual memory page and / or a physical memory area. The TLBI instruction is then forwarded by another processing element. By adding a delay to the SRQ cleanup cycle, the difference between the SRQ cleanup cycle and a cleanup cycle associated with the TLBI is reduced.

[0013] Further features, as well as the structure and operation of various embodiments, are described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The various aspects, features, and embodiments of a processor, a processor system, and / or a method for processing data become clearer when read in conjunction with the provided figures. The figures provide embodiments to illustrate aspects, features, and / or different embodiments of the processor, the processor system, and methods for handling and processing data, but the claims should not be limited precisely to the system, embodiments, methods, processes, and / or units in the form shown, and the features and / or processes shown may be used individually or in combination with other features and / or processes.It should be mentioned that a numbered element is numbered according to the figure in which it first appears, and is often, though not always, also designated by this number in subsequent figures, with identical reference numbers often, though not always, representing identical parts of the illustrative embodiments of the invention. Fig. Figure 1 shows a general computing or data processing system according to one embodiment. Fig. Figure 2 shows a block diagram according to one embodiment. Fig. Figure 3 illustrates an exemplary implementation of the memory coherence test with regard to the deactivation of a page mapping according to one embodiment. Fig. Figure 4 illustrates another exemplary implementation of the memory coherence test with regard to the deactivation of a page mapping according to one embodiment. Fig. Figure 5 illustrates another exemplary implementation of the memory coherence test with regard to the deactivation of a page mapping according to one embodiment. Fig. Figure 6 illustrates an exemplary flowchart for checking memory coherence with respect to disabling a page mapping according to one embodiment. DETAILED DESCRIPTION

[0015] The following description serves to illustrate the general basic ideas of the invention and is not to be understood as a limitation of the concepts claimed herein. Numerous details are set forth in the following detailed description to facilitate an understanding of a processor, its architectural structure, and its operation. However, it should be clear to those skilled in the art that different and numerous embodiments of the processor, the architectural structure, and the operating methods can be implemented without these specific details, and that the claims and the invention should not be limited to the embodiments, subassemblies, features, processes, methods, aspects, characteristics, or details described and shown in detail herein.Furthermore, individual features described herein can be used in combination with other described features in any of the various possible combinations and permutations.

[0016] Unless otherwise specifically defined, all terms shall be interpreted as broadly as possible, including meanings derived from the description as well as meanings that are familiar to the person skilled in the art and / or defined in dictionaries, textbooks, etc.

[0017] The term "workload" of a processor refers to the number of instructions executed by the processor during a specific period of time or at a specific time.

[0018] A processor can execute instructions in a series of small steps. To increase the number of instructions processed by the processor (and thus its speed), the processor can, in some cases, operate in a pipeline system. In a pipeline system, individual stages are provided within the processor, with each stage performing one or more of the steps required to execute an instruction. For example, the pipeline (in addition to other circuitry) can be located in a part of the processor called a processor core. Some processors can have multiple processor cores (e.g., in a multiprocessor system), and in certain cases, each processor core can have multiple pipelines. When a processor core has multiple pipelines, groups of instructions (called output groups) can be issued to the multiple pipelines in parallel and executed by each pipeline simultaneously.The pipeline can contain multiple stages, such as a decoding stage, an allocation stage, an execution stage, and so on. The execution stage can contain execution units that perform different types of operations as specified by the instructions. For example, a load-store unit (LSU) is an execution unit that processes, for instance, load instructions and store instructions.

[0019] For example, the physical address of a memory instruction being executed can be stored as entries in a memory sort queue (SRQ) in a memory storage unit (LSU). The SRQ can reside in an L1 data cache of a processor core. The entries in the SRQ can be committed memory instructions. Committed memory instructions are those executed by a processor or processing element, and their execution cannot be undone. Other processing elements are unaware of the execution until the SRQ is flushed to memory (e.g., a two-level (L2) cache) and thus cleared. After an SRQ is cleared or a committed memory instruction is flushed to memory (e.g., a two-level (L2) cache), the SRQ is no longer accessible.In a two-level (L2) cache, a value stored or updated at a memory address specified by the written memory instruction can be visible to all processors within the multiprocessor system. For example, if the L2 cache is a global memory accessible to all processing elements, then after the SRQ entries are swapped to the L2 cache, the values ​​updated by the memory instructions of the retrieved SRQ entries can be visible to all processing elements. According to one or more exemplary embodiments, the L2 cache can sometimes be a memory located near a processing element, and higher-level caches can also be used; for example, a level three (L3) cache can be the global memory accessible to multiple processing elements.SRQ entries can be moved from an SRQ in local storage to global storage.

[0020] For example, a mapping between a virtual address and a physical address can become invalid in response to certain events. For instance, when data is swapped to or from physical memory (e.g., when a new process is started or a context switch occurs), the entries in the TLBs must be updated to reflect the presence of the new data, and the swapped-out data associated with the corresponding TLB entries must be disabled. For example, an instruction to disable TLB entries can be called a TLB Disable instruction (TLBI). When a mapping between a virtual address and a physical address becomes invalid, a TLBI instruction is issued to all cores to swap out TLB entries corresponding to the disabled mapping.For example, one core can be tasked with broadcasting the TLBI instruction to the other cores within the multiprocessor system.

[0021] For example, a first processor can disable a specific TLB entry in its TLB, where this specific TLB entry can map a specific virtual address to a specific physical address. The first processor can halt any process and / or any instructions related to that physical address (e.g., writing committed memory instructions to that physical address) and broadcast the TLBI instruction over a connection reachable by all processors within the multiprocessor system. A second processor within the multiprocessor system can receive the TLBI instruction from the connection and, in response, swap all SRQ entries related to that physical address into memory. After clearing, the second processor can disable entries in its own TLB that contain that physical address.In response to the cleanup, the second processor can send an acknowledgment back to the first processor to indicate that the second processor has completed its TLB deactivation. The first processor can wait for acknowledgment from the second processor and from all other processors within the multiprocessor system before deactivating the relevant TLB entry in its own TLB.

[0022] For example, a transmission of the TLBI instruction might include the time it takes to travel from the first processing unit to the connection, traverse the connection, and then travel from the connection to the second processing unit. During this transmission time, the second processor can clear its SRQ at its normal SRQ clearing rate or within its clearing cycle. A problem can arise if the SRQ entry associated with the specific physical address of the TLBI instruction has not been swapped out by the second processing unit before the second processing unit sends back the acknowledgment to the first processing unit that forwarded the TLBI instruction. For example, SRQ logic or an SRQ algorithm can be used to detect whether the second processing unit has completed clearing the SRQ.If the SRQ logic is corrupted and incorrectly determines whether an SRQ is fully cleared, there is a possibility that the second processing element could send the acknowledgment before it has fully cleared the SRQ. If the second processing element is unable to fully clear its SRQ before sending the acknowledgment, the value displayed by the TLBI instruction at the specified physical address may be outdated after all processing elements have performed their deactivations according to the broadcast TLBI instruction. After all processing elements have completed their respective deactivations, they may be reading an outdated value from the physical address in question.This results in a conflict between the cleanup of the TLB entry (or page mapping), the commit to the SRQ, and the deactivation, whereby this conflict can impair the visibility of the memory instruction results for all processing elements in the multiprocessor system.

[0023] The methods and systems described herein can widen the window of conflict between a TLB entry deactivation cycle (e.g., TLBI) and a processor cleanup cycle in a multiprocessor system to verify the proper functioning of the processors' SRQ logic. Widening the window for this conflict can increase the chances of an SRQ cleanup event (or cycle) overlapping with a TLBI deactivation cycle. For example, the SRQ cleanup cycle in a processor can be delayed to slow down SRQ cleanup. Cleanup cycles can typically be shorter than deactivation cycles (e.g., SRQ cleanup may be faster than swapping out the TLBI instruction). Furthermore, the processing time of memory instructions can be variable, as memory instructions can be variable (e.g.,(because they depend on other threads and results from other processor cores), making the cleanup cycle unpredictable. By adding a delay to the SRQ cleanup cycle, a processor can additionally be able to detect potential problems and corruption in the SRQ logic. For example, without the delay, a processor's SRQ can be cleaned relatively quickly, and there's a greater chance that the SRQ is already empty when a processor receives a TLBI instruction from the connection. Since the SRQ logic indicates that the SRQ is empty, the processor can send an acknowledgment without verifying whether the SRQ is actually empty, and the SRQ logic cannot be checked.The additional time provided by the delayed cleanup cycle increases the likelihood that an empty SRQ will be present when the processor receives a TLBI instruction. This allows for a check to determine if the SRQ logic is functioning correctly to confirm cache coherence. If the SRQ is not empty, the processor's SRQ logic can instruct it to check and clean the SRQ before sending an acknowledgment. If the SRQ logic successfully instructs the processor to clean the SRQ before sending the acknowledgment, the processor's cache coherence check can be considered successful. If the SRQ logic fails to instruct the processor to clean the SRQ before sending the acknowledgment, the processor's cache coherence check can be considered a failure.

[0024] Fig. Figure 1 illustrates an information processing system 101, which can be a simplified example of a computer system capable of performing the data processing operations described herein. The computer system 101 can contain one or more processors 100 connected to a host bus 102. The processor(s) 100 can be, for example, a standard microprocessor, a custom processor, a custom programmable gate array (FPGA), an application-specific integrated circuit (ASIC), discrete logic, etc., or generally any instruction-executing unit. By way of example, the processor(s) 101 can be the multi-core processor(s) containing two or more processor cores. An L2 (layer two) cache memory 104 can be connected to a host bus 102. An I / O bridge (e.g.,A host-to-PCI bridge (106) can be connected to main memory (108), with each I / O bridge containing cache and main memory functions and providing bus control for data transfers between a PCI bus (110), the processor (100), the L2 cache (104), main memory (108), and the host bus (102). The main memory (108) can be connected to the I / O bridge (106) as well as to the host bus (102). Likewise, other memory types, such as random access memory (RAM) and / or various volatile and / or non-volatile memory units, can also be connected to the host bus (102) and / or the I / O bridge (106). For example, storage units connected to the host bus 102 can include electrically erasable programmable read-only memory (EEPROM), programmable read-only flash memory (PROM), battery backup RAM, hard disk drives, etc.Non-volatile memory units connected to the host bus 102 can contain executable firmware and programming instructions containing any non-volatile data, which can be executed to cause the processor 100 to perform specific functions, such as the procedures described herein. Units such as one or more I / O components used exclusively by one or more processors 100 can be connected to the PCI bus 110. A service processor interface and an ISA access pass-through 112 can provide an interface between the PCI bus 110 and the PCI bus 114. In this way, the PCI bus 114 can remain separate from the PCI bus 110. Units such as flash memory 118 are connected to the PCI bus 114.In one implementation, the flash memory 118 can contain BIOS code, which includes code executable by a processor and essential for a variety of basic system functions and system boot functions.

[0025] The PCI bus 114 can provide an interface for a variety of units shared by the one or more host processors 100 and the service processor 116, including, for example, flash memory 118. The PCI-to-ISA bridge 135 provides bus control for handling transfers between the PCI bus 114 and the ISA bus 140, the UBS (Universal Serial Bus) functionality 145, and the power-saving functionality 155, and may include other functional elements not shown, such as a real-time clock (RTC), DMA control, interrupt support, and system management bus support. Non-volatile RAM 120 can be connected to the ISA bus 140. The service processor 116 can include a bus 122 (e.g. a JTAG and / or an I2C bus) for transferring data to one or more processors 100 during the initialization steps.Bus 122 can also be connected to L2 cache 104, I / O bridge 106, and main memory 108 to provide a data transfer path between the processor, service processor, L2 cache, host-to-PCI bridge, and main memory 108. Service processor 116 also has access to power resources to shut down information processing unit 101.

[0026] Peripheral units and input / output (I / O) units can be connected to various interfaces (e.g., a parallel interface 162, a serial interface 164, a keyboard interface 168, and a mouse interface 170) connected to the ISA bus 140. Alternatively, many I / O units can be controlled by a (not shown) super I / O controller connected to the ISA bus 140. Interfaces that allow one or more processors 100 to exchange data with external units include, but are not limited to, serial interfaces such as RS-232, USB (Universal Serial Bus), SCSI (Small Computer Systems Interface), RS-309, or wireless data transmission interfaces such as WLAN, Bluetooth, NFC (Near Field Communication), or other wireless interfaces.

[0027] To connect computer system 101 to another computer system for copying files over a network, as in one example, the I / O component 130 can contain a LAN card connected to the PCI bus 110. To connect computer system 101 to an ISP to establish an internet connection using a telephone line, the modem 175 is connected to the serial port 164 and the PCI-to-ISA bridge 135.

[0028] Fig. Figure 1 shows an information processing system that uses one or more processors 100, but the information processing system can take many forms. For example, the information processing system 101 can take the form of a desktop, a server, a portable computer, a laptop, a notebook, or a computer or data processing system with another form factor. The information processing system 101 can also take other form factors, such as a personal digital assistant (PDA), a game console, an ATM, a mobile phone, a data transmission unit, or other units that contain a processor and memory.

[0029] Fig. Figure 2 shows a block diagram of a processor 200 according to one embodiment. The processor 200 can include at least one memory 202, an instruction cache 204, an instruction retrieval unit 206, a branch prediction unit 208, and a processor pipeline or a processing pipeline 210. The processor 200 can be integrated into a computer processor or distributed within a computer system. Instructions and data can be stored in the memory 202, and the instruction cache 204 can access instructions in the memory 202 and store the instructions to be retrieved. The memory 202 can contain any type of volatile or non-volatile memory, for example, a cache memory. The memory 202 and the instruction cache 204 can contain multiple cache levels. The processor 200 can also include a data cache (not shown).According to one embodiment, the instruction cache 204 can be configured to provide instructions in an 8-way associative mapping structure. Alternatively, any other desired configuration and size can be used. For example, the instruction cache 204 can be implemented as a fully associative mapping, an n-way associative mapping, or a direct mapping configuration.

[0030] In Fig. Figure 2 shows a simplified example of the instruction retrieval unit 206 and the processing pipeline 210. According to various embodiments, the processor 200 can contain multiple processing pipelines 210 and instruction retrieval units 206. According to one embodiment, the processing pipeline 210 contains a decoding unit 20, an output unit 22, an execution unit 24, and write-back logic 26. According to some examples, the instruction retrieval unit 206 and / or the branch prediction unit 208 can also be part of the processing pipeline 210. The processing pipeline 210 can also contain other features, such as error checking and processing logic, sorting buffers, one or more parallel paths through the processing pipeline 210, and other features currently or subsequently known in the art. While in Fig. Figure 2 shows a forward path through processor 200, however, other feedback and signal paths may be interposed between elements of processor 200.

[0031] Jump instructions (or "branches") can be either unconditional, meaning the jump is executed every time the instruction is encountered in the program, or conditional, meaning the jump is executed or not depending on a condition. The Processor 200 can provide conditional jump instructions to move from one instruction to a target instruction (and thereby skip any intervening instructions) if a condition is met. If the condition is not met, the next instruction following the jump instruction can be executed without jumping to the target instruction. In most cases, the instructions following the conditional jump are not known for certain until the underlying condition for the branch has been resolved.The branch prediction unit 208 can attempt to predict the outcome of conditional jump instructions in a program before the jump instruction is executed. If a branch is predicted incorrectly, all speculation beyond the branch point in the program must be discarded. For example, when a conditional jump instruction appears, the processor 200 can predict which instruction will be executed once the outcome of the jump condition is known. Instead of halting the processing pipeline 210 when the conditional jump instruction is issued, the processor can continue issuing instructions, starting with the next predicted instruction.

[0032] In a conditional jump, control can be transferred to the target address depending on the results of a preceding instruction. Conditional jumps can be resolved or unresolved, depending on whether the result of the preceding instruction is known at the time of jump execution. If the jump condition is resolved, it is known whether the jump should be executed. If the conditional jump is not executed, the next sequence of instructions immediately following the jump instruction is executed. If the conditional jump is executed, the sequence of instructions beginning with the target address is executed.

[0033] The instruction retrieval unit 206 retrieves instructions from the instruction cache 204 according to an instruction address for further processing by the decoding unit 20. The decoding unit 20 decodes instructions and forwards the decoded instructions, parts of instructions, or other decoded data to the output unit 22. The decoding unit 20 can also detect jump instructions that were not predicted by the jump prediction unit 208. The output unit 22 analyzes the instructions or other data and, based on the analysis, sends the decoded instructions, parts of instructions, or other data to one or more execution units in the execution unit 24. The execution unit 24 executes the instructions and determines whether the predicted jump direction is incorrect. The jump direction can be "taken" by retrieving subsequent instructions from the target address of the jump instruction.Otherwise, the jump direction cannot be "taken", and subsequent instructions are retrieved from memory locations following the jump instruction. If an incorrectly predicted jump instruction is detected, the instructions following the incorrectly predicted branch can be deleted from the various units of Processor 200.

[0034] The execution unit 24 can contain multiple execution units, such as fixed-point execution units, floating-point execution units, load / store execution units (or load-store units, LSUs), and vector multimedia execution units. The execution unit 24 can also contain specialized branch prediction units for predicting the destination of a multiple branch. The write-back logic 26 writes the results of instruction execution back to a destination resource 220. The destination resource 220 can be any type of resource, including registers, cache memory, other memory, I / O circuits for exchanging data with other units, other processing circuits, or other destinations for executed instructions or data. One or more of the processor pipeline units can also provide information about the execution of conditional branch instructions to the branch prediction unit 208.

[0035] For example, an execution time slice can be defined as a set of data processing circuits or hardware units connected in series with a processor core. An execution time slice can be a pipeline or a pipeline-like structure. Multiple execution time slices can be used as part of multi-threaded processing within a single processor core or within a multi-processor system. In modern computer architecture, there can be multiple execution units within an execution time slice, including linear logic units (LSUs), vector scalar units (VSUs), arithmetic logic units (ALUs), and other execution units.An LSU typically contains one or more memory queues, each with entries for tracking memory instructions and retaining memory data, and one or more load queues, each with entries for tracking load instructions and retaining load data.

[0036] According to one embodiment, the processor 200 can perform branch prediction to speculatively retrieve instructions following conditional branch instructions. The branch prediction unit 208 is designed to perform such branch predictions. According to one embodiment, the instruction cache 204 can send a signal to the branch prediction unit 208 indicating that the instruction address is being retrieved, enabling the branch prediction unit 208 to determine which target branch addresses to select for making a branch prediction. The branch prediction unit 208 can be connected to various parts of the processing pipeline 210, such as the execution unit 24, the decoding unit 20, the recording buffer, etc., to determine whether the predicted branch direction is correct or incorrect.

[0037] To enable multithreaded processes, instructions from different threads can be nested in a specific way at a point in the overall processor pipeline. In one example technique for nesting instructions from different threads, instructions are nested cycle by cycle based on nesting rules. For example, instructions from different threads can be nested so that a processor can execute an instruction from a first thread in the first clock cycle, then an instruction from a second thread in the second clock cycle, and then another instruction from the first thread in the third clock cycle, and so on. Some nesting techniques allow each thread to be assigned a priority, and then instructions from the different threads can be nested based on these assigned priorities.For example, if a first thread is assigned a higher priority than a second thread, a nesting rule can require that twice as many instructions from the first thread with the higher priority be inserted into the nested data stream compared to the second thread with the lower priority. Different nesting rules can be defined, such as rules for resolving threads with the same priority or rules for periodically nesting instructions from less important threads (e.g., executing an instruction from a low-priority thread every X cycles).

[0038] By nesting threads based on priorities, processor resources can be allocated according to the assigned priority. However, thread priorities sometimes fail to account for processor events such as incorrect branch predictions, which can affect a thread's ability to progress through a processor pipeline. These events can sometimes negatively impact the performance of the processor resources allocated to individual instruction threads in a multithreaded processor. For example, priority-based techniques that allocate threads with fewer instructions in the decode, rename, and instruction queue stages of the pipeline may sometimes be ineffective in reducing incorrect path assignments in the pipeline caused by incorrect predictions (e.g., incorrectly speculated instructions).These incorrect path instructions can tie up the retrieval bandwidth and other valuable resources of the processor, such as instruction queues and other functional units.

[0039] The productivity and / or performance of processor 200 can be increased by reducing the number of incorrect path instructions in the processing pipeline. For example, threads with more frequent incorrect predictions in processing pipeline 210 can be deferred (e.g., retrieved more slowly by the instruction retrieval unit), resulting in a reduction of incorrect path instructions in processing pipeline 210. Furthermore, a number of instructions following an initial processing pipeline 210 with unfinished or unresolved jump instructions can be traced to prevent the execution of an excessively large number of potentially incorrect path instructions.

[0040] According to one embodiment, the processor 200 can be an SMT processor configured to run multithreaded applications. The processor 200 can use one or more instruction queues 212 to collect instructions from one or more different threads. The instruction retrieval unit 206 can retrieve instructions stored in the instruction cache 204 and populate the instruction queues 212 with the retrieved instructions. The performance of the processor 200 can depend on how the instruction retrieval unit 206 populates these instruction queues 212. The instruction retrieval unit 206 can be configured to assign and manage priorities to the different threads and, based on these priorities, decide which instructions and / or threads should be retrieved and sent to the instruction queues 212.The processor 200 can also include a thread scheduler 214, which is configured to schedule instructions in the instruction queues 212 and distribute them to the processing pipeline 210. For example, the processor 200 can be a multi-core processor containing two or more processor cores, each of which can be configured to process a specific thread.

[0041] According to one example, in response to the fact that the execution unit 24 is a load-store unit (LSU) 228, a circuit 230 can be embedded or integrated into the LSU 228 to insert an SRQ pick-up delay. The SRQ pick-up delay can, for example, be a delay or slowdown of a pick-up cycle from a memory sort queue (SRQ) 229 (or a memory queue) for the memory (e.g., the L2 cache, which may be part of the target resource 220). According to another example, the circuit 230 can be enabled (e.g., turned on) or disabled (e.g., turned off) by the processor 200. Enabling and disabling the circuit 230 can be based on the operating state of the processor 200 and / or other processors or processor cores within a multiprocessor system.Processor 200 can, for example, activate circuit 230 to delay the extraction cycle of SRQ 229 in LSU 228. Conversely, processor 200 can deactivate circuit 230 to prevent the extraction cycle of SRQ 229 in LSU 220 from being delayed. By delaying the extraction cycle of SRQ 229, processor 200 can take additional time to check whether the cache coherence (e.g., load and store coherence) of processor 102 is successful or a failure. This additional time increases, for example, the chance that committed memory instructions remain in SRQ 229, allowing processor 102 to check the SRQ logic used to detect empty SRQs and trigger the sending of acknowledgment signals.If the SRQ is always empty when a TLBI instruction arrives, the SRQ logic cannot be checked. If the TLBI instruction is received and the SRQ logic activates a processing element to clear the SRQ before sending an acknowledgment signal, the cache coherence of processor 102 can be considered successful. If the TLBI instruction is received and the SRQ logic does not activate SRQ clearing, and a processing element sends an acknowledgment signal without clearing the SRQ, the cache coherence of processor 102 can be considered a failure.

[0042] Fig. Figure 3 illustrates an exemplary implementation of the memory coherence check with regard to disabling page mapping. According to one example, the processor 200 (see Fig. 1 and Fig. 2) N processing elements such as processing elements 310, 320, 340, designated as core 0, core 1 and core N, are included. In Fig. Figure 3 shows three processor cores, but the processor can contain 200 additional processor cores. A connecting element 301 (e.g., a bus, a mesh network, a crosspoint, etc.) can connect core 0, core 1, and core N, as well as other cores within the processor 200. Core 0 can contain a load-memory unit (LSU) 312, a level-two (L2) cache 316, and a TLB 318. The LSU 312 can contain an SRQ 314. The TLB 318 can contain multiple entries that indicate mappings between virtual memory addresses assigned to core 0 and a physical memory address. For example, the TLB 318 can contain entries labeled M1, M2, M3, and M4. Core 0 can contain a Load Storage Unit (LSU) 312, a Level Two (L2) Cache 316, and a TLB 318. The LSU 312 can contain an SRQ 314. Core 1 can contain an LSU 322, an L2 Cache 326, and a TLB 328.The LSU 322 can contain an SRQ 324. The TLB 328 can contain multiple entries indicating mappings between virtual memory addresses allocated to Core 1 and a physical memory address. Core N can contain an LSU 342, an L2 cache 346, and a TLB 348. The LSU 342 can contain an SRQ 344. The TLB 348 can contain multiple entries indicating mappings between virtual memory addresses allocated to Core N and a physical memory address. According to one or more exemplary embodiments, the L2 caches 316, 326, and 346 can be individual memory banks of a global L2 cache accessible to Core 0, Core 1, and Core N, respectively.

[0043] For example, upon an event such as a context switch, core 0 can disable an entry in TLB 318 and broadcast a TLBI (address mapping buffer disabling) instruction 302 on connection 301. TLBI instruction 302 can be an instruction to other processing elements to disable one or more TLB entries in their respective TLBs that correspond to a specific virtual address (e.g., P3) and / or a specific physical address (e.g., F4). For instance, TLBI instruction 302 can be an instruction to other processing elements, other than core 0, to disable TLB entries that map virtual addresses to the physical address F4.

[0044] The TLBI instruction 302 can be forwarded from core 0 to connection 301, then within connection 301, and then from connection 301 to processing elements such as core 1 and core N. Therefore, the transmission time of the TLBI instruction 302 can be the sum of the time intervals it takes for the TLBI instruction to travel from core 0, via connection 301, to a receiving core (e.g., core 1, core N). It should be noted that the transmission time of the TLBI instruction 302 can vary between different processing elements based on the distance between the receiving core and the core that issued the TLBI instruction 302, as well as different process variants, the hardware capabilities of the cores, the data traffic on the connection, and other factors. Fig. Although kernel 0 is shown as the processing element that issues a TLBI statement in Figure 3, other processing elements such as kernel 1 and kernel N can also be configured to issue TLBI statements regarding the deactivation of TLB entries in their respective TLBs.

[0045] Core 1 can receive TLBI instruction 302 from connection 301 and, in response, clear SRQ 324 or swap out entries remaining in SRQ 324 of LSU 322, where the SRQ entries to be swapped out may be committed memory instructions. In response to the complete clearing of SRQ 324, Core 1 can send an acknowledgment signal (ACK) 304 to Core 0 to inform Core 0 that SRQ 324 has been cleared. In response to the sending of ACK 304, Core 1 can disable all TLB entries in TLB 328 that relate to virtual address P3 and / or the virtual address F4 specified in TLBI instruction 302. For example, Core 1 can disable an entry M2 in TLB 328 that maps virtual address P3 to virtual address F4. Core 0 can respond to receiving the ACK signal 304 from core 1 and all other cores (e.g.,The ACK signal 306 from core N disables all TLB entries in TLB 318 that pertain to virtual address P3 and / or virtual address F4. For example, core 0 can wait for ACK signals from all cores before resuming normal operation. In response to receiving ACK signals, core 0 can, for instance, reassign page P3 to a different physical address and update TLB 318 with the new assignment.

[0046] For example, core 1 might execute logic 327 to detect if SRQ 324 is empty and, in response to the finding that SRQ 324 is empty, trigger an action to send ACK 304 to core 0. However, if logic 327 is corrupted, core 1 might incorrectly detect that SRQ 324 is empty, even though SRQ 324 might not be. If SRQ 324 is not empty, but core 1 sends ACK 302 to core 0, a problem can occur if a corrupted memory instruction has not been properly swapped out from SRQ 324. For example, if a committed memory instruction to store on F4 remains in SRQ 324, but the corruption in logic 327 causes an error in detecting the presence of the committed memory instruction in SRQ 324, core 1 can proceed to send the ACK signal 304 to core 0 and disable entry M2.As a result of this error, core 0 and other cores besides core 1 cannot see that a value in F4 is being updated, because the remaining committed memory instruction in SRQ 324 has not been cleared.

[0047] To reduce the probability that an SRQ is not properly cleared due to the error, core 0, core 1, and core N can each have a delay circuit (e.g., the one in Fig. The circuit shown (230) can be used, which can be configured to apply a delay of 330 to a cleanup cycle of SRQs 314, 324, and 344, respectively. For example, the delay circuit can be integrated into LSU 312, 322, and 342. The delay 330 can be a specific number of cycles added to a standard cleanup cycle of an SRQ (e.g., SRQs 314, 324, and 344), so that SRQs 314, 324, and 344 are cleaned more slowly in response to the delay 330. By slowing down the SRQ cleanup cycle, the chance that SRQ entries will still be present in the SRQ at the time a TLBI instruction is received can be increased. In other words, the delay of 330 allows the receiving processing element (e.g.,This allows more time for core 1 or core N (which receives TLBI instruction 302) to detect SRQ entries related to TLBI instruction 302 and take appropriate action to resolve the situation. For example, core 1 might detect an SRQ entry related to F4 in SRQ 324 and move the detected entry from SRQ 324 to the L2 cache 326. In response to the clearing of SRQ 324 (e.g., moving the entry until SRQ 324 is empty), core 1 might disable TLB entries related to TLBI instruction 302 in TLB 328.

[0048] Fig. Figure 4 illustrates another exemplary implementation of the memory coherence check with regard to disabling page mapping according to one embodiment. In a Fig. In the example shown, scenario 401 represents a core 1 that processes TLBI instruction 302 without applying a delay of 330, and scenario 402 represents a core 1 that processes TLBI instruction 302 with an application and a delay of 330. In scenario 401, when core 1 receives TLBI instruction 302 from the connection, SRQ 324 is empty, and SRQ entries E1, E2, and E3 have already been swapped out, for example, to the L2 cache 326. If SRQ entry E3 refers to TLBI instruction 302 (e.g., writing to physical address F4) and SRQ entry E3 was swapped out before core 1 receives TLBI instruction 302, then E3 has been swapped out correctly. However, if in scenario 401 logic 327 (see Fig. 3) If the SRQ is damaged, the statement that SRQ 324 is empty may be false. If the SRQ entry E3 is located in SRQ 324, but core 1 incorrectly confirms that SRQ 324 is empty, the entry E3 cannot be moved until core 1 sends the ACK signal 304 to link 301.

[0049] In scenario 402, when TLBI instruction 302 is received by core 1 from the link, SRQ 324 is not empty, and SRQ entry E3 is still present in SRQ 324 because a cleanup cycle of SRQ 324 has been applied with a delay of 330. Core 1 can determine that SRQ entry E3 relates to TLBI instruction 302 and can move SRQ entry E3 out of SRQ 324 before sending the ACK signal 304 to link 301. Applying the delay of 330 causes SRQ 324 to slow down its cleanup, giving core 1 additional time to detect SRQ entries within SRQ 324. For example, in scenario 401, core 1 can rely on the display from logic 327 that SRQ 325 is empty and send the ACK signal 304 to connection 301 without checking whether there are any entries in SRQ 324.By delaying the SRQ cleanup cycle, the probability of SRQ 324 being empty can be reduced, and thus core 1 can be caused to clean up SRQ 324 before the ACK signal 304 is sent.

[0050] For example, the delay could be a specific number of cycles added to a standard cleanup cycle of SRQ 324, and the number of cycles in delay 330 could be proportional to the time required to propagate TLBI instruction 302 through all cores in the multiprocessor system. Alternatively, the number of cycles in delay 330 could be the product of the number of cycles required to clean up each entry in SRQ 324 and the size of SRQ 324 (e.g., the number of allowed entries or the maximum number of entries in SRQ 324). For example, the standard number of cycles required to clean up each entry in SRQ 324 could be two (e.g., by removing one SRQ entry every two cycles), and the number of allowed entries in SRQ 324 could be 64.Thus, the number of cycles in the delay 330 can be any multiple of 64. According to another example, the circuit 230 (see . Fig. 2) Includes a random number generator for generating a random number between one and a multiple of the number of permissible entries in SRQ 324. The generated random number can be defined as the number of cycles in the delay 330 and / or as the number of delayed cycles for each SRQ entry. For example, the delay 330 applied to a first SRQ entry can be a first number of cycles, and the delay 330 applied to a second SRQ entry can be a second number of cycles. As an example, the circuit 230 can include linear feedback shift registers (LFSRs) to implement the generation of random numbers. The number of cycles defining the delay 330 can be arbitrary and based on a desired execution of the processor 102 (see Fig. 2) be configurable or programmable.

[0051] Fig. Figure 5 illustrates another exemplary implementation of the memory coherence test with regard to the deactivation of page mapping according to one embodiment. In the case described in Fig. In the example shown, a TLBI cycle 500 can capture a time span from T0 to T3. The TLB cycle 500 can contain the transmission time of a TLBI instruction from a first processing element to a second processing element. An SRQ cleanup cycle 502 can capture the time span from a time T0 to a time T1. The SRQ cleanup cycle can be shorter than the TLB cycle, allowing an SRQ to be cleared faster compared to transmitting a TLBI instruction. Thus, if the SRQ is cleared faster than the transmission of the TLBI instruction takes, the probability of the SRQ being empty when the TLBI instruction arrives at a processing element is higher. If the SRQ is empty when the TLBI arrives, SRQ logic intended to detect an empty SRQ cannot be checked because the SRQ is already empty (e.g., there is no occupied SRQ).

[0052] After applying the delay 330 to the SRQ cleanup cycle 502, the SRQ can be cleared more slowly within a delayed SRQ cleanup cycle between time T0 and time T1. This can reduce the difference between the TLBI cycle 500 and the SRQ cleanup cycle 502. In other words, the amount of overlap between the SRQ cleanup cycle 502 and the TLBI cycle 500 can be increased in response to the applied delay 330. The additional time between T1 and T2 resulting from the delay 330 increases the probability that more SRQ entries remain in the SRQ, so that the SRQ is not empty when the TLBI instruction arrives, and increases the probability that the SRQ logic (which is designed, for example, to detect an empty SRQ and send an acknowledgment) can be checked.It should be noted that the delay 330 can be variable, so that the SRQ cleanup cycle 502 can be delayed by different amounts depending on the desired execution mode.

[0053] Fig. Figure 6 illustrates an exemplary flowchart for checking memory coherence with respect to disabling during page conversion according to one embodiment. Process 600 may comprise one or more operations, actions, or functions, illustrated by one or more blocks 602, 604, and / or 606. Although discrete blocks are shown, different blocks can be further subdivided into smaller blocks, combined into fewer blocks, excluded, executed in parallel, or in a different order, depending on the desired execution mode.

[0054] Process 600 can begin in block 602. In block 602, a processing element can delay a cleanup cycle of a memory sort queue (SRQ) of the processing element. The processing element can be one of multiple processing elements in a data processing system. Process 600 can continue from block 602 to block 604. In block 604, the processing element can clean the SRQ during the delayed cleanup cycle. For example, the delay can be proportional to the time required to transmit the TLBI instruction from the other processing element over the connection to the processing element. Alternatively, the delay can be based on the product of the number of cycles required to move each entry in the SRQ and the size of the SRQ.According to another example, a number of cycles can be based on a random number between one and a multiple of the size of the SRQ.

[0055] Process 600 can proceed from block 604 to block 606. In block 606, the processing element can receive a TLBI (Total Address Translation Buffer) disable instruction from a link connecting multiple processing elements. The TLBI instruction can be an instruction to disable an address translation buffer (TLB) entry corresponding to a virtual memory page and / or a physical memory segment. The TLBI instruction can be broadcast by another processing element connected to the link. By adding a delay to the SRQ (Save Assignment Question) cleanup cycle, the difference between the SRQ cleanup cycle and a TLBI disablement cycle can be reduced.

[0056] For example, upon receiving the TLBI instruction, the processing element can determine whether to send an acknowledgment signal or clear the SRQ. The processing element can, for instance, detect an SRQ entry in the SRQ related to the TLB entry to be disabled and remove the detected SRQ entry from the SRQ. For example, upon the processing element completely clearing the SRQ, the processing element can send the acknowledgment signal to the processing element that broadcast the TLBI instruction over the connection. For example, upon receiving the TLBI instruction, the processing element can determine that the SRQ is empty and send the acknowledgment signal to the processing element that broadcast the TLBI instruction over the connection.

[0057] The present invention may be a system, a method, and / or a computer program product. The computer program product may comprise a computer-readable storage medium (or media) containing computer-readable program instructions to induce a processor to execute aspects of the present invention.

[0058] A computer-readable storage medium can be a physical unit capable of retaining and storing instructions for use by a system to execute instructions. For example, a computer-readable storage medium can be an electronic storage unit, a magnetic storage unit, an optical storage unit, an electromagnetic storage unit, a semiconductor storage unit, or any suitable combination thereof, without limitation. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a removable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), and erasable programmable read-only memory (EPROM).Flash memory), static random-access memory (SRAM), removable compact storage disk-read-only memory (CD-ROM), a DVD (digital versatile disc), a memory stick, a floppy disk, a mechanically coded unit such as punched cards or raised structures in a groove on which instructions are stored, and any suitable combination thereof. A computer-readable storage medium shall not, in its use herein, be understood as volatile signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses traveling through an optical fiber cable), or electrical signals transmitted by a wire.

[0059] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to individual data processing units or, via a network such as the internet, a local area network, a wide area network, and / or a wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission lines, wireless transmission, routing computers, firewalls, switching units, gateway computers, and / or edge servers. A network adapter card or network interface in each data processing unit receives computer-readable program instructions from the network and forwards them for storage on a computer-readable storage medium within the respective data processing unit.

[0060] Computer-readable program instructions for executing the steps of the present invention can be assembly instructions, ISA (Instruction Set Architecture) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., as well as conventional procedural programming languages ​​such as C or similar languages. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.In the latter case, the remotely located computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be established with an external computer (for example, via the internet using an internet service provider). In some embodiments, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can execute the computer-readable program instructions by using state information from the computer-readable program instructions to personalize the electronic circuits to perform aspects of the present invention.

[0061] Aspects of the present invention are described herein with reference to flowcharts and / or block diagrams or charts of methods, devices (systems), and computer program products according to embodiments of the invention. It is pointed out that each block of the flowcharts and / or block diagrams or charts, as well as combinations of blocks in the flowcharts and / or block diagrams or charts, can be executed by means of computer-readable program instructions.

[0062] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a specialized computer, or another programmable data processing device to create a machine, such that the instructions executed via the processor of the computer or other programmable data processing device generate a means of implementing the functions / steps specified in the block(s) of the flowcharts and / or block diagrams or charts.These computer-readable program instructions may also be stored on a computer-readable storage medium capable of controlling a computer, programmable data processing device and / or other units to function in a particular manner, such that the computer-readable storage medium on which instructions are stored has a manufactured product, including instructions that implement aspects of the function / step specified in the block(s) of the flowchart and / or block diagrams or charts.

[0063] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing device, or other unit to cause the execution of a series of process steps on the computer or other programmable device or other unit in order to generate a process executed on a computer, such that the instructions executed on the computer, other programmable device, or other unit implement the functions / steps specified in the block(s) of the flowcharts and / or block diagrams or charts.

[0064] The flowcharts and block diagrams or charts in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this context, each block in the flowcharts or block diagrams or charts can represent a module, segment, or part of instructions that includes one or more executable instructions for performing the specified logical function(s). In some alternative embodiments, the functions specified in the block may occur in a different order than shown in the figures. For example, two blocks shown consecutively may in reality be executed essentially simultaneously, or the blocks may sometimes be executed in reverse order depending on the corresponding functionality.It should also be noted that each block of the block diagrams or charts and / or flowcharts, as well as combinations of blocks in the block diagrams or charts and / or flowcharts, can be implemented by special hardware-based systems that perform the specified functions or steps, or execute combinations of special hardware and computer instructions.

Claims

[1] Method for disabling page replication entries in a data processing system with a plurality of processing elements, the method comprising: Applying a delay (330) to a store reorder queue (SRQ) (229) cleanup cycle of a processing element; Clean up SRQ (229) during the delayed cleanup cycle; and Receiving a translation lookaside buffer invalidation (TLBI) instruction (302) from a connection linking the plurality of processing elements, wherein the TLBI instruction (302) is an instruction to invalidate an address lookaside buffer (TLB) entry (318, 328) corresponding to a virtual memory page and / or a physical memory section, wherein the TLBI instruction (302) is broadcast by another processing element, where by applying the delay (330) to the cleanup cycle of the SRQ (229) the amount of overlap between the cleanup cycle of the SRQ (229) and a deactivation cycle associated with the TLBI (302) is increased. [2] The method of claim 1, further comprising: Detecting an SRQ entry in the SRQ (229) associated with the TLB entry (318,328) to be deactivated; and Clean up the detected SRQ entry from SRQ (229). [3] Method according to claim 1, further comprising sending an acknowledgment signal to the other processing element via the connection in response to a complete clearing of the SRQ (229). [4] Method according to claim 1, wherein the delay (330) is proportional to a time period required to transmit the TLBI instruction (302) from the other processing element via the connection to the processing element. [5] Method according to claim 1, wherein the delay (330) is based on a product of a number of cycles required to clear each entry from the SRQ (229) and a size of the SRQ (229). [6] Method according to claim 1, wherein a number of cycles in the delay (330) is based on a random number between one and a multiple of the size of the SRQ (229). [7] Data processing system comprising: a first processing element; a second processing element; a connection that is linked to the first processing element and the second processing element; wherein the first processing element is configured to broadcast an address translation buffer (TLBI) disable instruction (302) over the connection, wherein the TLBI instruction (302) is an instruction to disable an address translation buffer (TLB) entry (318,328) corresponding to a virtual memory page and / or a physical memory section. where the second processing element is configured such that it: a cleanup cycle of a memory sort queue (SRQ) (229) of the second processing element with a delay (330); the SRQ (229) is cleared during the delayed clearing cycle; and the TLBI instruction (302) is received from the connection, whereby by applying the delay (330) to the cleanup cycle of the SRQ (229) the amount of overlap between the cleanup cycle of the SRQ (229) and a deactivation cycle associated with the TLBI (302) is increased. [8] Data processing system according to claim 7, wherein the second processing element is configured such that it: detects an SRQ entry in SRQ (229) that is associated with the TLB entry (318,328) to be deactivated; and the detected SRQ entry was removed from SRQ (229). [9] Data processing system according to claim 7, wherein the second processing element is configured to send an acknowledgment signal to the first processing element via the connection in response to the complete clearing of the SRQ (229). [10] Data processing system according to claim 7, wherein the delay (330) is proportional to a time period required to transmit the TLBI instruction (302) from the first processing element via the connection to the second processing element. [11] Data processing system according to claim 7, wherein the delay (330) is based on a product of a number of cycles required to clear each entry from the SRQ (229) and a size of the SRQ (229). [12] Data processing system according to claim 7, wherein a number of cycles in the delay (330) is based on a random number between one and a multiple of the size of the SRQ (229). [13] Data processing system according to claim 12, wherein the one or more execution units comprise a random number generator configured to generate the random number. [14] Data processing system according to claim 13, wherein the random number generator is implemented by one or more linear feedback shift registers. [15] Processing element that includes: a processor pipeline (210) comprising one or more load store units (LSUs) (228) configured to execute load and store instructions, wherein the one or more LSUs (228) are configured to: impose a delay (330) on a cleanup cycle of a memory sort queue (SRQ) (229) of the processing element; clean up the SRQ (229) during the delayed cleanup cycle; and a disable address translation buffer (TLBI) instruction (302) from a connection linking a plurality of processing elements, wherein the TLBI instruction (302) is an instruction to disable an address translation buffer (TLB) entry (318,328) corresponding to a virtual memory page and / or a physical memory section, wherein the TLBI instruction (302) is broadcast by another processing element, where by applying the delay (330) to the cleanup cycle of the SRQ (229) the amount of overlap between the cleanup cycle of the SRQ (229) and a deactivation cycle associated with the TLBI (302) is increased. [16] Processing element according to claim 15, wherein the one or more LSUs (228) are configured such that they: detect an SRQ entry in SRQ (229) that is associated with the TLB entry (318,328) to be deactivated; and the detected SRQ entry was removed from SRQ (229). [17] Processing element according to claim 15, wherein the one or more LSUs (228) are configured to send an acknowledgment signal via the connection to the other processing element in response to a complete clearing of the SRQ (229). [18] Processing element according to claim 15, wherein the delay (330) is proportional to a time period required to transmit the TLBI instruction (302) from the other processing element via the connection to the processing element. [19] Processing element according to claim 15, wherein the delay (330) is based on a product of a number of cycles required to clear each entry from the SRQ (229) and a size of the SRQ. [20] Processing element according to claim 15, wherein a number of cycles in the delay (330) is based on a random number between one and a multiple of the size of the SRQ (229).

Citation Information

Patent Citations

  • Lazy flushing of translation lookaside buffers

    US20050114607A1

  • Apparatus and method for low-overhead synchronous page table updates

    US20170286300A1

  • Hazard detection of out-of-order execution of load and store instructions in processors without using real addresses

    US20190108034A1

  • System and method for invalidating an entry in a translation unit

    US6338128B1