Data cache access method and apparatus

By employing a blocking strategy for cache sets and paths in multi-core processors, data cache access address conflicts are precisely handled, resolving the memory access pipeline blocking problem caused by address conflicts in multi-core processors and improving the parallelism and performance of the processor.

CN119336659BActive Publication Date: 2025-12-26BEIJING VCORE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411898639.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-12-26
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

In multi-core processors, memory access pipeline blockage caused by data cache access address conflicts reduces processor parallelism and performance improvement.

Method used

A blocking strategy based on cache sets and paths is adopted. By combining cache sets and paths for blocking judgment, blocking is only implemented in specific pipeline stages, reducing the frequency of unnecessary blocking.

Benefits of technology

Precise handling of data cache access address conflicts reduces blocking caused by conflicts, improves the parallel processing capability of data cache access, and enhances processor performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119336659B_ABST
    Figure CN119336659B_ABST
Patent Text Reader

Abstract

The application provides a data cache access method and device, which is applied to the technical field of computer processors. The method comprises the following steps: obtaining an access request; when it is not determined that the access request operates on a way of data cache access, adopting a cache group blocking strategy to perform blocking judgment; when it is determined that the access request operates on a way of data cache access, adopting the cache group blocking strategy and a cache way blocking strategy to perform blocking judgment; wherein the cache group blocking strategy is that if group addresses are the same, the request is blocked, otherwise the request is not blocked; and the cache way blocking strategy is that if way addresses are the same, the request is blocked, otherwise the request is not blocked.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer processors, and particularly relates to a data cache access method and device. BACKGROUND

[0002] With the evolution of processor technology from single-core to multi-core, the load and store instructions in the processor usually access the data cache first to improve data access efficiency, however, cache invalidation may occur during the access process, and in order to ensure data cache consistency, there are external consistency requests from other processor cores, and there are also various situations that may cause data cache access address conflicts.

[0003] The current relatively simple method for handling data cache access address conflicts is to block the access operation of the subsequent conflict address. When cache access invalidation occurs, data is retrieved from the main memory or other lower-level storage systems, and the original cache block in the data cache is replaced and filled according to strategies such as least recently used (LRU). For cache consistency problems in multi-core processors, the inter-core external consistency request mechanism is used to maintain consistency.

[0004] However, due to the long processing time of data cache access, especially cache invalidation, this blocking operation will cause the memory pipeline to be blocked, thereby greatly reducing the parallelism of program execution and ultimately hindering the improvement of processor performance. SUMMARY

[0005] The present application provides a data cache access method and device to solve the problem of data cache access address conflicts.

[0006] The present application provides a data cache access method, comprising: obtaining an access request; in the case where the access request operates on a data cache access path is not determined, a cache set blocking strategy is used for blocking judgment; in the case where the access request operates on a data cache access path is determined, a cache set blocking strategy and a cache path blocking strategy are used for blocking judgment; wherein the cache set blocking strategy is that if the set address is the same, the request is blocked, otherwise the request is not blocked; the cache path blocking strategy is that if the path address is the same, the request is blocked, otherwise the request is not blocked.

[0007] According to the data cache access method provided in the application, the cache group blocking strategy is used for blocking judgment in the case that the access request operates on the road of the data cache access; the cache group blocking strategy and the cache road blocking strategy are used for blocking judgment in the case that the access request operates on the road of the data cache access, including: when the access request enters the main pipeline, if the group address of the access request satisfies the first condition, the access request is blocked; wherein the first condition includes that the group address of the access request is same as the group address of the flow water level after the main pipeline access data cache mark and label, or the group address of the access request is same as the group address of the flow water level after the replacement pipeline reads the label and data, or the group address of the access request is same as the group address of the filling pipeline.

[0008] According to the data cache access method provided in the application, the cache group blocking strategy is used for blocking judgment in the case that the access request operates on the road of the data cache access; the cache group blocking strategy and the cache road blocking strategy are used for blocking judgment in the case that the access request operates on the road of the data cache access, including: when the access request enters the replacement pipeline to replace the target cache block, if the group address of the access request is same as the group address of the flow water level after the main pipeline access request accesses data cache mark and label, or is same as the group address of the flow water level after the main pipeline access request accesses data cache mark and label and is the same road, the access request is blocked.

[0009] According to the data cache access method provided in the application, the cache group blocking strategy is used for blocking judgment in the case that the access request operates on the road of the data cache access; the cache group blocking strategy and the cache road blocking strategy are used for blocking judgment in the case that the access request operates on the road of the data cache access, including: when the access request enters the filling pipeline, if the access request is same as the group address of the flow water level after the main pipeline access request reads data and is the same road, the access request is blocked.

[0010] According to the data cache access method provided in the application, the cache group blocking strategy is used for blocking judgment in the case that the access request operates on the road of the data cache access; the cache group blocking strategy and the cache road blocking strategy are used for blocking judgment in the case that the access request operates on the road of the data cache access, comprising: when the access request enters the memory invalidation queue, if the access request and any one of the existing items in the memory invalidation queue are the same group address and are the same road of replacement filling, the access request is blocked.

[0011] The application further provides a data cache access device, comprising the following modules: an acquisition module and a processing module; the acquisition module is used for acquiring an access request; the processing module is used for, in the case that the access request operates on the road of the data cache access, adopting a cache group blocking strategy for blocking judgment; in the case that the access request operates on the road of the data cache access, adopting the cache group blocking strategy and a cache road blocking strategy for blocking judgment; wherein the cache group blocking strategy is that if the group address is the same, the request is blocked, otherwise the request is not blocked; the cache road blocking strategy is that if the road address is the same, the request is blocked, otherwise the request is not blocked.

[0012] According to the data cache access device provided in the application, the processing module is used for, when the access request enters the main pipeline, if the group address of the access request satisfies a first condition, blocking the access request; wherein the first condition comprises that the group address of the access request is the same as the group address of the flow water level after the data cache access tag of the main pipeline is accessed, or the group address of the access request is the same as the group address of the flow water level after the data of the replacement pipeline is read and tagged, or the group address of the access request is the same as the group address of the filling pipeline.

[0013] According to the data cache access device provided in the application, the processing module is used for, when the access request enters the replacement pipeline to replace the target cache block, if the group address of the access request is the same as the group address of the flow water level of the data cache access tag and the tag of the main pipeline access request, or is the same as the group address of the flow water level after the flow water level of the data cache access tag and the tag of the main pipeline access request, and is the same road, blocking the access request.

[0014] According to the data cache access device provided in the application, the processing module is used for, when the access request enters the filling pipeline, if the group address of the access request is the same as the group address of the flow water level of the data of the main pipeline access request, and is the same road, blocking the access request.

[0015] According to the data cache access device provided in the application, the processing module is configured to block the access request when the access request enters the memory access invalidation queue and the access request and any existing item in the memory access invalidation queue are replacement fills of the same group address and are replacement fills of the same path.

[0016] The application further provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the data cache access method according to any of the above when executing the program.

[0017] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program is executable on a processor to implement the data cache access method according to any of the above.

[0018] The application further provides a computer program product including a computer program, and the computer program is executable on a processor to implement the data cache access method according to any of the above.

[0019] The data cache access method and device provided in the application can accurately process data cache access address conflicts, reduce blocking caused by the conflicts, improve the parallel processing capability of data cache access address conflicts, and ultimately improve the performance of a processor. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0021] Figure 1 is a structural schematic diagram of a cache memory provided in the application;

[0022] Figure 2 is a pipeline architecture diagram of a data cache system provided in the application;

[0023] Figure 3 is a flowchart of a data cache access method provided in the application;

[0024] Figure 4 is a structural schematic diagram of a data cache access device provided in the application;

[0025] Figure 5 is a structural schematic diagram of an electronic device provided in the application. DETAILED DESCRIPTION

[0026] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0027] It should be noted that in the embodiments of the present application, the words such as "exemplary" or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or having more advantages than other embodiments or design schemes. Rather, the words such as "exemplary" or "for example" are used in the specific manner to present the relevant concept.

[0028] It should be noted that in the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive containing, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but can also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0029] In order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, the same items or similar items with basically the same functions and roles are distinguished by using "first", "second", etc. The skilled in the art can understand that "first", "second", etc. are not limited in number and execution order.

[0030] Some exemplary embodiments are described in the embodiments of the present application for illustrative purposes. It should be understood that the present application can be implemented in other ways not specifically shown in the drawings.

[0031] With the development of computer processor technology from single-core to multi-core, the performance of processor storage system has a more significant impact on the overall performance. When the processor executes the load instruction and store instruction, it involves the access to the data cache (DCache). Both instructions can cause the data cache access miss, and for the miss case, the data needs to be retrieved from the next level of storage system, which involves the replacement of the original block in the data cache and the refill operation with the retrieved data. Moreover, the data cache of the multi-core processor will also receive external consistency requests (Probe) from other processor cores to maintain cache coherence.

[0032] These different types of access to the data cache, the cache block address accessed may be the same case, such as the same tag and index bits, meaning the same cache block access; or although the cache block address is different, the index bits of the access address are the same, that is, the same cache set operation. As long as the cache set address is the same (i.e. the index bits are the same), it is possible to operate on the same cache block, which may cause data cache access conflict and lead to data and state inconsistency, so the data cache access address conflict must be processed.

[0033] The simplest way to handle data cache access address conflict is the blocking strategy, that is, only the operation of the conflict is processed before the previous operation of the conflict is solved, and the access operation of the conflict address is blocked. This way is relatively easy to implement, but has obvious disadvantages. Because the data cache access processing time is often long, especially for handling data cache miss. During this period, if other operations are blocked from execution, the memory pipeline will be blocked, which greatly reduces the parallelism of processor program execution, thereby hindering the improvement of processor performance.

[0034] The processing performance of data cache access request is crucial to the performance of processor, and there are many points to be balanced in the design process. Since the access addresses of multiple Load instructions, multiple Store instructions to the data cache, the replacement addresses and filling addresses of the data returned by the lower storage system to the data cache, and the addresses of the external consistency requests of other processor cores to the data cache in the multi-core processor may have the same cache block address or operation on the same cache group, thereby causing access address conflict. The processing of these conflict operations can be performed in series or in parallel, and the processing efficiency and parallelism of data cache access address conflict are particularly important for the out-of-order execution performance of high-performance processor and the multi-core parallelism of multi-core processor. This means that when designing the data cache access request processing mechanism, it is necessary to consider how to avoid the problems caused by conflicts while improving the processing efficiency and maintaining high parallelism as much as possible to improve the overall performance of the processor.

[0035] To solve the above problems, an embodiment of the present application provides a data cache access method, which combines cache set (Cache Set) and way (Way) to make blocking judgment, and only implements blocking at a specific pipeline stage, so as to make the blocking more fine-grained and reduce the frequency of unnecessary blocking. In order to clearly describe the method provided by the embodiment of the present application, the technologies involved in the embodiment of the present application are introduced in detail below.

[0036] 1. Function and principle of Cache (cache memory)

[0037] Cache is located between the processor and the memory, and is usually composed of SRAM (static random access memory). The operating speed of the processor is extremely fast, and the access speed of the memory is much slower. When the processor directly accesses data from the memory, it needs to wait for many clock cycles. Cache has fast access speed, and it can store part of the data that the processor has just used or will be used repeatedly. In this way, when the processor needs this part of data again, it can directly retrieve it from Cache, without the need to obtain it from the memory with long delay, thereby effectively reducing the waiting time of the processor and greatly improving the overall efficiency of the system.

[0038] 2. Structure of Cache

[0039] Cache is mainly composed of two parts, namely tag (Tag) part and data (Data) part. The Data part is responsible for saving a piece of continuous address data, and the Tag part is used to store the public address of the continuous data.

[0040] A tag and all its corresponding data form a single line, which is called a cache line. The data portion of a cache line is called a data block. If data can be stored in multiple locations within the cache, these multiple cache lines, accessible through the same address index, are called a cache set.

[0041] 3. Classification of Cache Composition Methods

[0042] There are three ways to configure a cache: direct associativity, set associativity, and full associativity. This application focuses on set associativity. In fact, direct associativity and full associativity can be regarded as special set associativity configurations with a path count of 1 and a path count equal to the number of cache lines, respectively.

[0043] 4. Address processing in group-associative mode

[0044] like Figure 1 As shown, in a set-associative cache structure, the processor's memory access address is split into three parts: tag, index, and block offset, used to locate data in the cache. The cache consists of multiple cache ways and cache sets. Each cache set contains multiple cache lines. The cache line number identifies the specific line in the cache, and the byte number locates the specific byte within that cache line. Tag RAM stores the tag corresponding to the data, indicating which location in main memory the data in the cache line originated from. Data RAM stores the data loaded from main memory. By comparing the tag in the access address with the tag in the tag RAM, a hit can be determined. If a hit occurs, data can be read from the data RAM. If multiple ways exist, the way number of the data storage needs to be determined.

[0045] An index can be used to find a set of cache lines, or a cache set, in the cache. Then, the tag portion read using the index is compared with the tag in the access address. Only when the two are equal does it mean that this cache line is the required one.

[0046] There are many access data in a Cache Line, with the help of Block Offset part in memory address and access width of access instruction, the real wanted data can be found accurately, which can locate to each byte.

[0047] 5. Tag bit (Meta) and state of Cache Line in multi-core processor

[0048] There is also a tag bit (Meta) in Cache Line of multi-core processor, which mainly marks whether the data is valid, whether the data is dirty (i.e. whether it has been modified), and marks the cache consistency state of multi-core processor. If there is ECC (Error Checking and Correction) check, it is usually placed in the tag storage array (Meta Array).

[0049] Each Cache block of data Cache in multi-core processor has the following three states:

[0050] INV (Invalid state): indicates that the corresponding data Cache block is in invalid state, i.e. the data of the block is currently unavailable.

[0051] SHD (Shared state): indicates that the corresponding data Cache block is in shared state, in which state the processor core can directly hit when reading the block, but if you want to write the block, write invalid will occur.

[0052] EXC (Exclusive state): indicates that the corresponding data Cache block is in exclusive state, at this time the processor reads and writes the block directly hit. And each Cache block also has a dirty bit (w bit) indicating whether the block has been written, if the block has been written, the w bit is set to 1. It should be noted that the write invalid operation of a Cache block will first set it to EXC state when retrieving the corresponding Cache block from the lower storage system, only when the Cache block is really written, the w bit will be set to 1, at this time the EXC state plus the w bit, which represents the M (Modify) state.

[0053] The following will be combined Figure 2 In the system of data cache processing, the functions, operation processes and mutual relations of different pipelines, queues and storage arrays are described in detail.

[0054] 1. Access pipeline of load instruction (Load)

[0055] Access procedure and clock cycle: The access procedure of Load instruction is divided into 3 beats (each beat corresponds to a clock cycle of the processor). The first beat reads Meta and Tag, which may be used to determine the location of the data to be obtained in the cache and other related information; the second beat reads Data; and the third beat performs Sel Data operation, i.e., selects the required part from the read data.

[0056] Hit and miss processing: If the Load instruction hits the data cache, the required data is directly returned. If the Load instruction misses the data cache, it will enter the Miss Queue after the 3 beats. Meanwhile, on the Load instruction access pipeline, the Way of the cache block to be replaced is determined, and the information of the replaced block is sent to the Miss Queue for subsequent processing.

[0057] 2. Main pipeline (Main Pipe)

[0058] Processing task and access procedure: The main pipeline is mainly responsible for processing Store instruction and Probe, and their access procedures are both 4 beats. The first beat reads Meta and Tag, which is similar to the first beat of Load instruction, and is used to preliminarily locate the data. The second beat reads Data. The third beat obtains the result of reading Data and combines it with the data to be written. The fourth beat updates Meta, Tag and Data according to the operation result. For Probe, if it is necessary to initiate write-back data to the next level storage system, an access request to the Writeback Queue is generated at this level, and the Writeback Queue is written.

[0059] Subsequent processing of Store instruction: After the Store instruction is submitted, the data in the Store instruction is moved to the Store Buffer. The Store Buffer combines the Store write requests in units of cache lines, and when it is close to full, it writes the combined multiple Store write requests to the data cache.

[0060] Miss Handling: Similar to Load instruction, if a Store instruction misses the data cache, it will enter the Miss Queue after 3 ticks. The way of the cache block to be replaced is decided in the Store instruction pipeline, and the information of the replaced block is sent to the Miss Queue.

[0061] 3. Refill Pipe

[0062] Refill Function and Operation Flow: The main function of the Refill Pipe is to refill the Load instruction that misses. It only takes 1 tick to write the data to the DCache after getting the refill data, thus completing the refill of the missing data due to the miss of the Load instruction.

[0063] 4. Replace Pipe

[0064] Replace Operation Flow: The Replace Pipe is responsible for reading the data of the block to be replaced, which takes 2 ticks. The first tick is to read the Meta and Data, which may be to obtain the relevant information of the block to be replaced for subsequent processing. The second tick invalidates the Meta and generates a Release request to the Writeback Queue, thus completing the processing of the block to be replaced and making its data state meet the replacement requirements.

[0065] 5. Miss Queue

[0066] Function and Processing Flow: The function of the Miss Queue is to retrieve the cache block or permission that needs to be refilled. For example, when a Store instruction accesses the cache block in the SHD state, it may need to retrieve the EXC permission. After retrieving these required contents, they are sent to the Refill Pipe to complete the data refill operation.

[0067] 6. Probe Queue

[0068] Functionality and Data Flow: The Probe Queue primarily accepts external consistency requests (Probes) and then forwards them to the Main Pipe. These external consistency requests are requests from other processor cores to access the data cache in order to maintain cache coherence for the multi-core processor.

[0069] 7. Writeback Queue

[0070] Processing tasks: The Writeback Queue mainly handles write-back (Release) and responses to external consistency requests (ProbeAck), playing a related role in processing write-back operations and responses to external consistency requests throughout the entire data processing flow.

[0071] 8. Data RAM, Tag RAM, and MetaArray

[0072] Storage Media and Characteristics: Data RAM and tag RAM are generally implemented using SRAM (Static Random Access Memory). MetaArray is generally implemented using registers. There is a relationship where writing a tag always writes the meta tag, but sometimes only the meta tag is written without writing the tag. This is related to different operational scenarios and data processing needs; for example, in some cases, only the tag-related information needs to be updated without updating the tag itself.

[0073] like Figure 3 As shown, this application provides a data cache access method, which can be applied to a data cache access device. The data cache access method may include steps S301-S302:

[0074] S301, The data cache access device obtains an access request.

[0075] Optionally, the above access request can be a data cache modification request or a data cache non-modification request. The data cache modification request can be one of the following: a data storage instruction access request, an external consistency request, a replacement request, a fill request, a memory access invalidation queue request, and a write-back queue request. The data cache non-modification request can be a data fetch instruction access request.

[0076] It should be noted that the modified data cache request refers to the request that needs to be blocked, and the unmodified data cache request refers to the request that does not need to be blocked.

[0077] It should be noted that the present application does not block the access request of the fetch instruction pipeline, because the access of the fetch instruction does not involve the modified data cache, and therefore does not need to be blocked, so as to ensure the execution efficiency of the fetch instruction.

[0078] S302, the data cache access device determines whether the access request operates on the way of the data cache access, and judges the blocking according to the cache set blocking strategy; if it is determined that the access request operates on the way of the data cache access, the blocking is judged according to the cache set blocking strategy and the cache way blocking strategy.

[0079] The cache set blocking strategy is to block the request if the set address is the same, and otherwise not to block the request; and the cache way blocking strategy is to block the request if the way address is the same, and otherwise not to block the request.

[0080] Optionally, when the access request enters the main pipeline, if the set address of the access request satisfies a first condition, the access request is blocked; wherein the first condition includes that the set address of the access request is the same as the set address of the flow water level after the main pipeline access data cache mark and label, or the set address of the access request is the same as the set address of the flow water level after the replacement pipeline reads the label and data, or the set address of the access request is the same as the set address of the filling pipeline.

[0081] For example, the first condition can be that the set address of the access request is the same as the set address of the 2nd, 3rd and 4th taps of the main pipeline, or the set address of the access request is the same as the set address of the 2nd tap of the replacement pipeline, or the set address of the access request is the same as the set address of the filling pipeline.

[0082] Specifically, when the store instruction (Store) access request and the external consistency request (Probe) enter the main pipeline MainPipe, if the set address of the access request is the same as the set address of the 2nd, 3rd and 4th taps of the main pipeline or the set address of the 2nd tap of the replacement pipeline or the set address of the filling pipeline, the access request will be blocked. Because the cache way (Way) operated at this time is not determined, the blocking is judged according to the set address which is a relatively coarse-grained condition.

[0083] The first stage of the main pipeline usually does not need to be blocked. The reason is that even if the store instruction is hit at the first stage of the main pipeline, and the reading of the meta and tag is interrupted by a write between the reading of the meta and tag and the reading of the data, the data read will not be inconsistent with the meta and tag read in the previous stage.

[0084] Optionally, when the access request enters the replacement pipeline to replace the target cache block, if the set address of the access request is the same as the set address of the data cache tag and tag pipeline stage accessed by the main pipeline access request, or is the same as the set address of the pipeline stage after the data cache tag and tag pipeline stage accessed by the main pipeline access request and is the same way, the access request is blocked.

[0085] For example, when the access request enters the replacement pipeline to replace the target cache block, if the set address of the access request is the same as the set address of the first stage of the main pipeline access request, or is the same as the set address of the second, third, or fourth stage of the main pipeline access request and is the same way, the access request is blocked.

[0086] Specifically, when the replacement pipeline Replace Pipe enters to replace the target cache block A, if the set address is the same as the first stage of the main pipeline access request, or is the same as the set address of the second, third, or fourth stage of the main pipeline access request and is the same way, the request to replace A will be blocked. Here, it is considered that the cache way (Way) of the replacement block has been determined when the load and store instructions are invalidated, so in the case of different main pipeline stages, different precision blocking judgments are used according to whether the way is determined (the first stage according to the set address, and the second, third, or fourth stage according to the set + way).

[0087] Optionally, when the access request enters the refill pipeline, if the access request is the same as the set address of the data pipeline stage accessed by the main pipeline access request and is the same way, the access request is blocked.

[0088] For example, when the access request enters the refill pipeline, if the access request is the same as the set address of the second stage of the main pipeline access request and is the same way, the access request is blocked.

[0089] Specifically, when the refill pipeline Refill Pipe enters, if the refill request is the same as the set address of the second stage of the main pipeline access request and is the same way, it will be blocked. Because the Refill Pipe determines the written way (Way) and only affects the data read by the Main Pipe, it only needs to make such a blocking judgment with the second stage of the main pipeline reading data.

[0090] Optionally, when the access request enters the access invalidation queue, if the access request and any one of the existing items in the access invalidation queue are replacement fills of the same group address and are replacement fills to the same way, the access request is blocked.

[0091] Specifically, when the Miss Queue is enqueued, if the invalidation request to be entered and the existing item in the Miss Queue are replacement fills of the same group address and are replacement fills to the same way, the invalidation request is blocked from entering the Miss Queue; for requests of the same address, a flexible processing mode of being directly passed to the Load pipeline or being merged into an access invalidation queue item to wait for backfill is adopted.

[0092] It should be noted that in the access of the data cache, there are only 3 writes, which are the 4th beat of the Main Pipe, the 2nd beat of the Replace Pipe and the Refill Pipe. For the 3 write cases, the relationship between each pipeline read operation is analyzed, and the above blocking strategy is used to ensure that there is no problem of data inconsistency caused by write operation interrupting read operation between different pipelines. For example, the Set blocking mode of the Main Pipe can avoid related problems; the blocking strategy between the Replace Pipe and the Main Pipe ensures that the Store and the Replace of the same address cannot exist in the two pipelines at the same time, so as to ensure that the read mark Meta and the read tag Tag and the read Data cannot be interrupted; the Refill Pipe will not have the write interrupting read situation because it is blocked by the 2nd beat Set + Way of the Main Pipe when entering.

[0093] In the embodiments of the present application, the data cache access address conflict can be accurately processed, the blocking situation caused by the conflict is reduced, the parallel processing capability of the data cache access address conflict is improved, and finally the performance of the processor is improved.

[0094] The above mainly introduces the scheme provided by the embodiments of the present application from the perspective of method. In order to realize the above functions, it contains the hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed in the present text, the embodiments of the present application can be realized in the form of hardware or the combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraint conditions of the technical scheme. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0095] It should be noted that the apparatus in the embodiments of this application includes a virtual device and a physical device. The virtual device can be a data cache access device, and the physical device can include electronic devices, computer storage media, and computer program products.

[0096] The data cache access method provided in this application can be executed by a data cache access device or a control module for data cache access within that device. This application uses the execution of the data cache access method by a data cache access device as an example to illustrate the data cache access device provided in this application.

[0097] It should be noted that the embodiments of this application can divide the data cache access device into functional modules according to the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. Optionally, the module division in the embodiments of this application is illustrative and is only a logical functional division; other division methods may be used in actual implementation.

[0098] like Figure 4 As shown in the figure, this application embodiment provides a data cache access device 400. The data cache access device 400 includes an acquisition module 401 and a processing module 402. The acquisition module 401 can be used to acquire access requests; the processing module 402 can be used to perform blocking judgment using a cache group blocking strategy when the path operated by the access request on the data cache is not determined; and to perform blocking judgment using both the cache group blocking strategy and the cache path blocking strategy when the path operated by the access request on the data cache is determined. The cache group blocking strategy blocks the request if the group addresses are the same, and does not block the request otherwise; the cache path blocking strategy blocks the request if the path addresses are the same, and does not block the request otherwise.

[0099] Optionally, the processing module 402 described above can be used to block the access request when the access request enters the main pipeline if the group address of the access request meets a first condition; wherein the first condition includes: the group address of the access request is the same as the group address of the pipeline level after the access data cache tag and label in the main pipeline, or the group address of the access request is the same as the group address of the pipeline level after the replacement pipeline reads the tag and data, or the group address of the access request is the same as the group address of the filling pipeline.

[0100] Optionally, the processing module 402 can be configured to, when the access request enters the replacement pipeline to replace the target cache block, if the group address of the access request is the same as the group address of the data cache tag and label pipeline accessed by the main pipeline access request, or is the same as the group address of the pipeline after the data cache tag and label pipeline accessed by the main pipeline access request and is the same way, block the access request.

[0101] Optionally, the processing module 402 can be configured to, when the access request enters the fill pipeline, if the access request is the same as the group address of the data pipeline accessed by the main pipeline access request and is the same way, block the access request.

[0102] Optionally, the processing module 402 can be configured to, when the access request enters the memory invalidation queue, if the access request is the same group address and the same way as the replacement fill of any one of the existing items in the memory invalidation queue, block the access request.

[0103] In the embodiment of the application, the data cache access address conflict can be accurately processed, the blocking caused by the conflict can be reduced, the parallel processing capability of the data cache access address conflict can be improved, and the performance of the processor can be improved.

[0104] Figure 5 An example of a schematic diagram of the physical structure of an electronic device is shown in Figure 5 As shown, the electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 can communicate with each other through the communications bus 540. The processor 510 can invoke a logical instruction in the memory 530 to execute a data cache access method, which includes: obtaining an access request; in a case where the way operated by the access request to the data cache is not determined, performing blocking judgment by using a cache group blocking strategy; in a case where the way operated by the access request to the data cache is determined, performing blocking judgment by using the cache group blocking strategy and a cache way blocking strategy; wherein the cache group blocking strategy is to block the request if the group addresses are the same, and otherwise not to block the request; and the cache way blocking strategy is to block the request if the way addresses are the same, and otherwise not to block the request.

[0105] Further, the logic instructions in the memory 530 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0106] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, so that the computer can execute the data cache access method provided by the above-mentioned methods. The method comprises: obtaining an access request; in the case where it is not determined that the access request operates on a way of data cache access, adopting a cache group blocking strategy to perform blocking judgment; in the case where it is determined that the access request operates on a way of data cache access, adopting the cache group blocking strategy and a cache way blocking strategy to perform blocking judgment; wherein the cache group blocking strategy is that if the group addresses are the same, the request is blocked, otherwise the request is not blocked; and the cache way blocking strategy is that if the way addresses are the same, the request is blocked, otherwise the request is not blocked.

[0107] In another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the data cache access method provided by the above-mentioned methods. The method comprises: obtaining an access request; in the case where it is not determined that the access request operates on a way of data cache access, adopting a cache group blocking strategy to perform blocking judgment; in the case where it is determined that the access request operates on a way of data cache access, adopting the cache group blocking strategy and a cache way blocking strategy to perform blocking judgment; wherein the cache group blocking strategy is that if the group addresses are the same, the request is blocked, otherwise the request is not blocked; and the cache way blocking strategy is that if the way addresses are the same, the request is blocked, otherwise the request is not blocked.

[0108] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0109] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary general hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0110] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A data cache access method, characterized by, The method comprises the following steps: acquiring an access request, wherein the access request is a modified data cache request or an unmodified data cache request; the modified data cache request is a request that needs to be blocked, and the modified data cache request is one of a memory instruction access request, an external consistency request, a replacement request, a padding request, a memory invalidation queue request and a write-back queue request; the unmodified data cache request is a request that does not need to be blocked; in a case where a way operated by the access request on a data cache is not determined, a cache set blocking strategy is used for blocking judgment; in a case where the way operated by the access request on the data cache is determined, the cache set blocking strategy and a cache way blocking strategy are used for blocking judgment; the cache set blocking strategy is that if a set address is the same, the request is blocked, otherwise the request is not blocked; the cache way blocking strategy is that if a way address is the same, the request is blocked, otherwise the request is not blocked.

2. The data cache access method of claim 1, wherein, The method of using the cache set blocking strategy to perform the blocking judgment in the case where the way operated by the access request on the data cache is not determined, and using the cache set blocking strategy and the cache way blocking strategy to perform the blocking judgment in the case where the way operated by the access request on the data cache is determined, comprises the following steps: when the access request enters a main pipeline, if a set address of the access request satisfies a first condition, the access request is blocked; the first condition comprises that the set address of the access request is the same as a set address of a flow water level after a data cache tag of the main pipeline is accessed, or the set address of the access request is the same as a set address of a flow water level after a data of a replacement pipeline is read and tagged, or the set address of the access request is the same as a set address of a padding pipeline.

3. The data cache access method of claim 1, wherein, The method of using the cache set blocking strategy to perform the blocking judgment in the case where the way operated by the access request on the data cache is not determined, and using the cache set blocking strategy and the cache way blocking strategy to perform the blocking judgment in the case where the way operated by the access request on the data cache is determined, comprises the following steps: when the access request enters a replacement pipeline to replace a target cache block, if a set address of the access request is the same as a set address of a flow water level of a data cache tag accessed by a main pipeline access request, or the set address of the access request is the same as a set address of a flow water level after the flow water level of the data cache tag accessed by the main pipeline access request and is the same way, the access request is blocked.

4. The data cache access method of claim 1, wherein, The method of using the cache set blocking strategy to perform the blocking judgment in the case where the way operated by the access request on the data cache is not determined, and using the cache set blocking strategy and the cache way blocking strategy to perform the blocking judgment in the case where the way operated by the access request on the data cache is determined, comprises the following steps: when the access request enters a padding pipeline, if the access request is the same as a set address of a flow water level of a data read by a main pipeline access request and is the same way, the access request is blocked.

5. The data cache access method of claim 1, wherein, The blocking judgment is performed by using the cache group blocking strategy in the case that the access request operates on the data cache access road is not determined, and the blocking judgment is performed by using the cache group blocking strategy and the cache road blocking strategy in the case that the access request operates on the data cache access road is determined, comprising: When the access request enters the memory invalidation queue, if the access request and any one of the existing items in the memory invalidation queue are replacement padding of the same group address and are replacement padding of the same road, the access request is blocked.

6. A data cache access device, characterized by Comprising: An acquisition module and a processing module; The acquisition module is configured to acquire an access request, the access request being a modified data cache request or an unmodified data cache request; the modified data cache request is a request that needs to be blocked, and the modified data cache request is one of a store instruction access request, an external consistency request, a replacement request, a padding request, a memory invalidation queue request, and a write-back queue request; the unmodified data cache request is a request that does not need to be blocked. The processing module is configured to perform blocking judgment by using a cache group blocking strategy in the case that the access request operates on the data cache access road is not determined, and perform blocking judgment by using the cache group blocking strategy and the cache road blocking strategy in the case that the access request operates on the data cache access road is determined; the cache group blocking strategy is that if the group address is the same, the request is blocked, otherwise the request is not blocked; the cache road blocking strategy is that if the road address is the same, the request is blocked, otherwise the request is not blocked.

7. The data cache access apparatus according to claim 6, wherein, The processing module is configured to, when the access request enters the main pipeline, if the group address of the access request satisfies a first condition, block the access request; the first condition includes that the group address of the access request is the same as the group address of the flow water level after the main pipeline access data cache tag and label, or the group address of the access request is the same as the group address of the flow water level after the replacement pipeline reads the label and data, or the group address of the access request is the same as the group address of the padding flow water pipeline.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the data cache access method in any one of claims 1 to 5. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the data cache access method in any one of claims 1 to 5.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the data cache access method in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and readable storage medium

    CN115454887A

  • Access system and method for data storage

    WO2016192045A1