Cached data reading method and apparatus, device, and storage medium
By managing memory access instructions in a concise and clear multi-level pipeline queue architecture in the secondary cache, the problem of high complexity in the calculation advance is solved in the prior art, and the precise control of the fixed advance amount of wake-up requests is achieved, circuit cost and power consumption are reduced, and accuracy and coverage are improved.
Patent Information
- Application Number
- PCT/CN2024/136265
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-26
AI Technical Summary
In the prior art, due to the large number of states of pipeline queues and wake-up request queues, the complexity of the calculation is high in advance is increased, and the cost and power consumption of the circuit are reduced, while the accuracy and coverage are reduced.
By managing memory access instructions in a concise and clear multi-level pipeline queue architecture in the secondary cache, the pipeline queue architecture and the design of each processing time in the pipeline line is based on the pipeline queue architecture and instructions, accurate and stable control of the fixed advance amount of wake-up requests is achieved.
It reduces the complexity and cost of the circuit, reduces power consumption, and improves the accuracy and coverage of wake-up instructions, significantly saving the delay in the memory access instruction reading process.
Smart Images

Figure CN2024136265_26062025_PF_FP_ABST
Abstract
Description
Cache data reading method, device, equipment and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 20, 2023, with application number 202311764013.1 and invention name “Cache data reading method, device, equipment and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a cache data reading method, apparatus, device, and storage medium. Background Art
[0003] Modern processors generally have three levels of cache: L1 cache, L2 cache and L3 cache. When the processor's memory access instruction misses the L1 cache, L2 cache needs to refill the data requested by the memory access instruction into the L1 cache to ensure normal data access.
[0004] At present, if the processor is awakened to access the L1 cache L1 after the L2 cache L2 refills the data into the L1 cache L1, the access process will incur a long delay. In order to reduce the delay, the L2 cache L2 can read the request status in its own pipeline queue and refill request queue, calculate the time to send the wake-up request in advance based on the request status, and send the wake-up request in advance, that is, wake up the processor to access the L1 cache L1 in advance, instead of waiting for the refill to be completed before waking up, thereby reducing the access delay.
[0005] However, in the above process, due to the large number of states of the pipeline queue and the wake-up request queue, the calculation of the advance issuance time is extremely complex, which makes the cost and power consumption of the implementation circuit high. If the number of considered states is reduced in order to reduce complexity, the accuracy and coverage will decrease. Summary of the Invention
[0006] The embodiments of the present application provide a cache data reading method, apparatus, device and storage medium to solve the problems in the related art.
[0007] In a first aspect, an embodiment of the present application provides a method for reading cached data, the method comprising:
[0008] When it is determined that the memory access instruction of the processor does not hit in the first-level cache, controlling the memory access instruction to enter the pipeline queue from the starting data bit of the pipeline queue of the second-level cache; the pipeline queue includes a plurality of data bits arranged in sequence;
[0009] After a first number of data bits have passed, obtaining a hit result of the memory access instruction in the secondary cache;
[0010] If the hit result is a hit, a wake-up instruction is immediately generated through the secondary cache and issued by the wake-up queue, and after an interval of a second number of data bits, refill data is obtained through the secondary cache, and the memory access instruction is issued as a refill instruction by the refill queue;
[0011] The wake-up instruction is used to wake up the operation of reading data from the first-level cache through the memory access instruction; the refill instruction is used to write the refill data into the first-level cache for reading by the memory access instruction.
[0012] In a second aspect, an embodiment of the present application provides a cache data reading device, the device comprising:
[0013] an enqueue module, configured to, when determining that a memory access instruction of the processor does not hit in the first-level cache, control the memory access instruction to enter the pipeline queue from the starting data bit of the pipeline queue of the second-level cache; the pipeline queue includes a plurality of data bits arranged in sequence;
[0014] A judgment module, configured to obtain a hit result of the memory access instruction in the secondary cache after an interval of a first number of data bits;
[0015] a dequeue module, configured to, if the hit result is a hit, immediately generate a wake-up instruction through the secondary cache and issue it from the wake-up queue, and after an interval of a second number of data bits, obtain refill data through the secondary cache and issue the memory access instruction as a refill instruction from the refill queue;
[0016] The wake-up instruction is used to wake up the operation of reading data from the first-level cache through the memory access instruction; the refill instruction is used to write the refill data into the first-level cache for reading by the memory access instruction.
[0017] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor;
[0018] a memory for storing instructions executable by the processor;
[0019] The processor is configured to execute the instructions to implement the method of the first aspect.
[0020] In a fourth aspect, an embodiment of the present application further provides a computer program comprising a computer-readable code, which, when executed on a computing and processing device, causes the computing and processing device to execute the method of the first aspect.
[0021] In the fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores the computer program as described in the fourth aspect. When the instructions in the computer-readable storage medium are executed by the processor of an electronic device, the electronic device is able to perform the method of the first aspect.
[0022] In an embodiment of the present application, when it is determined that the memory access instruction of the processor does not hit in the first-level cache, the memory access instruction is controlled to enter the pipeline queue from the starting data bit; and after an interval of a first number of data bits, the hit result of the memory access instruction in the second-level cache is obtained; if the hit result is a hit, a wake-up instruction is immediately generated through the second-level cache and issued by the wake-up queue, and after an interval of a second number of data bits, refill data is obtained through the second-level cache, and the memory access instruction is issued as a refill instruction by the refill queue. The present application implements the management of memory access instructions through the concise and clear multi-level pipeline queue architecture of the second-level cache. Based on the architecture of the pipeline queue and the design of each processing timing of the instruction in the pipeline, it can achieve accurate and stable control of the fixed advance amount of the wake-up request, thereby ensuring the accuracy and coverage of the wake-up instruction. The entire process does not require reading the status of the requests at each level of the pipeline, and calculating the advance issuance time in real time based on the status of the refill queue request, so the complexity is extremely low, reducing the cost and power consumption of the circuit.
[0023] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0025] FIG1 is an architecture diagram of an implementation scenario provided by an embodiment of the present application;
[0026] FIG2 is a flowchart of a method for reading cached data provided by an embodiment of the present application;
[0027] FIG3 is a flowchart showing the specific steps of a cache data reading method provided in an embodiment of the present application;
[0028] FIG4 is a block diagram of a cache data reading device provided in an embodiment of the present application;
[0029] FIG5 schematically shows a block diagram of a computing and processing device for executing the method according to the present application;
[0030] FIG6 schematically illustrates a storage unit for holding or carrying program codes for implementing the method according to the present application;
[0031] FIG7 is a block diagram of an electronic device provided in an embodiment of the present application;
[0032] FIG8 is a block diagram of another electronic device according to another embodiment of the present application. Specific embodiments
[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In the embodiments of this application, the term "multiple" refers to two or more, and other quantifiers are similar.
[0035] Referring to Figure 1, Figure 1 is an architecture diagram of an implementation scenario provided by an embodiment of the present application. In order to improve execution efficiency and reduce the interaction between the processor and the memory, modern processors can integrate a multi-level cache architecture on the processor. The common architecture is the three-level cache structure of Figure 1, including: Level 1 cache L1, Level 2 cache L2 and Level 3 cache L3. Level 1 cache L1 is the cache closest to the processor, with the smallest capacity and the fastest speed; Level 2 cache L2 has a larger capacity but is slower than Level 1 cache L1. Level 2 cache L2 is the buffer of Level 1 cache L1. The function of Level 2 cache L2 is to store data that is needed for processor processing but cannot be stored by Level 1 cache L1; Level 3 cache L3 has the largest capacity and is also the slowest level. Level 3 cache L3 and memory can be regarded as buffers of Level 2 cache L2.
[0036] When the processor is running, it will first search the L1 cache for the required data according to the memory access instruction, then the L2 cache, and then the L3 cache. If the data is not found in the L3 cache, it will be retrieved from the main memory. The longer the search path, the longer it takes. Therefore, if certain data needs to be retrieved frequently, it is best to ensure that this data is in the L1 cache, so that the speed will be very fast. Among them, the memory access instruction is an instruction to obtain data from a specified address in the main memory, or to store data to a specified address in the main memory.
[0037] Specifically, the secondary cache L2 can adopt a 5-level pipeline architecture, that is, it includes a pipeline queue with 5 sequentially arranged data bits, each data bit corresponds to a pipeline moment, different data bits correspond to different pipeline moments, and the request is used to enter the pipeline from the initial data bit of the pipeline queue, and change the data bit as time goes by; the secondary cache L2 can also include a wake-up queue and a refill queue, the refill queue is used to send a refill request to the first-level cache L1, so as to refill the data required for the memory access request but not stored in the first-level cache L1 from the second-level cache L2 into the first-level cache L1; the wake-up queue is used to send a wake-up request to the first-level cache L1 to wake up the processor to access the first-level cache L1 based on the memory access instruction.
[0038] FIG2 is a flowchart of a method for reading cached data provided by an embodiment of the present application. As shown in FIG2 , the method may include:
[0039] Step 101: When it is determined that a memory access instruction of a processor does not hit in the first-level cache, control the memory access instruction to enter the pipeline queue from the starting data bit of the pipeline queue of the second-level cache; the pipeline queue includes multiple data bits arranged in sequence.
[0040] In the embodiment of the present application, referring to Figure 1, the processor reads data through memory access instructions in the order of the first-level cache L1, the second-level cache L2, the third-level cache L3, and the memory. If the processor's memory access instruction misses in the first-level cache L1, it means that the data requested by the memory access instruction is not stored in the first-level cache L1. At this time, the first-level cache L1 can pass the memory access instruction to the second-level cache L2 to determine whether the second-level cache L2 stores the data. If so, the processor will ask the second-level cache L2 to refill the data into the first-level cache L1.
[0041] Specifically, the L2 cache can employ a multi-stage pipeline architecture, comprising a pipeline queue with multiple sequentially arranged data bits. Each data bit corresponds to a pipeline moment, with different data bits corresponding to different pipeline moments. The pipeline queue's function is to receive and issue instructions and maintain the instruction status according to the timing. Memory access instructions can enter the pipeline at the starting data bit (S1) of the pipeline queue.
[0042] Step 102: After a first number of data bits have passed, obtain a hit result of the memory access instruction in the secondary cache.
[0043] In an embodiment of the present application, after the memory access instruction enters the pipeline of the secondary cache from the starting data bit (S1) of the pipeline queue, it is necessary to further check whether the secondary cache has the data requested to be read by the memory access instruction. The process of determining whether the data exists can be understood as determining the hit result of the memory access instruction in the secondary cache. The hit result includes a hit or a miss. A hit represents that the secondary cache has the data requested to be read by the memory access instruction; a miss represents that the secondary cache does not have the data requested to be read by the memory access instruction.
[0044] Specifically, judging the hit result of the memory access instruction in the secondary cache requires traversing the data in the secondary cache, so it takes a certain amount of time, which is fixed. Referring to Figure 1, the embodiment of the present application can obtain the hit result of the memory access instruction in the secondary cache after the memory access instruction enters the pipeline from the starting data bit (S1) of the pipeline queue, and after the first number of data bits (two data bits), (obtain the hit result at the moment corresponding to the S3 data bit). Since the hit result is obtained, and the hit result includes a hit or a miss, the embodiment of the present application can subsequently execute the corresponding instruction control operation according to the hit result.
[0045] It should be noted that the interval of the first number of data bits (Figure 1 shows an interval of 2 data bits) refers to the time required to wait for the execution of the operation of determining the hit result of the memory access instruction in the secondary cache. Since this time is fixed, and the data bits in the pipeline queue of the secondary cache represent the moment, this time can be converted into the first number of data bits. The interval of the first number of data bits starting from the starting data bit (S1) represents the completion of the execution process of waiting for the determination of whether a hit is achieved after the memory access instruction enters the pipeline of the secondary cache, thereby obtaining the hit result of the memory access instruction in the secondary cache.
[0046] Step 103: If the hit result is a hit, a wake-up instruction is immediately generated through the secondary cache and issued by the wake-up queue, and after an interval of a second number of data bits, refill data is obtained through the secondary cache, and the memory access instruction is issued as a refill instruction by the refill queue.
[0047] The wake-up instruction is used to wake up the operation of reading data from the first-level cache through the memory access instruction; the refill instruction is used to write the refill data into the first-level cache for reading by the memory access instruction.
[0048] In the embodiment of the present application, the hit result is a hit, indicating that the secondary cache has the data requested to be read by the memory access instruction, so the primary cache seeks to let the secondary cache refill the data into the primary cache, and the wake-up instruction issued by the secondary cache can wake up the memory access instruction that missed the primary cache. After waking up, the memory access instruction can quickly read the data from the primary cache, so that the memory access instruction of the processor can normally realize the function of reading data.
[0049] However, if the memory access instruction is awakened to access the first-level cache after the second-level cache refills the data into the first-level cache, the data that has been refilled will have to wait for the memory access instruction to wake up, and the whole process will produce a long delay (that is, after refilling the data, it takes a certain amount of time to wake up the memory access instruction, and the memory access instruction still needs several cycles to actually read the refilled data). In order to reduce this delay, the embodiment of the present application can control the second-level cache to allow it to issue a wake-up instruction in advance before issuing the refill instruction, and ensure that the advance amount is fixed and accurate.
[0050] In order to achieve this purpose, referring to Figure 1, in an embodiment of the present application, for the case where the hit result is a hit, at the moment when the hit result is obtained (the moment corresponding to data bit S3), a wake-up instruction can be immediately generated by the second-level cache and sent by the wake-up queue to the first-level cache, and at the moment when the hit result is obtained (the moment corresponding to data bit S3) after an interval of a second number of data bits (an interval of 2 data bits), the refill data requested to be read by the memory access instruction is obtained at the moment corresponding to data bit S5, and the memory access instruction is sent as a refill instruction by the refill queue to the first-level cache at the moment corresponding to data bit S5, wherein the refill instruction carries the refill data.
[0051] It should be noted that the interval of the second number of data bits refers to the time length represented by waiting for the second number of data bits. The waiting time length is the time length taken by the secondary cache to read the refill data. The waiting time length is fixed, so the time length can be converted into the second number of data bits. The interval of the second number of data bits (interval of 2 data bits) from the moment when the hit result is obtained (the moment corresponding to data bit S3) represents the process of waiting for the secondary cache to obtain the refill data.
[0052] Since the first-level cache receives the wake-up instruction at the moment corresponding to data bit S3, it can start executing the wake-up of the memory access instruction in advance. The first-level cache can immediately read the refill data for the subsequent refill received, so there is no need to wait for the time spent on the wake-up of the memory access instruction, which significantly saves the delay in the memory access instruction reading process.
[0053] For example, referring to Figure 1, when the hit result is a hit, the actual sending time of the wake-up request is the moment corresponding to data bit S3. Since the refill request has to wait in the refill queue for the length of time represented by a data bit (ensuring that the refill request at the exit of the refill queue is issued in time to reduce the probability of congestion in the refill queue), the actual sending time of the refill request is the moment corresponding to data bit S6 (not drawn). It can be seen that the embodiment of Figure 1 of the present application can ensure that for each refill request, there is a wake-up request issued three data bits in advance.
[0054] The embodiment of the present application realizes the management of memory access instructions through the concise and clear multi-level pipeline queue architecture of the secondary cache, and based on the fixed time spent to obtain the hit result and the fixed time spent to obtain the refill data, based on the pipeline queue, it is designed to obtain the hit result of the memory access instruction and send a wake-up request at a fixed data bit, and to obtain the refill data at another fixed data bit and send the refill request through the refill queue; based on the pipeline queue architecture and the design of each processing time of the instruction in the pipeline, it is possible to achieve accurate and stable control of the fixed advance amount of the wake-up request, thereby ensuring the accuracy and coverage of the memory access instruction reading process. The entire process does not require reading the status of the requests at each level of the pipeline, and the real-time calculation of the advance issuance time based on the status of the refill queue request, so the complexity is extremely low, reducing the cost and power consumption of the circuit.
[0055] In summary, in the embodiment of the present application, when it is determined that the memory access instruction of the processor does not hit in the first-level cache, the memory access instruction is controlled to enter the pipeline queue from the starting data bit; and after an interval of a first number of data bits, the hit result of the memory access instruction in the second-level cache is obtained; if the hit result is a hit, a wake-up instruction is immediately generated through the second-level cache and issued by the wake-up queue, and after an interval of a second number of data bits, refill data is obtained through the second-level cache, and the memory access instruction is issued as a refill instruction by the refill queue. The present application implements the management of memory access instructions through the concise and clear multi-level pipeline queue architecture of the second-level cache. Based on the architecture of the pipeline queue and the design of each processing timing of the instruction in the pipeline, it can achieve accurate and stable control of the fixed advance amount of the wake-up request, thereby ensuring the accuracy and coverage of the wake-up instruction. The entire process does not require reading the status of the requests at each level of the pipeline, and calculating the advance issuance time in real time based on the status of the refill queue request, so the complexity is extremely low, reducing the cost and power consumption of the circuit.
[0056] FIG3 is a flowchart of specific steps of a cache data reading method provided by an embodiment of the present application. As shown in FIG3 , the method may include:
[0057] Step 201: When it is determined that the memory access instruction of the processor does not hit in the first-level cache, control the memory access instruction to enter the pipeline queue from the starting data bit of the pipeline queue of the second-level cache; the pipeline queue includes multiple data bits arranged in sequence.
[0058] This step may be specifically referred to the above step 101 and will not be described in detail here.
[0059] Step 202: After a first number of data bits have passed, obtain a hit result of the memory access instruction in the secondary cache.
[0060] This step may be specifically referred to the above step 102 and will not be described in detail here.
[0061] Optionally, the pipeline queue includes five sequentially arranged data bits, each data bit corresponds to a pipeline time, and different data bits correspond to different pipeline times; step 202 may specifically include:
[0062] Sub-step 2021: After an interval of 2 data bits, obtain the hit result of the memory access instruction in the secondary cache.
[0063] In an embodiment of the present application, referring to Figure 1, in an implementation scenario, the pipeline queue may include 5 sequentially arranged data bits: data bit S1-data bit S5. Based on the analysis of the time required to determine the hit result of the memory access instruction in the secondary cache, it is found that the judgment process consumes the time represented by 2 consecutive data bits. Therefore, the processing timing of the instruction designed in this application is: the hit result can be obtained at an interval of 2 data bits from the starting data bit (S1). This indicates that after the memory access instruction enters the pipeline of the secondary cache, the execution process of waiting for the hit determination is completed, thereby obtaining the hit result of the memory access instruction in the secondary cache.
[0064] Step 203: If the hit result is a hit, a wake-up instruction is immediately generated through the secondary cache and issued by the wake-up queue, and after an interval of a second number of data bits, refill data is obtained through the secondary cache, and the memory access instruction is issued as a refill instruction by the refill queue.
[0065] The wake-up instruction is used to wake up the operation of reading data from the first-level cache through the memory access instruction; the refill instruction is used to write the refill data into the first-level cache for reading by the memory access instruction.
[0066] This step may be specifically referred to the above step 103 and will not be described in detail here.
[0067] Optionally, the pipeline queue includes five sequentially arranged data bits, each data bit corresponds to a pipeline moment, and different data bits correspond to different pipeline moments; step 203 may specifically include:
[0068] Sub-step 2031: After an interval of 2 data bits, obtain the refill data through the secondary cache.
[0069] In an embodiment of the present application, referring to Figure 1, in one implementation scenario, the pipeline queue may include 5 sequentially arranged data bits: data bit S1-data bit S5. Based on the analysis of the time required for the secondary cache to obtain refill data, it is found that the acquisition process consumes the time represented by 2 consecutive data bits. Therefore, the processing timing of the instruction designed in this application is: from the moment of acquisition to the moment of the hit result (the moment corresponding to data bit S3), the refill data can be obtained with an interval of 2 data bits, which represents the process of the instruction in the pipeline waiting for the secondary cache to obtain the refill data.
[0070] Optionally, the method may further include:
[0071] Step 204: If the hit result is a miss, the memory access instruction is controlled to leave the pipeline queue, and after waiting for the third-level cache to refill the second-level cache, the memory access instruction is controlled to enter the pipeline queue from the starting data bit, and at the same time, a wake-up instruction is generated through the second-level cache and issued by the wake-up queue; and after the second number of data bits, the refill data is obtained through the second-level cache, and the memory access instruction is issued as a refill instruction by the refill queue.
[0072] In an embodiment of the present application, the hit result is a miss, indicating that the data requested to be read by the memory access instruction is not stored in the second-level cache. At this time, it is necessary to check whether the third-level cache has the data. If the data is stored in the third-level cache, the third-level cache is allowed to refill the data into the second-level cache, and then the second-level cache is allowed to refill the data into the first-level cache for the memory access instruction to read; if the data is not stored in the third-level cache, the data is read from the memory and refilled into the third-level cache, and then the third-level cache is allowed to refill the data into the second-level cache, and finally the second-level cache is allowed to refill the data into the first-level cache for the memory access instruction to read.
[0073] Specifically, when the hit result is a miss, since it is necessary to wait for the L3 cache to refill the data into the L2 cache, the memory access instruction can be controlled to leave the pipeline queue first to wait for the L3 cache to be refilled. Referring to Figure 1, after waiting for the L3 cache to refill the L2 cache, the memory access instruction is controlled to re-enter the pipeline queue from the starting data bit, and at the same time, a wake-up instruction is generated through the L2 cache and issued by the wake-up queue (the time of entering the wake-up queue and issuing the wake-up instruction is the time corresponding to data bit S1); and after an interval of a second number of data bits (an interval of 2 data bits), the refill data is obtained through the L2 cache, and the memory access instruction is issued as a refill instruction from the refill queue.
[0074] For example, referring to Figure 1, when the hit result is a miss, the actual sending time of the wake-up request is the moment corresponding to data bit S1. Since the refill request has to wait in the refill queue for the length of time represented by a data bit (ensuring that the refill request at the exit of the refill queue is issued in time to reduce the chance of congestion in the refill queue), the actual sending time of the refill request is the moment corresponding to data bit S4. It can be seen that the embodiment of Figure 1 of the present application can ensure that for each refill request, a wake-up request is issued three data bits in advance.
[0075] Optionally, after controlling the memory access instruction to leave the pipeline queue, the method may further include:
[0076] Step 205: Control the memory access instruction to enter the missing status register to wait.
[0077] After waiting for the L3 cache to refill the L2 cache, the method may further include:
[0078] Step 206 : Extract the memory access instruction from the missing status register and go to step 201 .
[0079] In steps 205-206, after the memory access instruction leaves the L2 cache pipeline, it enters the Miss-Status Handling Register (MSHR) to wait. The MSHR records each incomplete transaction, including the failed address, keyword information, and incompletely executed instructions. Once the L3 cache has refilled the L2 cache, the memory access instruction in the MSHR can be re-executed. The re-executed memory access instruction can re-enter the L2 cache pipeline queue from the starting data bit.
[0080] Optionally, the process of generating a wake-up instruction through the secondary cache and issuing it through the wake-up queue can be specifically implemented through the following sub-steps:
[0081] Sub-step A1: Generate the wake-up instruction through the secondary cache, and add the wake-up instruction to the wake-up queue.
[0082] Sub-step A2: When there are other wake-up instructions queued before the wake-up instruction, wait until the other wake-up instructions are issued, and then issue the wake-up instruction from the wake-up queue.
[0083] Sub-step A3: When no other wake-up instruction is queued before the wake-up instruction, directly send the wake-up instruction from the wake-up queue.
[0084] In an embodiment of the present application, for sub-steps A1-A3, the enqueuing strategy of the wake-up queue includes: for a memory access instruction that hits the secondary cache, the corresponding wake-up instruction is enqueued after a first number of data bits from the starting data bit (the enqueuing timing is data bit S3 of the pipeline queue in Figure 1); for a memory access instruction that does not hit the secondary cache, the corresponding wake-up instruction is enqueued at the starting data bit (the enqueuing timing is data bit S1 of the pipeline queue in Figure 1).
[0085] As for the dequeueing strategy of the wakeup queue, a dequeueing strategy can be adopted, which is always possible (if the queue is empty, dequeueing can be performed at the current data position). Specifically, if there are other wakeup instructions queued before the wakeup instruction, the wakeup instruction is sent from the wakeup queue after the other wakeup instructions are issued. If there are no other wakeup instructions queued before the wakeup instruction, the wakeup instruction is directly sent from the wakeup queue, thereby improving the efficiency of the wakeup queue in sending wakeup instructions.
[0086] It should be noted that the wake-up request does not need to wait for a data bit after entering the wake-up queue.
[0087] Optionally, the process of issuing the memory access instruction as a refill instruction from the refill queue can be specifically implemented through the following sub-steps:
[0088] Sub-step B1: Waiting for one data bit in the refill queue before issuing the refill instruction.
[0089] In the embodiment of the present application, the refill instruction is sent after waiting for one data bit in the refill queue, the purpose of which is to ensure that the refill request at the exit of the refill queue is sent in time, thereby reducing the probability of congestion in the refill queue.
[0090] For example, referring to Figure 1, when the hit result is a hit, the refill request enters the refill queue at data bit S5. Since the refill request has to wait in the refill queue for the length of time represented by a data bit, the actual timing of issuing the refill request is the moment corresponding to data bit S6 (not shown).
[0091] Optionally, the secondary cache includes a plurality of different ports, each of which has a corresponding pipeline queue, a refill queue, and a wake-up queue; and the process of generating a wake-up instruction through the secondary cache and issuing the wake-up instruction through the wake-up queue can be specifically implemented through the following sub-steps:
[0092] Sub-step C1: generating a wake-up instruction through the secondary cache, and determining a candidate port in the current data bit of the pipeline queue to which the wake-up instruction is to be sent.
[0093] Sub-step C2: selecting a target port that is allowed to send a wake-up instruction from all the candidate ports according to a preset arbitration strategy.
[0094] Sub-step C3: sending a target wake-up instruction corresponding to the target port through the target wake-up queue corresponding to the target port, and blocking the sending of wake-up instructions by other ports.
[0095] In an embodiment of the present application, for sub-steps C1-C3, the secondary cache may include multiple different ports, each port having a corresponding pipeline queue, refill queue and wake-up queue, and the pipeline queues, refill queues and wake-up queues of different ports are different.
[0096] For the case where the second-level cache has multiple ports and the first-level cache has only one port, there will be a moment when multiple ports of the second-level cache will send wake-up requests to the first-level cache separately, but the first-level cache can only receive one wake-up request at a time. This results in port 1 of the second-level cache successfully sending a wake-up request (the other ports need to wait), and then after three data bits, due to port competition, it is not necessarily port 1 that sends a refill request to the first-level cache (it may be other ports that send a refill request). This causes the wake-up request and refill request sent by the port to be out of sync, affecting the stability of the memory access request reading process.
[0097] To address this issue, the present invention can, when generating a wake-up instruction through the L2 cache, determine all candidate ports in the pipeline queue's current data bit (at the current moment) to which the wake-up instruction is to be sent. Then, according to a preset arbitration strategy, it can select a target port from among all candidate ports that is allowed to send the wake-up instruction. After selecting the target port, it sends the target wake-up instruction corresponding to the target port, blocking the remaining ports from sending the wake-up instruction.
[0098] It should be noted that the arbitration strategy can be implemented in various ways. For example, one implementation can randomly select one of all candidate ports as the target port. Another implementation can select the target port in a round-robin manner. For example, at the current moment, port 1 is used as the target port to send the wake-up request. At the next moment, if port 2 is among the candidate ports, port 2 will be used as the target port to send the wake-up request (if port 3 is not among the candidate ports but port 2 is, port 3 will be used as the target port to send the wake-up request, and so on). At the next moment, if port 3 is among the candidate ports, port 3 will be used as the target port to send the wake-up request (if port 4 is not among the candidate ports but port 3 is, port 4 will be used as the target port to send the wake-up request, and so on).
[0099] Based on sub-steps C1-C3, after the second number of data bits have passed, refill data is obtained from the secondary cache, and the memory access instruction is issued from the refill queue as a refill instruction. This process can be specifically implemented by the following sub-steps:
[0100] Sub-step D1: After the second number of data bits, the selected target port continues to be the port for sending refill instructions, and the target refill instructions corresponding to the target port are sent through the refill queue corresponding to the target port, and the sending of refill instructions by the remaining ports is blocked.
[0101] Furthermore, since the target port is selected through the arbitration policy to send the wake-up instruction, after the second number of data bits, the target port continues to send the target refill instruction, thereby blocking the other ports from sending the refill instruction. This ensures that each port in the L2 cache accurately sends the wake-up instruction and the refill instruction, and ensures that the lead time of the sent wake-up instruction is stable and accurate. Each port's sending of the wake-up instruction and the refill instruction is an independent process and is not affected by the other ports.
[0102] It should be noted that blocking of ports other than the target port can be achieved by setting a mask, in which the valid bits corresponding to the ports other than the target port are overwritten with 0.
[0103] In summary, in the related art, the secondary cache reads the request status in its own pipeline queue and refill request queue, calculates the time to send the wake-up request in advance based on the request status, and sends the wake-up request in advance. This method requires analysis 2 5×2×2×2×16=524,288 or more states, which has extremely high complexity, so only simplified solutions can be adopted during implementation, which reduces coverage and accuracy. However, the embodiment of the present application only needs to calculate the advance issuance time in real time based on the state of the request itself, and the entire process does not need to read the state of the pipeline request, so the complexity is extremely low, reducing the cost and power consumption of the circuit. For example, in the example of Figure 1, the solution of the embodiment of the present application can reduce the delay of 3 data bits for the reading of all memory access requests that do not hit in the first-level cache. In particular, for requests that do not hit in the first-level cache but hit in the second-level cache, the delay is only 10 data bits, so reducing the delay by 3 data bits is a very considerable benefit.
[0104] In summary, in an embodiment of the present application, when it is determined that the memory access instruction of the processor does not hit in the first-level cache, the memory access instruction is controlled to enter the pipeline queue from the starting data bit; and after an interval of a first number of data bits, the hit result of the memory access instruction in the second-level cache is obtained; if the hit result is a hit, a wake-up instruction is immediately generated through the second-level cache and issued by the wake-up queue, and after an interval of a second number of data bits, refill data is obtained through the second-level cache, and the memory access instruction is issued as a refill instruction by the refill queue. The present application implements the management of memory access instructions through the concise and clear multi-level pipeline queue architecture of the second-level cache. Based on the architecture of the pipeline queue and the design of each processing timing of the instruction in the pipeline, it can achieve accurate and stable control of the fixed advance amount of the wake-up request, thereby ensuring the accuracy and coverage of the wake-up instruction. The entire process does not require reading the status of the requests at each level of the pipeline, and calculating the advance issuance time in real time based on the status of the refill queue request, so the complexity is extremely low, reducing the cost and power consumption of the circuit.
[0105] FIG4 is a block diagram of a cache data reading device provided in an embodiment of the present application, the device comprising:
[0106] The enqueue module 301 is configured to, when it is determined that a memory access instruction of the processor does not hit in the first-level cache, control the memory access instruction to enter the pipeline queue from the starting data bit of the pipeline queue of the second-level cache; the pipeline queue includes a plurality of data bits arranged in sequence;
[0107] A determination module 302 is configured to obtain a hit result of the memory access instruction in the secondary cache after a first number of data bits have passed;
[0108] a dequeue module 303 configured to, if the hit result is a hit, immediately generate a wake-up instruction through the secondary cache and issue it from the wake-up queue, and after a second number of data bits, obtain refill data through the secondary cache and issue the memory access instruction as a refill instruction from the refill queue;
[0109] The wake-up instruction is used to wake up the operation of reading data from the first-level cache through the memory access instruction; the refill instruction is used to write the refill data into the first-level cache for reading by the memory access instruction.
[0110] Optionally, the device further includes:
[0111] a separation module, configured to control the memory access instruction to separate from the pipeline queue if the hit result is a miss;
[0112] a re-entry module, configured to control the memory access instruction to enter the pipeline queue from the starting data bit after waiting for the third-level cache to refill the second-level cache, and at the same time generate a wake-up instruction through the second-level cache and issue it from the wake-up queue;
[0113] The refill module is configured to obtain refill data through the secondary cache after an interval of the second number of data bits, and issue the memory access instruction as a refill instruction from the refill queue.
[0114] Optionally, the device further includes:
[0115] A waiting module, used for controlling the memory access instruction to enter a missing status register for waiting;
[0116] An extraction module is used to extract the memory access instruction from the missing status register, and enter the step of controlling the memory access instruction to enter the pipeline queue from the starting data bit, and at the same time generate a wake-up instruction and issue it from the wake-up queue.
[0117] Optionally, the dequeue module 303 includes:
[0118] An adding submodule, configured to generate the wake-up instruction through the secondary cache and add the wake-up instruction to the wake-up queue;
[0119] A first judgment submodule is configured to, if other wake-up instructions are queued before the wake-up instruction, wait for the other wake-up instructions to be issued and then issue the wake-up instruction from the wake-up queue;
[0120] The second judgment submodule is configured to directly send the wake-up instruction from the wake-up queue if no other wake-up instruction is queued before the wake-up instruction.
[0121] Optionally, the dequeue module 303 includes:
[0122] The sending submodule is used to wait for one data bit in the refill queue before sending the refill instruction.
[0123] Optionally, the pipeline queue includes 5 sequentially arranged data bits, each data bit corresponds to a pipeline moment, and different data bits correspond to different pipeline moments;
[0124] The judgment module 302 includes:
[0125] A first interval submodule, configured to obtain a hit result of the memory access instruction in the secondary cache after an interval of 2 data bits;
[0126] The dequeue module 303 includes:
[0127] The second interval submodule is configured to obtain the refill data through the secondary cache after an interval of 2 data bits.
[0128] Optionally, the secondary cache includes a plurality of different ports, each of the ports having a corresponding pipeline queue, a refill queue, and a wake-up queue;
[0129] The dequeue module 303 includes:
[0130] A candidate submodule, configured to generate a wake-up instruction through the secondary cache and determine a candidate port in the current data bit of the pipeline queue to which the wake-up instruction is to be sent;
[0131] An arbitration submodule, configured to select a target port that is allowed to send a wake-up instruction from all the candidate ports according to a preset arbitration strategy;
[0132] The first blocking submodule is configured to send a target wake-up instruction corresponding to the target port through a target wake-up queue corresponding to the target port, and block the sending of the wake-up instruction by other ports.
[0133] Optionally, the dequeue module 303 includes:
[0134] The second blocking submodule is used to continue to use the selected target port as the port for sending refill instructions after an interval of a second number of data bits, send the target refill instructions corresponding to the target port through the refill queue corresponding to the target port, and block the sending of refill instructions by the remaining ports.
[0135] In summary, in the embodiment of the present application, when it is determined that the memory access instruction of the processor does not hit in the first-level cache, the memory access instruction is controlled to enter the pipeline queue from the starting data bit; and after an interval of a first number of data bits, the hit result of the memory access instruction in the second-level cache is obtained; if the hit result is a hit, a wake-up instruction is immediately generated through the second-level cache and issued by the wake-up queue, and after an interval of a second number of data bits, refill data is obtained through the second-level cache, and the memory access instruction is issued as a refill instruction by the refill queue. The present application implements the management of memory access instructions through the concise and clear multi-level pipeline queue architecture of the second-level cache. Based on the architecture of the pipeline queue and the design of each processing timing of the instruction in the pipeline, it can achieve accurate and stable control of the fixed advance amount of the wake-up request, thereby ensuring the accuracy and coverage of the wake-up instruction. The entire process does not require reading the status of the requests at each level of the pipeline, and calculating the advance issuance time in real time based on the status of the refill queue request, so the complexity is extremely low, reducing the cost and power consumption of the circuit.
[0136] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0137] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0138] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0139] An embodiment of the present application provides a cache data reading device, comprising a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors to include methods for performing one or more of the above embodiments.
[0140] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0141] The various component embodiments of the present application can be implemented in hardware, or in a software module running on one or more processors, or in a combination thereof. It will be appreciated by those skilled in the art that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the computing processing equipment according to the embodiment of the present application. The application can also be implemented as a device or apparatus program (for example, a computer program and a computer program product) for performing a part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0142] For example, FIG5 illustrates a computing device that can implement the methods according to the present application. The computing device typically includes a processor 1010 and a computer program product or computer-readable medium in the form of a memory 1020. Memory 1020 can be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Memory 1020 has storage space 1030 for program code 1031 for executing any of the method steps described above. For example, storage space 1030 for program code can include individual program codes 1031 for implementing various steps in the method described above. These program codes can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. Such computer program products are typically portable or fixed storage units, as described with reference to FIG6 . This storage unit can have storage segments, storage space, and the like arranged similarly to memory 1020 in the computing device of FIG5 . The program code can, for example, be compressed in a suitable form. Typically, the storage unit includes computer-readable codes 1031 ′, ie, codes that can be read by a processor such as 1010 , which, when executed by a computing device, cause the computing device to perform the steps of the method described above.
[0143] 7 is a block diagram of an electronic device 600 according to an exemplary embodiment. For example, the electronic device 600 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0144] 7 , the electronic device 600 may include one or more of the following components: a processing component 602 , a memory 604 , a power component 606 , a multimedia component 608 , an audio component 610 , an input / output (I / O) interface 612 , a sensor component 614 , and a communication component 616 .
[0145] The processing component 602 generally controls the overall operation of the electronic device 600, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 602 may include one or more modules to facilitate interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate interaction between the multimedia component 608 and the processing component 602.
[0146] The memory 604 is used to store various types of data to support operations on the electronic device 600. Examples of such data include instructions for any application or method operating on the electronic device 600, contact data, phone book data, messages, pictures, multimedia, etc. The memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0147] The power supply assembly 606 provides power to the various components of the electronic device 600. The power supply assembly 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 600.
[0148] The multimedia component 608 includes a screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the electronic device 600 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0149] The audio component 610 is used to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC), which is used to receive external audio signals when the electronic device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 also includes a speaker for outputting audio signals.
[0150] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0151] The sensor assembly 614 includes one or more sensors for providing various aspects of status assessment for the electronic device 600. For example, the sensor assembly 614 can detect the open / closed state of the electronic device 600, the relative positioning of components, such as the display and keypad of the electronic device 600. The sensor assembly 614 can also detect changes in the position of the electronic device 600 or a component of the electronic device 600, the presence or absence of user contact with the electronic device 600, the orientation or acceleration / deceleration of the electronic device 600, and temperature changes of the electronic device 600. The sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0152] The communication component 616 is used to facilitate wired or wireless communication between the electronic device 600 and other devices. The electronic device 600 can access a wireless network based on a communication standard, such as WiFi, an operator network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0153] In an exemplary embodiment, the electronic device 600 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the methods provided in the embodiments of the present application.
[0154] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by a processor 620 of an electronic device 600 to perform the above method. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0155] FIG8 is a block diagram of an electronic device 700 according to an exemplary embodiment. For example, the electronic device 700 can be provided as a server. Referring to FIG8 , the electronic device 700 includes a processing component 722, which further includes one or more processors, and a memory resource represented by a memory 732 for storing instructions that can be executed by the processing component 722, such as an application. The application stored in the memory 732 can include one or more modules, each of which corresponds to a set of instructions. In addition, the processing component 722 is configured to execute instructions to perform the method provided in the embodiment of the present application.
[0156] The electronic device 700 may further include a power supply component 726 configured to perform power management of the electronic device 700, a wired or wireless network interface 750 configured to connect the electronic device 700 to a network, and an input / output (I / O) interface 758. The electronic device 700 may operate based on an operating system stored in the memory 732, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0157] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in the above embodiment when executed by a processor.
[0158] References herein to "one embodiment," "an embodiment," or "one or more embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Furthermore, please note that instances of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.
[0159] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0160] In the claims, any reference signs placed between brackets shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0161] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
[0162] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0163] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A method for reading cache data, the method comprising: When it is determined that the memory access instruction of the processor does not hit in the first-level cache, control the memory access instruction to enter the pipeline queue from the starting data bit of the pipeline queue of the second-level cache; the pipeline queue includes a plurality of data bits arranged in sequence; After a first number of data bits have been separated, obtaining a hit result of the memory access instruction in the secondary cache; If the hit result is a hit, a wake-up instruction is immediately generated through the secondary cache and issued by the wake-up queue, and after an interval of a second number of data bits, refill data is obtained through the secondary cache, and the memory access instruction is issued as a refill instruction by the refill queue; The wake-up instruction is used to wake up the operation of reading data from the first-level cache through the memory access instruction; and the refill instruction is used to write the refill data into the first-level cache for reading by the memory access instruction.
2. The cache data reading method according to claim 1, wherein: The method further comprises: If the hit result is a miss, controlling the memory access instruction to leave the pipeline queue; After waiting for the L3 cache to refill the L2 cache, controlling the memory access instruction to enter the pipeline queue from the starting data bit, and generating a wake-up instruction through the L2 cache and issuing it from the wake-up queue; After the second number of data bits are separated, refill data is obtained through the secondary cache, and the memory access instruction is issued from the refill queue as a refill instruction.
3. The cache data reading method according to claim 2, wherein: After controlling the memory access instruction to leave the pipeline queue, the method further includes: Controlling the memory access instruction to enter the missing status register to wait; After waiting for the L3 cache to refill the L2 cache, the method further includes: The memory access instruction is extracted from the missing status register, and the step of controlling the memory access instruction to enter the pipeline queue from the starting data bit is entered, and at the same time, a wake-up instruction is generated through the secondary cache and issued by the wake-up queue.
4. The cache data reading method according to claim 1 or 2, wherein: The generating a wake-up instruction through the secondary cache and issuing it through the wake-up queue includes: Generate the wake-up instruction through the secondary cache, and add the wake-up instruction to the wake-up queue; In the case where other wake-up instructions are queued before the wake-up instruction, after waiting for the other wake-up instructions to be issued, the wake-up instruction is issued from the wake-up queue; In the case that no other wake-up instructions are queued before the wake-up instruction, the wake-up instruction is directly issued from the wake-up queue.
5. The cache data reading method according to claim 1 or 2, wherein: The step of issuing the memory access instruction as a refill instruction from a refill queue includes: The refill instruction is sent out after waiting for one data bit in the refill queue.
6. The cache data reading method according to claim 1, wherein: The pipeline queue includes 5 sequentially arranged data bits, each of the data bits corresponds to a pipeline time, and different data bits correspond to different pipeline times; The step of obtaining a hit result of the memory access instruction in the secondary cache after the first number of data bits are separated includes: After an interval of 2 data bits, obtaining a hit result of the memory access instruction in the secondary cache; The step of obtaining refill data through the secondary cache after the interval of the second number of data bits comprises: After an interval of 2 data bits, the refill data is obtained through the secondary cache.
7. The cache data reading method according to claim 1 or 2, wherein: The secondary cache includes a plurality of different ports, each of which has a corresponding pipeline queue, a refill queue and a wake-up queue; The generating a wake-up instruction through the secondary cache and issuing it through the wake-up queue includes: Generate a wake-up instruction through the secondary cache, and determine a candidate port of the pipeline queue to which the wake-up instruction is to be sent; According to a preset arbitration strategy, a target port that is allowed to send a wake-up instruction is selected from all the candidate ports; The target wake-up instruction corresponding to the target port is sent through the target wake-up queue corresponding to the target port, and the sending of the wake-up instruction by other ports is blocked.
8. The cache data reading method according to claim 7, wherein: The method of acquiring refill data through the secondary cache after the interval of the second number of data bits, and issuing the memory access instruction as a refill instruction from the refill queue comprises: After an interval of the second number of data bits, the selected target port continues to be the port for sending refill instructions, and the target refill instructions corresponding to the target port are sent through the refill queue corresponding to the target port, and the sending of refill instructions by other ports is blocked.
9. The cache data reading method according to any one of claims 1 to 8, wherein: The hit result is a hit, indicating that the secondary cache has the data requested to be read by the memory access instruction; The hit result being a miss indicates that the secondary cache does not store the data requested to be read by the memory access instruction.
10. The cache data reading method according to claim 2, wherein: The first number of data bits is determined based on a time required for executing an operation of determining a hit result of the memory access instruction in the secondary cache; The second number of data bits is determined based on the length of time it takes for the secondary cache to read the refill data.
11. The cache data reading method according to claim 2, wherein: The method further comprises: When the third-level cache does not store the data requested by the memory access instruction, the data requested by the memory access instruction is read from the memory and refilled into the third-level cache, so that the third-level cache can refill the data requested by the memory access instruction into the second-level cache.
12. The cache data reading method according to claim 7, wherein: The step of selecting a target port that is allowed to send a wake-up instruction from all the candidate ports according to a preset arbitration strategy includes: One of the candidate ports is randomly selected as the target port, or the target port is selected from all the candidate ports in a round-robin manner.
13. A cache data reading device, the device comprising: An enqueue module, for controlling the memory access instruction of the processor to enter the pipeline queue from the starting data bit of the pipeline queue of the secondary cache when it is determined that the memory access instruction of the processor does not hit in the primary cache; the pipeline queue includes a plurality of data bits arranged in sequence; A judgment module, configured to obtain a hit result of the memory access instruction in the secondary cache after a first number of data bits have been separated; a dequeue module, configured to, if the hit result is a hit, immediately generate a wake-up instruction through the secondary cache and issue it from the wake-up queue, and obtain refill data through the secondary cache after an interval of a second number of data bits, and issue the memory access instruction as a refill instruction from the refill queue; The wake-up instruction is used to wake up the operation of reading data from the first-level cache through the memory access instruction; and the refill instruction is used to write the refill data into the first-level cache for reading by the memory access instruction.
14. The cache data reading device according to claim 13, wherein: The device also includes: A separation module, used for controlling the memory access instruction to separate from the pipeline queue if the hit result is a miss; A re-entry module, used for controlling the memory access instruction to enter the pipeline queue from the starting data bit after waiting for the third-level cache to refill the second-level cache, and generating a wake-up instruction through the second-level cache and issuing it from the wake-up queue; The refill module is used to obtain refill data through the secondary cache after the second number of data bits are separated, and to issue the memory access instruction as a refill instruction from the refill queue.
15. The cache data reading device according to claim 14, wherein: The device also includes: A waiting module, used for controlling the memory access instruction to enter a missing status register for waiting; The extraction module is used to extract the memory access instruction from the missing status register, and enter the step of controlling the memory access instruction to enter the pipeline queue from the starting data bit, and at the same time generate a wake-up instruction through the secondary cache and issue it from the wake-up queue.
16. The cache data reading device according to claim 13 or 14, wherein: The dequeue module comprises: An adding submodule, configured to generate the wake-up instruction through the secondary cache and add the wake-up instruction to the wake-up queue; A first judgment submodule is used for, when there are other wake-up instructions queued before the wake-up instruction, to wait for the other wake-up instructions to be issued and then issue the wake-up instruction from the wake-up queue; The second judgment submodule is used for directly sending the wake-up instruction from the wake-up queue when no other wake-up instruction is queued before the wake-up instruction.
17. The cache data reading device according to claim 13 or 14, wherein: The dequeue module comprises: The sending submodule is used to send out the refill instruction after waiting for one data bit in the refill queue.
18. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 12.
19. A computer program comprising computer readable code which, when executed on a computing processing device, causes the computing processing device to execute the method according to any one of claims 1 to 12.
20. A computer-readable storage medium storing the computer program according to claim 19, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and storage medium
CN115543938A
Cache access method and device, storage medium and electronic equipment
CN116909943A
Cache data reading method and device, equipment and storage medium
CN117453435A
Apparatus and method for distributed non-blocking multi-level cache
US6430654B1
Cited By
Data processing device and data processing method
CN120407440A
Graphics processor, instruction execution method, terminal equipment and medium
CN121437247A