Cache processing method and device, chip, equipment and medium
By introducing a high concurrency architecture into the cache processing system, and using parallel processing of data arrays and tag arrays, the problem of inefficiency of multiple data processing instructions in the cache system is solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202410088946.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-07-22
AI Technical Summary
As the processor transmission width increases, higher requirements are put forward for the design of cache, and the prior art is difficult to effectively improve the parallel processing efficiency of multiple data processing instructions.
A cache processing system with a high concurrency architecture is designed, including a data array, a P cache pipeline and a P tag array. Each cache pipeline is connected to a tag array, and data processing instructions are executed in parallel to improve data processing efficiency.
By processing multiple data processing instructions in parallel, the data processing efficiency of the cache system is significantly improved and the performance of the processor is improved.
Smart Images

Figure CN120353501A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular to the field of cache technologies. Specifically, it relates to a cache processing method, a cache processing device, a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Processors, caches, and memories are important components in computer devices. The processor is mainly responsible for executing computer instructions. The memory is generally a Dynamic Random Access Memory (DRAM), which is mainly responsible for temporarily storing data. The cache is generally a Static Random-Access Memory (SRAM). Its basic principle is to utilize the spatial locality of programs (that is, there is a high probability that a program will access data adjacent to the recently accessed data) and the temporal locality of programs (that is, there is a high probability that a program will access the recently accessed data). After reading the data currently accessed by the processor from the memory, it is written into the cache, and the adjacent data is also read from the memory in advance and written into the cache. When the processor accesses the data and its adjacent data again in the future, it can be quickly read directly from the cache. With the continuous increase in the issue width of the processor (the number of instructions that can be issued simultaneously in each clock cycle), higher requirements are imposed on the design of the cache. Summary of the Invention
[0003] Embodiments of this application provide a cache processing method, device, chip, device, and medium, which can utilize the designed cache processing system with a high-concurrency architecture to implement parallel processing of multiple data processing instructions and improve data processing efficiency.
[0004] On the one hand, embodiments of this application provide a cache processing method. This cache processing method is applied to a cache processing system. The cache processing system includes a data array, P cache pipelines, and P tag arrays, where P is a positive integer; one cache pipeline is connected to one tag array, and all P cache pipelines are connected to the data array; the data array is used to store data in the form of cache lines; the tag array is used to store cache tags of cache lines; this cache processing method includes:
[0005] Obtain a first data processing instruction in the first cache pipeline; the first data processing instruction carries a first access address, and the first access address is used to locate the first data indicated by the first data processing instruction;
[0006] Based on the first tag array connected to the first cache pipeline, determine a first cache line in the data array for storing the first data;
[0007] Perform a data processing operation on the first data in the first cache line according to the first data processing instruction;
[0008] Wherein, the first cache pipeline is any one of the P cache pipelines; the data processing instructions in each cache pipeline are executed in parallel based on the respective connected tag arrays.
[0009] Correspondingly, an embodiment of the present application provides a cache processing device, in which a cache processing system is provided. The cache processing system includes a data array, P cache pipelines and P tag arrays, where P is a positive integer; one cache pipeline is connected to one tag array, and all P cache pipelines are connected to the data array; the data array is used to store data in the form of cache lines; the tag array is used to store the cache tags of the cache lines; the cache processing device includes:
[0010] An acquisition module, configured to acquire the first data processing instruction in the first cache pipeline; the first data processing instruction carries a first access address, and the first access address is used to locate the first data indicated by the first data processing instruction;
[0011] A processing module, configured to determine, based on the first tag array connected to the first cache pipeline, the first cache line for storing the first data from the data array;
[0012] A processing module, configured to perform a data processing operation on the first data in the first cache line according to the first data processing instruction;
[0013] Wherein, the first cache pipeline is any one of the P cache pipelines; the data processing instructions in each cache pipeline are executed in parallel based on the respective connected tag arrays.
[0014] In one embodiment, the data array includes M-way data memories, and each data memory includes N cache lines; the data array is divided into N cache groups, each cache group corresponds to a cache index, and each cache group includes M cache lines, and the M cache lines belong to different data memories respectively; both M and N are positive integers; the first access address includes a first index field and a first tag field;
[0015] The processing module is further configured to call the tag reading unit in the first cache pipeline to determine the first cache group whose cache index matches the first index field from the N cache groups, and read the cache tags of each cache line in the first cache group from the first tag array to obtain M cache tags;
[0016] The processing module is further configured to call a tag comparison unit in the first cache pipeline to match the M cache tags read with the first tag domain to obtain a tag matching result. If the tag matching result indicates that there is a cache tag of a target cache line among the M cache tags read that matches the first tag domain, the target cache line is determined as the first cache line for storing the first data.
[0017] In one embodiment, the first data processing instruction is a data load instruction. Among the P cache pipelines, there are Q1 data load access pipelines, where Q1 is a positive integer and Q1 is less than or equal to P. The cache processing system further includes Q1 load issue units. One load issue unit is connected to one data load access pipeline. Each load issue unit is configured to send a data load instruction to the respective connected data load access pipeline. The first cache pipeline is any one of the Q1 data load access pipelines, and the first cache pipeline is connected to the first load issue unit among the Q1 load issue units. The first access address further includes a first offset field.
[0018] The processing module is further configured to call a way prediction unit in the first cache pipeline to predict a first data memory storing the first data from the M-way data memory.
[0019] The processing module is further configured to call a first data reading unit in the first cache pipeline to read a second cache line from the first data memory, and the second cache line belongs to the first cache group.
[0020] The processing module is further configured to call a first data selection unit in the first cache pipeline to obtain the cache tag of the first cache line determined based on the tag matching result. If the cache tag of the first cache line is the same as the cache tag of the second cache line, the first data is read from the second cache line according to the first offset field, and the first data is sent to the first load issue unit for data load processing.
[0021] In one embodiment, the first data processing instruction is a data store instruction, and the data store instruction further carries second data to be stored. Among the P cache pipelines, there are Q2 data store access pipelines, where Q2 is a positive integer and Q2 is less than or equal to P. The first cache pipeline is any one of the Q2 data store access pipelines. The first access address further includes a first offset field.
[0022] The processing module is further configured to call a second data reading unit in the first cache pipeline to read M cache lines in the first cache group from the M-way data memory respectively, and send the M cache lines read to a second data selection unit in the first cache pipeline.
[0023] The processing module is further configured to call the second data selection unit in the first cache pipeline to obtain the cache tag of the first cache line determined based on the tag matching result, determine the first cache line from the M cache lines read according to the cache tag of the first cache line, and perform a merging process on the determined first cache line and the second data according to the first offset field, so as to replace the first data in the first cache line with the second data to obtain merged data;
[0024] The processing module is further configured to call the data write unit in the first cache pipeline to write the merged data into the first cache line in the data array.
[0025] In one embodiment, the first data processing instruction is a data detection instruction, and the third data to be detected is also carried in the data detection instruction; among the P cache pipelines, there are Q2 data storage access pipelines, where Q2 is a positive integer and Q2 is less than or equal to P; the first cache pipeline is any one of the Q2 data storage access pipelines; the first access address further includes a first offset field; the cache processing system further includes a detection emission unit for sending the data detection instruction;
[0026] The processing module is further configured to call the second data reading unit in the first cache pipeline to read M cache lines in the first cache group from the M-way data memory respectively;
[0027] The processing module is further configured to call the second data selection unit in the first cache pipeline to obtain the cache tag of the first cache line determined based on the tag matching result, determine the first cache line from the M cache lines read according to the cache tag of the first cache line, and read the first data from the first cache line according to the first offset field;
[0028] The processing module is further configured to call the data comparison unit in the first cache pipeline to perform a consistency detection on the first data and the third data to obtain a detection result, and return the detection result to the detection emission unit that sends the first data processing instruction.
[0029] In one embodiment, the processing module is further configured to, if the tag matching result indicates that there is no cache line in the first cache group whose cache tag matches the first tag field, call the replacement selection unit in the first cache pipeline to determine the cache line to be replaced from the first cache group;
[0030] The processing module is further configured to call a preset register in the cache processing system to generate a data replacement instruction for the cache line to be replaced, and send the data replacement instruction to the second cache pipeline, so as to schedule the second cache pipeline to write the data in the cache line to be replaced into the cache data source, obtain a fourth data including the first data from the cache data source, backfill the fourth data into the cache line to be replaced in the data array, and determine the cache line to be replaced after backfill processing as the first cache line for storing the first data. The second cache pipeline is any one of the P data storage access pipelines.
[0031] In one embodiment, the processing module is further configured to call the second cache pipeline to obtain a data replacement instruction in the second cache pipeline, and based on the second tag array connected to the second cache pipeline, read the cache line to be replaced from the data array, write the data in the cache line to be replaced into the write queue, and write the data in the cache line to be replaced in the write queue into the cache data source.
[0032] In one embodiment, the processing module is further configured to call the first cache pipeline to generate a replacement cache tag based on the tag field in the first access address, and generate a tag update signal for the cache line to be replaced, where the tag update signal includes the replacement cache tag;
[0033] The processing module is further configured to call an update signal replication unit in the cache processing system to update the cache tag of the cache line to be replaced in each tag array to the replacement cache tag based on the tag update signal.
[0034] In one embodiment, the first cache pipeline is a data storage access pipeline; the cache processing system further includes a store buffer, a detection and emission unit, a preset register, and a preset scheduler; the store buffer is configured to send a data storage instruction to the preset scheduler, the detection and emission unit is configured to send a data detection instruction to the preset scheduler, and the preset register is configured to send a data replacement instruction to the preset scheduler;
[0035] The processing module is further configured to call a preset scheduler in the first cache pipeline to determine the sending order of S data processing instructions obtained by the preset scheduler according to the priority. The S data processing instructions include one or more of a data storage instruction, a data replacement instruction, and a data detection instruction. The data storage instruction, the data replacement instruction, and the data detection instruction each have their corresponding priorities; S is a positive integer. Select a first data processing instruction from the S data processing instructions according to the sending order, and send the first data processing instruction to the first cache pipeline.
[0036] In one embodiment, if the storage buffer obtains multiple data storage requests and the tag fields and index fields in the access addresses carried by the multiple data storage instructions are the same, the processing module is further configured to call the storage buffer in the first cache pipeline to merge the multiple data storage instructions into a new data storage instruction and send the new data storage instruction to a preset scheduler.
[0037] In one embodiment, the data array includes M data memories, each data memory includes one or more data storage banks, and each data storage bank is used to store one or more cache lines;
[0038] If the first data storage bank includes a first cache line and no other cache pipeline accesses the first data storage bank, the processing module is further configured to call the data reading unit in the first cache pipeline to read first data from the first cache line in the first data storage bank; wherein, the first data storage bank is any data storage bank in the data array, and the other cache pipelines are the cache pipelines other than the first cache pipeline among the P cache pipelines.
[0039] Correspondingly, an embodiment of the present application provides a chip, which includes: a processor and a cache processing device, the processor is configured to send a data processing instruction to the cache processing device, and the cache processing device is configured to execute the above cache processing method.
[0040] Correspondingly, an embodiment of the present application provides a computer device, which includes:
[0041] A chip, adapted to implement a computer program;
[0042] A computer-readable storage medium, the computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the chip to execute the above cache processing method.
[0043] Correspondingly, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is read and executed by the chip of the computer device, the computer device is caused to execute the above cache processing method.
[0044] Correspondingly, an embodiment of the present application provides a computer program product, which includes a computer program, and the computer program is stored in a computer-readable storage medium. The chip of the computer device reads the computer program from the computer-readable storage medium, and the chip executes the computer program, so that the computer device executes the above cache processing method.
[0045] In the embodiments of the present application, a cache processing system with a high-concurrency architecture is designed by using a data array (for storing data in the form of cache lines), P cache pipelines, and P tag arrays (for storing cache tags of cache lines, and the cache tags stored in the P tag arrays are the same). In this cache processing system, all P cache pipelines are connected to the data array, and one cache pipeline is connected to one tag array, so that the data processing instructions in each cache pipeline can be executed in parallel based on the respective connected tag arrays. Specifically, the first cache pipeline is any one of the P cache pipelines. Executing the first data processing instruction in the first cache pipeline based on the first tag array connected to the first cache pipeline includes: obtaining the first data processing instruction in the first cache pipeline, where the first data processing instruction includes a first access address for locating the first data, determining, based on the first tag array connected to the first cache pipeline, a first cache line in the data array for storing the first data, and performing a data processing operation on the first data in the first cache line according to the first data processing instruction. It can be seen that in the embodiments of the present application, P tag arrays are instantiated based on the P cache pipelines. In this way, the data processing instructions in the P cache pipelines can be processed in parallel based on the respective connected tag arrays, greatly improving the data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 Structural schematic of a cache processing system provided by an embodiment of the present application Figure 1 ;
[0048] Figure 2 Structural schematic diagram of a data array provided by an embodiment of the present application;
[0049] Figure 3 Schematic diagram of data mapping based on a memory address provided by an embodiment of the present application;
[0050] Figure 4 Sub-structural schematic of a cache processing system provided by an embodiment of the present application Figure 1 ;
[0051] Figure 5 Sub-structural schematic of a cache processing system provided by an embodiment of the present application Figure 2 ;
[0052] Figure 6 A substructure diagram of a cache processing system provided in an embodiment of the present application Figure 3 ;
[0053] Figure 7 A substructure diagram of a cache processing system provided in an embodiment of the present application Figure 4 ;
[0054] Figure 8 A schematic diagram of the structure of a data array provided in an embodiment of the present application;
[0055] Figure 9 It is a flowchart of a cache processing method provided in an embodiment of the present application;
[0056] Figure 10 A schematic diagram of a processing flow of a data loading instruction provided in an embodiment of the present application;
[0057] Figure 11 A schematic diagram of a processing flow of a data processing instruction provided in an embodiment of the present application;
[0058] Figure 12 It is a structural diagram of a cache processing device provided in an embodiment of the present application;
[0059] Figure 13 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0061] See also Figure 1 , Figure 1 A schematic diagram of a cache processing system provided in an embodiment of the present application Figure 1 ;like Figure 1 As shown, the cache processing system 10 includes a data array 11, P cache pipelines 12 (such as Figure 1 Cache pipeline 1, cache pipeline 2, ... cache pipeline P) and P tag arrays 13 (such as Figure 1 wherein a cache pipeline 12 is connected to a tag array 13, and P cache pipelines 12 are all connected to the data array 11; P is a positive integer.
[0062] The data array 11 is used to store data in the form of cache lines. That is, a cache line is the smallest loadable unit in the data array 11. The so-called loadable unit means that the data array 11 needs to store, read, and replace data in units of cache lines. A cache line can be used to store H data with adjacent memory addresses (usually the size of a cache line does not exceed 64 bytes), and H is a positive integer. Among them, the memory address is an address used to uniquely identify and access a specific location in the memory. The memory address is divided into three address ranges: a tag field, an index field, and an offset field. The so-called adjacent memory addresses mean that both the tag field and the index field in the memory addresses are the same. For example, when the processor accesses a certain data in the memory (such as the first data), multiple data with the same tag field and index field as those in the memory address of the first data can be obtained, and the first data and the multiple data (with a size not exceeding 64 bytes) are written into a cache line of the data array 11.
[0063] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a data array provided by an embodiment of the present application; as Figure 2 shown, the data array 11 includes M-way data memories 21 (such as Figure 2 the data memory 1, data memory 2,... data memory M in Figure 2 ), and each data memory 21 includes N cache lines (such as Figure 2 the cache line 1, cache line 2,... cache line N in
[0064] ). At the same time, the data array 11 is divided into N cache groups (such as Figure 2 the cache group 1, cache group 2,... cache group N in
[0064] ), and each cache group includes M cache lines, and the M cache lines belong to different data memories respectively. Among them, both M and N are positive integers. In addition, each cache group corresponds to a cache index. In one implementation, if H data with adjacent memory addresses are obtained from the memory, the index fields included and the same in the memory addresses of the H data are obtained, and the cache group whose cache index matches (such as matching when they are the same, and not matching when they are different) the index field is determined from the N cache groups, and the H data are stored in any cache line included in the cache group. That is to say, the data determined by a memory address will be stored in a fixed cache group and can be stored in any cache line included in the fixed cache group.
[0064] Each of the p tag arrays 13 is used to store cache tags of cache lines, and the cache tags stored in the p tag arrays 13 are the same. Among them, the cache tag is the same tag field included in the memory addresses of each piece of data in the corresponding cache line; for example, when H pieces of data with adjacent memory addresses are read and stored in the corresponding cache line (such as cache line 1), the same tag field included in the memory addresses of the H pieces of data is obtained, and this tag field is determined as the cache tag of cache line 1. Optionally, each tag array 13 may include an M-way tag memory. One-way tag memory corresponds to one-way data memory, and each way of tag memory is used to store cache tags of each cache line in the corresponding data memory; for example, tag memory 1 is used to store cache tags of each cache line in data memory 1, tag memory 2 is used to store cache tags of each cache line in data memory 2,... tag memory M is used to store cache tags of each cache line in data memory M.
[0065] Please refer to Figure 3 , Figure 3 which is a schematic diagram of data mapping based on memory address provided by an embodiment of the present application; as Figure 3 shown, the cache group to which the data indicated by the memory address is stored can be determined through the index field in the memory address (by matching the index field with the cache index); the cache line in the corresponding cache group to which the data indicated by the memory address is stored can be determined according to the tag field in the memory address (by matching the tag field with the cache tag), and the offset of the data indicated by the memory address in the corresponding cache line can be determined according to the offset field in the memory address. In this way, the data indicated by the memory address can be obtained from the corresponding cache line based on the offset.
[0066] In one implementation, each tag array can also be used to store the cache index corresponding to the cache group to which the cache line belongs. In this way, each cache pipeline can implement the matching of the cache index and the index field through the cache index stored in the connected tag array. Each tag array can also be used to store the valid bit of the cache line, and the valid bit is used to indicate whether the corresponding cache line is valid; if there is data stored in the corresponding cache line, then the corresponding cache line is valid, and if there is no data stored in the corresponding cache line, then the corresponding cache line is invalid.
[0067] A cache pipeline is an instruction processing pipeline. An instruction processing pipeline refers to splitting an instruction into multiple execution steps, and different functional circuit units respectively execute these multiple execution steps. P cache pipelines 12 may include Q1 data load access pipelines (load access unit, LAU); Q1 is a positive integer and Q1 is less than or equal to P. A data load access pipeline is an instruction processing pipeline for executing data load instructions. A data load instruction is an instruction that requests to load data, such as an instruction for a processor to request to load data into a register in the processor. The P cache pipelines 12 may further include Q2 data store access pipelines (store access unit, SAU); Q2 is a positive integer and Q2 is less than or equal to P. A data store access pipeline is an instruction processing pipeline for executing data store instructions, data replacement instructions, and data detection instructions. That is, in this application, data replacement instructions and data detection instructions are merged into the data store access pipeline for processing. A data store instruction is an instruction that requests to store data. A data replacement instruction is an instruction that requests to replace data. A data detection instruction is an instruction that requests to perform a consistency check.
[0068] Please refer to Figure 4 , Figure 4 which is a schematic diagram of a sub-structure of a cache processing system provided by an embodiment of this application Figure 1 ; as Figure 4 shown, the cache processing system 10 further includes Q1 load emission units 14 (such as Figure 4 the load emission unit 1, load emission unit 2,... load emission unit Q1 in
[0069] Please refer to Figure 5 , Figure 5 which is a schematic diagram of a sub-structure of a cache processing system provided by an embodiment of this application Figure 2 ; as Figure 5 shown, the data load access pipeline Q1 (any one of the P cache pipelines 12) is connected to the load emission unit Q1 (any one of the Q1 load emission units 14), and the load emission unit Q1 is used to send a data load instruction to the data load access pipeline Q1.
[0070] In addition, by Figure 5It can be known that the data loading access pipeline Q1 includes multiple circuit units: a first tag reading unit 31, a first tag comparison unit 32, a way prediction unit 33, a first data reading unit 34, a first data selection unit 35, and a first replacement selection unit 36. Taking the execution of a data loading instruction in the data loading access pipeline Q1 (assuming that the access address carried by the data loading instruction is the first access address, the first access address is used to locate the first data, and the first access address includes a first index field, a first tag field, and a first offset field) as an example, each circuit unit in the data loading access pipeline Q1 will be introduced below.
[0071] The first tag reading unit 31 is connected to the tag array Q1 (which is the tag array 13 among the P tag arrays 13 connected to the data loading access pipeline Q1). The first tag reading unit 31 is used to determine a first cache group in which the cache index matches the first index field from N cache groups, and send a first read signal for the first cache group to the tag array Q1.
[0072] The tag array Q1 is connected to the first tag comparison unit 32. The tag array Q1 is used to respond to the first read signal and return the cache tags (a total of M cache tags) of each cache line in the first cache group to the first tag comparison unit 32.
[0073] The first tag comparison unit 32 is connected to the first data selection unit 35 ( Figure 5 the connection line is not marked in the figure). The first tag comparison unit 32 is used to perform a matching process on the M cache tags and the first tag field to obtain a tag matching result. If the tag matching result indicates that there is a cache tag of a target cache line among the M cache tags that matches the first tag field, the tag matching result is sent to the first data selection unit 35.
[0074] The first tag comparison unit 32 is also connected to the first replacement selection unit 36. The first tag comparison unit 32 is also used to send the tag matching result to the first replacement selection unit 36 if the tag matching result indicates that there is no cache tag among the M cache tags that matches the first tag field.
[0075] The way prediction unit 33 (Way PredicateUunit, WPU) is connected to the first data reading unit 34. The way prediction unit 33 is used to determine the first data memory to be accessed by the first data reading unit 34. The first data memory refers to the data memory predicted by the way prediction unit 33 from the M-way data memory based on the way prediction mechanism and storing the first data.
[0076] Among them, the path prediction mechanism refers to a strategy that can effectively predict the possible paths to be accessed based on historical information and current instruction information. For example, the executed data load instructions and the data memories accessed by the executed data load instructions are recorded in the path prediction unit 33. For a newly initiated data load instruction, the tag field and index field in the access address carried by the newly initiated data load instruction are compared with the tag field and index field in the access address carried by the executed data load instruction. If both the tag field and the index field are the same, it indicates that the access addresses carried by the newly initiated data load instruction and the executed data load instruction fall within the same cache line, and the data memory accessed by the executed data load instruction is used as the prediction result.
[0077] The first data reading unit 34 is connected to each of the M-way data memories 21 in the data array 11. The first data reading unit 34 is configured to send a second read signal for the second cache line to the first data memory in the data array 11. The second cache line is a cache line in the first data memory included in the first cache group, that is to say, the second cache line belongs to the first cache group.
[0078] Each of the M-way data memories 21 in the data array 11 is connected to the first data selection unit 35. The first data memory in the data array 11 is configured to send the second cache line to the first data selection unit 35 in response to the second read signal.
[0079] The first data selection unit 35 is configured to obtain the cache tag of the target cache line determined based on the tag matching result (i.e., the cache tag that matches the first tag field). If the cache tag of the target cache line is the same as the cache tag of the second cache line, it indicates that the second cache line is the target cache line. According to the first offset field, the first data is read from the second cache line and sent to the load emission unit Q1 for data loading processing. For example, if a data load instruction requests to load data into a register of the processor, the load emission unit Q1 can send the first data to the register of the processor.
[0080] In one implementation, if the cache tag of the target cache line is different from the cache tag of the second cache line, the first data selection unit 35 can respectively read each cache line included in the first cache group from the M-way data memories, select the cache line whose cache tag matches the first tag field from each cache line included in the first cache group, and send the selected cache line to the load emission unit Q1 for data loading processing. In another implementation, if the cache tag of the target cache line is different from the cache tag of the second cache line, the path prediction unit 33 can re-predict the data memory, and the first data selection unit 35 re-executes the data array search in the re-predicted data memory.
[0081] The cache processing system 10 further includes a preset register 15, which can be a miss status handling register (MSHR). When a cache miss occurs (i.e., the data to be accessed is not found in the data array 11), a miss status handling register is used to record the cache miss status, and other instructions can continue to execute. After the cache miss is satisfied (i.e., the missing data is written into the data array 11), this miss status handling register is released. The so-called release means clearing the data in the miss status handling register so that other data can use this miss status handling register.
[0082] The first replacement selection unit 36 is connected to the preset register 15. The first replacement selection unit 36 is configured to determine the number of the cache line to be replaced and send a miss signal including the number of the cache line to be replaced to the preset register 15 if it is determined based on the tag matching result that there is no cache tag in the M cache tags that matches the first tag field (i.e., the cache line of the first data is not included in the data array 11). Wherein, the number of the cache line to be replaced may include a set number and a way number. The way number is used to indicate which way of the data memory the cache line to be replaced belongs to, and the set number is used to indicate which cache set the cache line to be replaced belongs to.
[0083] Wherein, the first replacement selection unit 36 may include a plurality of cache line selection units 361 and a multiplexer 362. One cache line selection unit 361 corresponds to one cache line selection policy. Each cache line selection unit 361 is configured to select a candidate cache line from the data array according to the corresponding cache line selection policy. The multiplexer 362 is configured to select the cache line to be replaced from the multiple candidate cache lines selected by the plurality of cache line selection units 361. The cache line selection policy may include at least one of the following: First In First Out (FIFO) policy, Least Recent Used (LRU) policy, Least Freqency Used (LFU) policy, Random (RAND) policy, Pseudo Least Recent Used (PLRU) policy, and the present application does not limit this.
[0084] The preset register 15 is configured to generate a data replacement instruction for the cache line to be replaced based on the number of the cache line to be replaced in response to the miss signal, and send the data replacement instruction to any one of the P cache pipelines 12 for data storage access pipeline to schedule the data storage access pipeline to write the data in the cache line to be replaced into the cache data source, send a cache read signal including a first access address to the cache data source to obtain a fourth data including a first data from the cache data source, and backfill the fourth data into the cache line to be replaced in the data array 11. In a specific implementation, the data storage access pipeline writes the data in the cache line to be replaced into the write queue 19 (located in the cache processing system 10). When the fourth data is backfilled into the cache line to be replaced in the data array 11, the preset register 15 wakes up the write queue 19 to send a cache write signal to the cache data source. The cache write signal includes the data in the cache line to be replaced, and the cache data source writes the data in the cache line to be replaced into the cache data source in response to the cache write signal. The fourth data may include, in addition to the first data, data adjacent to the memory address of the first data. In this way, the first data selection unit 35 can read the first data from the cache line to be replaced after the backfill process and send the first data to the load emission unit Q1 for data loading processing.
[0085] It should be noted that caches are generally divided into three levels: level 1 cache (L1 cache), level 2 cache (L2 cache), and level 3 cache (L3 cache). The order in which the processor accesses the cache is as follows: L1 cache, L2 cache, L3 cache. That is, if the L1 cache misses the corresponding data (i.e., does not store the corresponding data), then the L2 cache is accessed. If the L2 cache misses the corresponding data, then the L3 cache is accessed. If the L3 cache misses the corresponding data, then the memory is accessed. If the above cache processing system is applied to the L1 cache, the cache data source is the L2 cache; if the above cache processing system is applied to the L2 cache, the cache data source is the L3 cache; if the above cache processing system is applied to the L3 cache, the cache data source is the memory.
[0086] In the field of processors, to complete a certain computing task, first the processor loads data into the registers inside the processor through data loading instructions, then reads the data from the registers to the arithmetic unit of the processor for arithmetic operations, then writes the data obtained from the operations to the registers, and finally writes the data in the registers to the external memory. In this process, data loading is at the top of the entire dependency chain. If the data loading performance is not high, it will restrict the performance of the entire processor. Therefore, this application can improve the data loading efficiency and thus the performance of the processor by processing multiple data loading instructions in parallel through multiple data loading access pipelines. In addition, this application applies a path prediction mechanism in the data loading access pipeline, which can take advantage of parallel search in the tag array and data array, and the data array search is only executed on the predicted data memory, which can effectively improve the speed of concurrent processing.
[0087] Please refer to Figure 6 , Figure 6 which is a schematic diagram of a sub-structure of a cache processing system provided by an embodiment of this application Figure 3 ; as Figure 6 shown, the cache processing system 10 further includes a store buffer 16, a probe unit 17, and a preset scheduler 18. The preset register 15, the store buffer 16, and the probe unit 17 are all connected to the preset scheduler 18, and the preset scheduler 18 is connected to the data storage access pipeline Q2 (any one of the P cache pipelines 12).
[0088] The store buffer 16 is used to send data storage instructions to the preset scheduler 18. In one implementation, if the store buffer 16 obtains multiple data storage requests, and the tag fields and index fields in the access addresses carried by the multiple data storage instructions are the same, it indicates that the multiple data storage instructions access the same cache line. The multiple data storage instructions can be merged into a new data storage instruction (which may include the access addresses carried by the multiple data storage instructions respectively), and the new data storage instruction is sent to the preset scheduler 18.
[0089] The probe unit 17 is used to send data detection instructions to the preset scheduler 18. In one implementation, in order to detect whether the data stored in the cache data source is consistent with the corresponding data in the data array 11, the cache data source can send a data detection instruction to the probe unit 17, and the probe unit 17 forwards the data detection instruction to the preset scheduler 18.
[0090] The preset register 15 is used to send a data replacement instruction to the preset scheduler 18. For example, when any data load access pipeline or any data store access pipeline misses data, a miss signal carrying the number of the cache line to be replaced can be sent to the preset register 15. Based on this miss signal, the preset register 15 generates a data replacement instruction for the cache line to be replaced and sends the data replacement instruction to the preset scheduler 18.
[0091] The preset scheduler 18 is used to select a data processing instruction to be sent to the data store access pipeline Q2 from the S obtained data processing instructions (including one or more of a data store instruction, a data replacement instruction, and a data detection instruction). Specifically, in implementation, the data store instruction may have a first priority, and the data replacement instruction and the data detection instruction may have a second priority. The preset scheduler 18 can be a strict priority scheduler (Sp). The preset scheduler 18 will first select a data store instruction with the first priority from the S data processing instructions and send it to the data store access pipeline Q2. When there is no data store instruction, a data replacement instruction or a data detection instruction with the second priority among the S data processing instructions is sent to the data store access pipeline Q2.
[0092] In addition, from Figure 6 it can be known that the data store access pipeline Q2 includes multiple circuit units: a second tag reading unit 41, a second tag comparison unit 42, a second data reading unit 43, a second data selection unit 44, a data writing unit 45, a second replacement selection unit 46, and a data comparison unit 47. Here, taking the execution of a data processing instruction in the data store access pipeline Q2 (any one of a data store instruction, a data replacement instruction, and a data detection instruction, assuming that the access address carried by the data processing instruction is a first access address, the first access address is used to locate a first data, and the first access address includes a first index field, a first tag field, and a first offset field) as an example, each circuit unit in the data store access pipeline Q2 will be introduced.
[0093] The second tag reading unit 41 is connected to the tag array Q2 (the tag array 13 among the P tag arrays 13 that is connected to the data store access pipeline Q2). The second tag reading unit 41 is used to determine a first cache group whose cache index matches the first index field from the N cache groups and send a first read signal for the first cache group to the tag array Q2.
[0094] The tag array Q2 is connected to the second tag comparison unit 42. The tag array Q2 is used to respond to the first read signal and return the cache tags of each cache line in the first cache group (a total of M cache tags) to the second tag comparison unit 42.
[0095] The second tag comparison unit 42 is connected to the second data selection unit 44 ( Figure 6 The connection line is not marked in Figure 6 ). The second tag comparison unit 42 is configured to match the M cache tags with the first tag field to obtain a tag matching result. If the tag matching result indicates that there is a cache tag of a target cache line in the M cache tags that matches the first tag field, the tag matching result is sent to the second data selection unit 44.
[0096] The second tag comparison unit 42 is also connected to the second replacement selection unit 46. The second tag comparison unit 42 is further configured to, if the tag matching result indicates that there is no cache tag in the M cache tags that matches the first tag field, send the tag matching result to the second replacement selection unit 46.
[0097] The second data reading unit 43 is connected to each of the M-way data memories 21 in the data array 11. The second data reading unit 43 is configured to send third read signals for corresponding cache lines in the first cache group to the M-way data memories 21 in the data array 11 respectively.
[0098] The M-way data memories 21 in the data array 11 are connected to the second data selection unit 44. The M-way data memories 21 in the data array 11 are configured to, in response to the corresponding third read signals, send the corresponding cache lines in the first cache group to the second data selection unit 44 (a total of M cache lines included in the first cache group are received).
[0099] The second data selection unit 44 is configured to obtain the cache tag of the target cache line determined based on the tag matching result (i.e., the cache tag that matches the first tag field), and determine the target cache line from each of the cache lines in the first cache group.
[0100] The second data selection unit 44 is also connected to the data write unit 45. In one embodiment, if the data processing instruction in the data storage access pipeline Q2 is a data storage instruction, the second data selection unit 44 is further configured to merge the target cache line and the second data to be stored carried by the data storage instruction according to the first offset field, so as to replace the first data in the target cache line with the second data, obtain merged data (including the second data and the adjacent data of the first data (i.e., the data other than the first data in the target cache line)), and send the merged data to the data write unit 45. The data write unit 45 is connected to the data array 11, and the data write unit 45 is configured to write the merged data into the target cache line in the data array 11. In one embodiment, if the data processing instruction in the data storage access pipeline Q2 is a data replacement instruction, the second data selection unit 44 is further configured to send the target cache line to the data write unit 45, and the data write unit 45 is configured to write the data in the target cache line into the write-out queue 19, and the write-out queue 19 is configured to write the data in the target cache line into the cache data source. Specifically, when implementing, the preset register 15 can wake up the write-out queue 19 to send a cache write signal to the cache data source when backfilling data into the target cache line in the data array, and the cache write signal includes the data in the target cache line, and the cache data source responds to the cache write signal to write the data in the target cache line into the cache data source. That is to say, when the cache line is replaced, the data it includes will be written back to the cache data source by the write-out queue 19, and then based on the cache data consistency mechanism, the data can be written back to the memory.
[0101] The second data selection unit 44 is also connected to the data comparison unit 47. In one embodiment, if the data processing instruction in the data storage access pipeline Q2 is a data detection instruction, the second data selection unit 44 is further configured to read the first data from the target cache line according to the first offset field, and send the first data to the data comparison unit 47. The data comparison unit 47 is configured to perform a consistency check on the first data and the third data to be detected carried by the data detection instruction (i.e., check whether the first data and the third data are the same), obtain a detection result, return the detection result to the detection emission unit 17, and the detection emission unit 17 sends the detection result to the circuit unit (such as the cache data source) that initiates the data detection instruction.
[0102] The second replacement selection unit 46 is connected to the preset register 15. The second replacement selection unit 46 is configured to determine the number of the cache line to be replaced (which may include the set number and the way number, and the way number is used to indicate which data memory the cache line to be replaced belongs to, and the set number is used to indicate which cache set the cache line to be replaced belongs to) if it is determined based on the tag matching result that there is no cache tag in the M cache tags that matches the first tag field (that is, there is no cache line in the data array 11 that contains the first data), and send the miss signal including the number of the cache line to be replaced to the preset register 15.
[0103] Wherein, the second replacement selection unit 46 may include a plurality of cache line selection units 461 and a multiplexer 462 (Mux). One cache line selection unit 461 corresponds to one cache line selection policy. Each cache line selection unit 461 is configured to select a candidate cache line from the data array according to the corresponding cache line selection policy. The multiplexer 462 is configured to select the cache line to be replaced from the multiple candidate cache lines selected by the plurality of cache line selection units 461. The cache line selection policy may include at least one of the following: First In First Out (FIFO) policy, Least Recent Used (LRU) policy, Least Freqency Used (LFU) policy, Random (RAND) policy, Pseudo Least Recent Used (PLRU) policy, and the present application does not limit this.
[0104] The preset register 15 is configured to generate a data replacement instruction for the cache line to be replaced based on the number of the cache line to be replaced in response to the miss signal, and send the data replacement instruction to any one of the P cache pipelines 12, such as the data storage access pipeline Q2, to schedule the data storage access pipeline to write the data in the cache line to be replaced into the cache data source, send a cache read signal including the first access address to the cache data source to obtain the fourth data including the first data from the cache data source, and backfill the fourth data to the cache line to be replaced in the data array 11. Specifically, the data storage access pipeline writes the data in the cache line to be replaced into the write queue 19 (located in the cache processing system 10). When the fourth data is backfilled to the cache line to be replaced in the data array 11, the preset register 15 wakes up the write queue 19 to send a cache write signal to the cache data source, and the cache write signal includes the data in the cache line to be replaced. The cache data source writes the data in the cache line to be replaced into the cache data source in response to the cache write signal.
[0105] Since the storage buffer can merge multiple data storage instructions in units of cache lines, the number of data storage instructions is small. In this application, the data detection instruction and the data replacement instruction are merged into the data storage access pipeline, which can improve the utilization rate of the data storage access pipeline. Optionally, Q2 can be 1, that is, the cache processing system 10 includes one data load access pipeline.
[0106] Please refer to Figure 7 , Figure 7 which is a schematic diagram of a sub-structure of a cache processing system provided by an embodiment of this application Figure 4 ; as Figure 7 shown, the cache processing system 10 further includes an update signal replication unit 20. The update signal replication unit 20 is connected to each of the P tag arrays 13. The update signal replication unit 20 is configured to receive a tag update signal (used to request an update of the cache tag, which can be sent by any one of the cache pipelines), and after copying the received tag update signal into P copies, send them to the P tag arrays 13 respectively. In this way, the P tag arrays 13 will update the corresponding cache tags based on the received tag update signals. Schematically, each of the P tag arrays 13 is a single-port RAM (Random Access Memory), with only one read / write port, and the stored content (i.e., the cache tag) is exactly the same. Since the amount of information of the cache tag is small and the area occupied by the tag array is relatively small, this way of tag array replication can be achieved at a relatively small area cost, and at the same time can greatly improve the data processing efficiency (achieved by accessing multiple tag arrays in parallel).
[0107] Please refer to Figure 8 , Figure 8 which is a schematic diagram of the structure of a data array provided by an embodiment of this application. As Figure 8 shown, the data array 11 includes M data memories (such as Figure 8 the data memory 1, data memory 2,... data memory M in Figure 8 ), and each data memory includes one or more data banks (equivalent to a storage unit, such as
[0108] the data bank 1,... data bank 3 in ). Among them, each data bank is used to store one or more cache lines. If the data banks accessed by multiple access sources (i.e., cache pipelines) are different, these multiple access sources can access the data array 11 in parallel. Through the multi-bank technology, parallel reading of the data array can be realized, the reading efficiency of the cache line can be improved, which is beneficial to improving the performance of the processor.
[0108] The above-mentioned cache processing system can be deployed in a computer device to provide cache processing services for the computer device. The computer device can include a terminal device and a server. Among them, the terminal device can be an electronic device, including but not limited to mobile phones, tablet computers, desktop computers, laptop computers, palmtop computers, in-vehicle devices, augmented reality / virtual reality (AR / VR) devices, head-mounted displays, smart TVs, wearable devices, smart speakers, digital cameras, cameras, and other mobile Internet devices (MIDs) with network access capabilities, or terminal devices in scenarios such as trains, ships, and flights. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, vehicle-road collaboration, content delivery network (CDN), and big data and artificial intelligence platforms. The so-called cloud computing refers to the delivery and usage model of IT (Internet) infrastructure, which means obtaining the required resources in a on-demand and easily expandable manner through the network; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services in a on-demand and easily expandable manner through the network. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balance.
[0109] It should be noted that when the embodiments of the present application are applied to specific products or technologies, such as when obtaining the first data, permission or consent from the target object (i.e., the owner of the corresponding data) needs to be obtained, and the relevant data collection, use, and processing processes need to comply with the relevant laws, regulations, and standards of the region, conform to the principles of legality, legitimacy, and necessity, and do not involve obtaining data types prohibited or restricted by laws and regulations.
[0110] Please refer to Figure 9 , Figure 9It is a schematic flowchart of a cache processing method provided by an embodiment of this application. This cache processing method is applied to the above-mentioned cache processing system 10, that is, this cache processing method can be executed by the above-mentioned cache processing system 10. In this cache processing system 10, the data processing instructions in each cache pipeline can be executed in parallel based on their respective connected tag arrays; hereinafter, taking the first cache pipeline (any one of the P cache pipelines) as an example, this cache processing method will be described; this cache processing method mainly includes but is not limited to the following steps S901 to S903:
[0111] S901. Obtain the first data processing instruction in the first cache pipeline; the first data processing instruction carries a first access address, and the first access address is used to locate the first data indicated by the first data processing instruction.
[0112] In one embodiment, if the first cache pipeline is a data load access pipeline, the first cache pipeline is connected to a first load emission unit (any one of the Q1 load emission units 14), and the first load emission unit is used to send a data load instruction to the first cache pipeline. That is to say, the first data processing instruction in the first cache pipeline is a data load instruction.
[0113] In another embodiment, if the first cache pipeline is a data store access pipeline, the first cache pipeline is connected to a preset scheduler, and the preset scheduler is connected to a store buffer, a detection emission unit, and a preset register. The store buffer is used to send a data store instruction to the preset scheduler, the detection emission unit is used to send a data detection instruction to the preset scheduler, and the preset register is used to send a data replacement instruction to the preset scheduler. The data store instruction, the data replacement instruction, and the data detection instruction each have their corresponding priorities. The preset scheduler can be called to determine the sending order of the S data processing instructions obtained by the preset scheduler according to the priorities, and select the first data processing instruction from the S data processing instructions according to the sending order, and send the first data processing instruction to the first cache pipeline. That is to say, the first data processing instruction in the first cache pipeline is any one of the data store instruction, the data replacement instruction, and the data detection instruction. Specifically, the data store instruction may have a first priority, and the data replacement instruction and the data detection instruction may have a second priority. The data store instruction with the first priority among the S data processing instructions will be preferentially sent to the first cache pipeline. When there is no data store instruction, the data replacement instruction or the data detection instruction with the second priority among the S data processing instructions (which can be determined according to the acquisition order) will be sent to the first cache pipeline.
[0114] The first data processing instruction carries a first access address, which is used to locate the first data indicated by the first data processing instruction. The first access address includes a first tag field, a first index field, and a first offset field. In one implementation, the first data indicated by the first data processing instruction can be accessed from a cache (which may include an L1 cache, an L2 cache, and an L3 cache) based on the first access address. If the cache miss occurs, the first data indicated by the first data processing instruction is accessed from the memory based on the first access address.
[0115] In one embodiment, if the store buffer obtains multiple data store requests, and the tag fields and index fields in the access addresses carried by the multiple data store instructions are the same, the multiple data store instructions are merged into a new data store instruction, and the new data store instruction is sent to a preset scheduler through the store buffer. The new data store instruction may include the access addresses carried by the multiple data store instructions respectively. When the first data processing instruction is the new data store instruction, each access address in the new data store instruction can be used as the first access address.
[0116] S902. Determine a first cache line for storing the first data from the data array based on a first tag array connected to a first cache pipeline.
[0117] The data array includes M-way data memories, and each data memory includes N cache lines; the data array is divided into N cache groups, each cache group corresponds to a cache index respectively, and each cache group includes M cache lines, and the M cache lines belong to different data memories respectively; both M and N are positive integers.
[0118] In one embodiment, determining a first cache line for storing the first data from the data array based on a first tag array connected to a first cache pipeline includes: determining a first cache group whose cache index matches the first index field (matches if consistent, does not match if inconsistent) from the N cache groups, reading the cache tags of each cache line in the first cache group from the first tag array to obtain M cache tags. Matching the read M cache tags with the first tag field (matches if consistent, does not match if inconsistent) to obtain a tag matching result. If the tag matching result indicates that there is a cache tag of a target cache line that matches the first tag field among the read M cache tags, the target cache line is determined as the first cache line for storing the first data. If the tag matching result indicates that there is no cache tag that matches the first tag field among the read M cache tags, it means that the data array does not include a cache line storing the first data (i.e., the data array misses).
[0119] S903. Perform a data processing operation on the first data in the first cache line according to the first data processing instruction.
[0120] In one embodiment, if the first cache pipeline is a data load access pipeline, the first data memory predicted to store the first data is retrieved from the M-way data memory, and this prediction can be achieved based on a way prediction mechanism. The second cache line is read from the first data memory, and the second cache line belongs to the first cache group. The cache tag of the first cache line (i.e., the cache tag of the target cache line) determined based on the tag matching result is obtained. If the cache tag of the first cache line is the same as the cache tag of the second cache line, it indicates that the second cache line is the first cache line. In this way, the data memory storing the corresponding data can be predicted in advance through the way prediction mechanism, which can improve the parallel loading efficiency. The first data can be read from the read second cache line based on the first offset, and the first data is returned to the first load issue unit for data loading processing. This data loading processing means returning the first data to the circuit unit receiving the first data. For example, if the first data processing instruction requests to load data into a register of the processor, the first load issue unit can send the first data to the register of the processor.
[0121] In one embodiment, if the first cache pipeline is a data store access pipeline, M cache lines in the first cache group are respectively read from the M-way data memory. The cache tag of the first cache line (i.e., the cache tag of the target cache line) determined based on the tag matching result is obtained, and the first cache line is determined from the read M cache lines according to the cache tag of the first cache line. If the first data processing instruction is a data load instruction (which also carries the second data to be stored), then according to the first offset field, the determined first cache line and the second data are merged to replace the first data in the first cache line with the second data, obtaining merged data (including the second data and the adjacent data of the first data (i.e., the data other than the first data in the first cache line)), and the merged data is written into the first cache line in the data array.
[0122] In summary, please refer to Figure 10 , Figure 10 which is a schematic diagram of the processing flow of a data storage instruction provided by an embodiment of the present application. The meanings of the following steps are introduced Figure 10 as follows:
[0123] S1. The data storage instruction (carrying an access address, and the access address includes an index field, a tag field, and an offset field) is initiated by the storage buffer, passes through a preset scheduler, and is input into the data store access pipeline.
[0124] S2. The second tag reading unit initiates a first read request to the connected tag array.
[0125] S3. The connected tag array returns tag information (including one or more cached tags) for the first read request to the second tag comparison unit, and the second tag comparison unit compares the cached tags with the tag fields to determine whether the data array hits the data indicated by the data storage instruction.
[0126] S4. If there is a hit, the second data reading unit issues a second read request to the data array.
[0127] S5. The data array returns to the second data selection unit a cache line containing the data indicated by the data storage instruction for the second read request.
[0128] S6. The second data selection unit sends the cache line to the data writing unit, and the data writing unit replaces the data indicated by the data storage instruction in the cache line with the data to be stored carried by the data storage instruction to obtain merged data, and writes the merged data into the data array.
[0129] In the embodiment of the present application, the data storage instruction is divided into multiple execution steps, so that different functional circuit units included in the data storage access pipeline can be introduced to execute the multiple execution steps, which is beneficial to realizing parallel operation of the data storage instruction and making full use of computing resources.
[0130] In one embodiment, if the first cache pipeline is a data storage access pipeline, M cache lines in the first cache group are respectively read from the M-way data memory. Obtain the cache tag of the first cache line determined based on the tag matching result (i.e., the cache tag of the target cache line), and determine the first cache line from the M cache lines read according to the cache tag of the first cache line. If the first data processing instruction is a data detection instruction (which also carries the third data to be detected), then according to the first offset field, the first data is read from the first cache line. The first data and the third data are subjected to consistency detection to obtain a detection result. The detection result is returned to the detection emission unit that issues the first data processing instruction, and the detection emission unit returns the detection result to the circuit unit (such as the cache data source) that receives the detection result.
[0131] In one embodiment, if the first cache pipeline is a data storage access pipeline, M cache lines in the first cache group are respectively read from the M-way data memory. Obtain the cache tag of the first cache line determined based on the tag matching result (i.e., the cache tag of the target cache line), and determine the first cache line from the M cache lines read according to the cache tag of the first cache line. If the first data processing instruction is a data replacement instruction, the data in the first cache line can be written into the write queue, and the write queue writes the data in the first cache line into the cache data source.
[0132] In one embodiment, if the tag matching result indicates that there is no cache line in the first cache group whose cache tag matches the first tag domain, a cache line to be replaced is determined from the first cache group. The cache line to be replaced can be a cache line that does not store data, or a cache line that has not been accessed for a long time (e.g., exceeding a preset duration), etc. A data replacement instruction for the cache line to be replaced (which may include the number of the cache line to be replaced) is generated. The data replacement instruction is sent to the second cache pipeline, and the data replacement instruction in the second cache pipeline is obtained. Based on the second tag array connected to the second cache pipeline, the cache line to be replaced is determined from the data array; the implementation process is the same as the process of determining the first cache line for storing the first data from the data array based on the first tag array connected to the first cache pipeline, which will not be elaborated here. The cache line to be replaced is read from the data array, the data in the cache line to be replaced is written into the write queue, and the data in the cache line to be replaced in the write queue is written into the cache data source. Specifically, when the cache line to be replaced is backfilled with new data, the data in the cache line to be replaced in the write queue is written into the cache data source, so that the data in the cache line to be replaced included in the data array is written back to the cache data source.
[0133] In one embodiment, a replacement cache tag may also be generated based on the tag domain in the first access address, that is, the tag domain in the first access address is used as the replacement cache tag. A tag update signal for the cache line to be replaced is generated, and the replacement cache tag is included in the tag update signal. Based on the tag update signal, the cache tag of the cache line to be replaced in each tag array is updated to the replacement cache tag. For example, the tag update signal can be copied into P copies by an update signal copying unit and sent to P tag arrays respectively, so that each tag array responds to the tag update signal and updates the cache tag of the cache line to be replaced to the replacement cache tag.
[0134] In one embodiment, the data array includes M-way data memories, each data memory includes one or more data storage banks, and each data storage bank is used to store one or more cache lines. Reading the first data from the first cache line includes: if the first data storage bank (any data storage bank in the data array) includes the first cache line and no other cache pipeline accesses the first data storage bank, the first data is read from the first cache line in the first data storage bank. Here, the other cache pipelines are the cache pipelines other than the first cache pipeline among the P cache pipelines.
[0135] In summary, please refer to Figure 11 , Figure 11 which is a schematic diagram of the processing flow of a data processing instruction provided by an embodiment of this application. The meanings of the following steps are introduced Figure 11 as follows:
[0136] S1. A data loading instruction (carrying an access address, which includes an index field, a tag field, and an offset field) is issued by a load emission unit and enters a data loading access pipeline. A first tag reading unit in the data loading access pipeline initiates a read request for a connected tag array.
[0137] S2. The connected tag array returns tag information (including one or more cached tags) to a first tag comparison unit for the read request. The first tag comparison unit compares the cached tags with the tag field to determine whether the data array hits the data indicated by the data loading instruction.
[0138] S3. When the first tag comparison unit finds a miss, a first replacement selection unit selects a cache line to be replaced.
[0139] S4. The first replacement selection unit sends a miss signal and the number of the cache line to be replaced to a preset register.
[0140] S5. In response to the miss signal, the preset register sends a cache read signal to a cache data source.
[0141] S6. The preset register obtains the data indicated by the data loading instruction returned by the cache data source in response to the cache read signal, and backfills the data indicated by the data loading instruction to the cache line to be replaced in the data array.
[0142] S7. The preset register generates a data replacement instruction for the cache line to be replaced based on the number of the cache line to be replaced. After being arbitrated by a preset scheduler, the data replacement instruction enters the data loading access pipeline.
[0143] S8. A second data reading unit initiates a read access to the cache line to be replaced in the data array.
[0144] S9. The data array returns the cache line to be replaced to the second data reading unit. The second data reading unit sends the cache line to be replaced to a data writing unit, and the data writing unit writes the data in the cache line to be replaced into a write-out queue, waiting to be sent to a lower-level cache data source.
[0145] S10. When the data indicated by the data loading instruction is backfilled to the cache line to be replaced in the data array, the write-out queue writes the data in the cache line to be replaced into the cache data source.
[0146] In the embodiment of the present application, it is beneficial to use a preset register to handle cache misses, so that the cache processing system can continue to respond to other data processing instructions, and the parallel processing efficiency can be improved.
[0147] It can be seen that the present application proposes a cache processing system with a high-concurrency architecture. The cache processing system for accessing caches (including a data array and a tag array) includes P cache pipelines. The parallel access of the P cache pipelines to the tag array is achieved by replicating the tag array, and the parallel access of the P cache pipelines to the data array is achieved by the multi-bank technology, enabling multiple data processing instructions to be processed in parallel, thereby effectively improving the data processing efficiency. At the same time, the data detection instruction and the data replacement instruction are merged into the data storage access pipeline, which can improve the utilization rate of the data storage access pipeline.
[0148] The device of the embodiment of the present application is provided below. Next, in combination with the cache processing method provided in the above embodiment of the present application, the related device of the embodiment of the present application will be introduced accordingly.
[0149] Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of a cache processing device provided by an embodiment of the present application. As Figure 12 shown, the cache processing device 1200 is provided with a cache processing system. The cache processing system includes a data array, P cache pipelines, and P tag arrays, where P is a positive integer. One cache pipeline is connected to one tag array, and all P cache pipelines are connected to the data array. The data array is used to store data in the form of cache lines, and the tag array is used to store the cache tags of the cache lines. The cache processing device 1200 can be used to execute the corresponding steps in the cache processing method provided by the embodiment of the present application. Specifically, the cache processing device 1200 may specifically include:
[0150] An acquisition module 1201, configured to acquire a first data processing instruction in a first cache pipeline. The first data processing instruction carries a first access address, and the first access address is used to locate the first data indicated by the first data processing instruction.
[0151] A processing module 1202, configured to determine a first cache line for storing the first data from the data array based on the first tag array connected to the first cache pipeline.
[0152] The processing module 1202 is configured to perform a data processing operation on the first data in the first cache line according to the first data processing instruction.
[0153] Wherein, the first cache pipeline is any one of the P cache pipelines. The data processing instructions in each cache pipeline are executed in parallel based on the respective connected tag arrays.
[0154] In one embodiment, the data array includes M data memories, and each data memory includes N cache lines; the data array is divided into N cache groups, each cache group corresponds to a cache index respectively, and each cache group includes M cache lines, and the M cache lines belong to different data memories respectively; both M and N are positive integers; the first access address includes a first index field and a first tag field;
[0155] The processing module 1202 is specifically configured to call a tag reading unit in the first cache pipeline (if the first cache pipeline is a data load access pipeline, then the tag reading unit is the first tag reading unit; if the first cache pipeline is a data store access pipeline, then the tag reading unit is the second tag reading unit) to determine a first cache group whose cache index matches the first index field from the N cache groups, and read the cache tags of each cache line in the first cache group from the first tag array to obtain M cache tags;
[0156] The processing module 1202 is further specifically configured to call a tag comparison unit in the first cache pipeline (if the first cache pipeline is a data load access pipeline, then the tag comparison unit is the first tag comparison unit; if the first cache pipeline is a data store access pipeline, then the tag comparison unit is the second tag comparison unit) to perform matching processing on the read M cache tags and the first tag field to obtain a tag matching result. If the tag matching result indicates that there is a cache tag of a target cache line that matches the first tag field among the read M cache tags, then determine the target cache line as the first cache line for storing the first data.
[0157] In one embodiment, the first data processing instruction is a data load instruction, and among the P cache pipelines, there are Q1 data load access pipelines, where Q1 is a positive integer and Q1 is less than or equal to P; the cache processing system further includes Q1 load issue units, one load issue unit is connected to one data load access pipeline, and each load issue unit is used to send a data load instruction to the data load access pipeline it is connected to; the first cache pipeline is any one of the Q1 data load access pipelines, and the first cache pipeline is connected to the first load issue unit among the Q1 load issue units; the first access address further includes a first offset field;
[0158] The processing module 1202 is further specifically configured to call a way prediction unit in the first cache pipeline to predict a first data memory that stores the first data from the M data memories;
[0159] The processing module 1202 is further specifically configured to call a first data reading unit in the first cache pipeline to read a second cache line from the first data memory, and the second cache line belongs to the first cache group;
[0160] The processing module 1202 is further specifically configured to call the first data selection unit in the first cache pipeline to obtain the cache tag of the first cache line determined based on the tag matching result. If the cache tag of the first cache line is the same as the cache tag of the second cache line, then according to the first offset field, read the first data from the second cache line, and send the first data to the first load and issue unit for data loading processing.
[0161] In one embodiment, the first data processing instruction is a data storage instruction, and the data storage instruction further carries the second data to be stored; among the P cache pipelines, there are Q2 data storage access pipelines, where Q2 is a positive integer and Q2 is less than or equal to P; the first cache pipeline is any one of the Q2 data storage access pipelines; the first access address further includes a first offset field;
[0162] The processing module 1202 is further specifically configured to call the second data reading unit in the first cache pipeline to respectively read M cache lines in the first cache group from the M-way data memory, and send the M read cache lines to the second data selection unit in the first cache pipeline;
[0163] The processing module 1202 is further specifically configured to call the second data selection unit in the first cache pipeline to obtain the cache tag of the first cache line determined based on the tag matching result, determine the first cache line from the M read cache lines according to the cache tag of the first cache line, and according to the first offset field, perform a merging process on the determined first cache line and the second data, so as to replace the first data in the first cache line with the second data to obtain merged data;
[0164] The processing module 1202 is further specifically configured to call the data write unit in the first cache pipeline to write the merged data into the first cache line in the data array.
[0165] In one embodiment, the first data processing instruction is a data detection instruction, and the data detection instruction further carries the third data to be detected; among the P cache pipelines, there are Q2 data storage access pipelines, where Q2 is a positive integer and Q2 is less than or equal to P; the first cache pipeline is any one of the Q2 data storage access pipelines; the first access address further includes a first offset field; the cache processing system further includes a detection and issue unit for sending data detection instructions;
[0166] The processing module 1202 is further specifically configured to call the second data reading unit in the first cache pipeline to respectively read M cache lines in the first cache group from the M-way data memory;
[0167] The processing module 1202 is further specifically configured to call a second data selection unit in the first cache pipeline to obtain a cache tag of a first cache line determined based on a tag matching result, determine a first cache line from the read M cache lines according to the cache tag of the first cache line, and read first data from the first cache line according to a first offset field;
[0168] The processing module 1202 is further specifically configured to call the data comparison unit in the first cache pipeline to perform consistency detection on the first data and the third data, obtain a detection result, and return the detection result to the detection emission unit that sends the first data processing instruction.
[0169] In one embodiment, the processing module 1202 is further specifically used to call a replacement selection unit in the first cache pipeline (if the first cache pipeline is a data loading access pipeline, the replacement selection unit is a first replacement selection unit; if the first cache pipeline is a data storage access pipeline, the replacement selection unit is a second replacement selection unit) to determine a cache line to be replaced from the first cache group if the tag matching result indicates that the first cache group does not have a cache line whose cache tag matches the first tag field;
[0170] The processing module 1202 is further specifically used to call a preset register in the cache processing system to generate a data replacement instruction for the cache line to be replaced, and send the data replacement instruction to the second cache pipeline to schedule the second cache pipeline to write the data in the cache line to be replaced into the cache data source. The second cache pipeline is any one of the P cache pipelines. The data storage access pipeline obtains fourth data including the first data from the cache data source, backfills the fourth data into the cache line to be replaced in the data array, and determines the cache line to be replaced after the backfill processing as the first cache line for storing the first data.
[0171] In one embodiment, the cache processing system also includes a write-out queue; the processing module 1202 is also used to call the second cache pipeline to obtain a data replacement instruction in the second cache pipeline; based on the second tag array connected to the second cache pipeline, read the cache line to be replaced from the data array; write the data in the cache line to be replaced into the write-out queue; write the data in the cache line to be replaced in the write-out queue to the cache data source.
[0172] In one embodiment, the processing module 1202 is further specifically configured to call the first cache pipeline to generate a replacement cache tag based on the tag field in the first access address, and generate a tag update signal for the cache line to be replaced, wherein the tag update signal includes the replacement cache tag;
[0173] The processing module 1202 is further specifically configured to call the update signal replication unit in the cache processing system to update the cache tags of the cache lines to be replaced in each tag array to replacement cache tags based on the tag update signal.
[0174] In one embodiment, the first cache pipeline is a data storage access pipeline; the cache processing system further includes a storage buffer, a detection emission unit, a preset register, and a preset scheduler; the storage buffer is configured to send a data storage instruction to the preset scheduler, the detection emission unit is configured to send a data detection instruction to the preset scheduler, and the preset register is configured to send a data replacement instruction to the preset scheduler;
[0175] The processing module 1202 is further specifically configured to call the preset scheduler in the first cache pipeline to determine the sending order of the S data processing instructions obtained by the preset scheduler according to the priority. The S data processing instructions include one or more of a data storage instruction, a data replacement instruction, and a data detection instruction, and the data storage instruction, the data replacement instruction, and the data detection instruction each have their corresponding priorities; S is a positive integer. The first data processing instruction is selected from the S data processing instructions according to the sending order, and the first data processing instruction is sent to the first cache pipeline.
[0176] In one embodiment, if the storage buffer obtains multiple data storage requests, and the tag fields and index fields in the access addresses carried by the multiple data storage instructions are the same, the processing module 1202 is further specifically configured to call the storage buffer in the first cache pipeline to merge the multiple data storage instructions into a new data storage instruction, and send the new data storage instruction to the preset scheduler.
[0177] In one embodiment, the data array includes M-way data memories, each way of data memory includes one or more data storage banks, and each data storage bank is configured to store one or more cache lines;
[0178] If the first data storage bank includes a first cache line and no other cache pipeline accesses the first data storage bank, the processing module 1202 is further specifically configured to call the data reading unit in the first cache pipeline to read the first data from the first cache line in the first data storage bank; wherein, the first data storage bank is any data storage bank in the data array, and the other cache pipelines are the cache pipelines other than the first cache pipeline among the P cache pipelines.
[0179] In the embodiment of the present application, a cache processing system with a high-concurrency architecture is designed by using a data array (for storing data in the form of cache lines), P cache pipelines, and P tag arrays (for storing cache tags of cache lines, and the cache tags stored in the P tag arrays are the same). In this cache processing system, the P cache pipelines are all connected to the data array, and one cache pipeline is connected to one tag array, so that the data processing instructions in each cache pipeline can be executed in parallel based on the tag arrays connected to them respectively. Specifically, the first cache pipeline is any one of the P cache pipelines. Executing the first data processing instruction in the first cache pipeline based on the first tag array connected to the first cache pipeline includes: obtaining the first data processing instruction in the first cache pipeline, where the first data processing instruction includes a first access address for locating the first data, and based on the first tag array connected to the first cache pipeline, determining the first cache line for storing the first data from the data array, and performing a data processing operation on the first data in the first cache line according to the first data processing instruction. It can be seen that in the embodiment of the present application, P tag arrays are instantiated based on the P cache pipelines. In this way, the data processing instructions in the P cache pipelines can be processed in parallel based on the tag arrays connected to them respectively, greatly improving the data processing efficiency.
[0180] Please refer to Figure 13 , Figure 13 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 1300 includes a computer-readable storage medium 1301 and a chip 1302. Among them, the computer-readable storage medium 1301 and the chip 1302 can be connected through a bus or other means. The chip 1302 includes a processor 1303 and a cache processing device 1304, and a cache processing system is provided in the cache processing device 1304. The chip 1302 is used to call the cache processing system provided in the cache processing device 1304 to execute the steps executed by the cache processing system in the foregoing method embodiment. The processor 1303 is used to send a data processing instruction to the cache processing device 1304 to trigger the cache processing system in the cache processing device 1304 to execute the steps executed by the cache processing system in the foregoing method embodiment. The chip 1302 may further include a memory, and the memory is used to provide the data stored in the cache processing system. The computer-readable storage medium 1301 is used to store a computer program, and the computer program includes program instructions. In one implementation, the computer device 1300 (the chip 1302 therein) is used to call the program instructions stored in the computer-readable storage medium 1301 to perform the following operations:
[0181] Obtain the first data processing instruction in the first cache pipeline; the first data processing instruction carries a first access address, and the first access address is used to locate the first data indicated by the first data processing instruction;
[0182] Based on the first tag array connected to the first cache pipeline, determine the first cache line for storing the first data from the data array;
[0183] Perform a data processing operation on the first data in the first cache line according to the first data processing instruction;
[0184] Wherein, the first cache pipeline is any one of the P cache pipelines; the data processing instructions in each cache pipeline are executed in parallel based on their respective connected tag arrays.
[0185] In one implementation, the data array includes an M-way data memory, and each data memory includes N cache lines; the data array is divided into N cache groups, each cache group corresponds to a cache index respectively, and each cache group includes M cache lines, and the M cache lines belong to different data memories respectively; both M and N are positive integers; the first access address includes a first index field and a first tag field;
[0186] When the chip 1302 is used to determine the first cache line for storing the first data from the data array based on the first tag array connected to the first cache pipeline, the following steps are specifically executed:
[0187] Determine the first cache group whose cache index matches the first index field from the N cache groups;
[0188] Read the cache tags of each cache line in the first cache group from the first tag array to obtain M cache tags;
[0189] Perform a matching process on the read M cache tags and the first tag field to obtain a tag matching result;
[0190] If the tag matching result indicates that there is a cache tag of the target cache line among the read M cache tags that matches the first tag field, then determine the target cache line as the first cache line for storing the first data.
[0191] In one implementation, the first data processing instruction is a data load instruction. Among the P cache pipelines, there are Q1 data load access pipelines, where Q1 is a positive integer and Q1 is less than or equal to P. The cache processing system further includes Q1 load issue units. One load issue unit is connected to one data load access pipeline. Each load issue unit is used to send a data load instruction to the data load access pipeline it is connected to. The first cache pipeline is any one of the Q1 data load access pipelines, and the first cache pipeline is connected to the first load issue unit among the Q1 load issue units. The first access address further includes a first offset field.
[0192] When the chip 1302 is used to perform a data processing operation on the first data in the first cache line according to the first data processing instruction, the following steps are specifically executed:
[0193] Predict the first data memory that stores the first data in the M-way data memory.
[0194] Read the second cache line from the first data memory. The second cache line belongs to the first cache group.
[0195] Obtain the cache tag of the first cache line determined based on the tag matching result.
[0196] If the cache tag of the first cache line is the same as the cache tag of the second cache line, read the first data from the second cache line according to the first offset field.
[0197] Send the first data to the first load issue unit for data load processing.
[0198] In one implementation, the first data processing instruction is a data store instruction, and the data store instruction further carries the second data to be stored. Among the P cache pipelines, there are Q2 data store access pipelines, where Q2 is a positive integer and Q2 is less than or equal to P. The first cache pipeline is any one of the Q2 data store access pipelines. The first access address further includes a first offset field.
[0199] When the chip 1302 is used to perform a data processing operation on the first data in the first cache line according to the first data processing instruction, the following steps are specifically executed:
[0200] Read M cache lines in the first cache group from the M-way data memory respectively.
[0201] Obtain the cache tag of the first cache line determined based on the tag matching result.
[0202] Determine the first cache line from the M cache lines read according to the cache tag of the first cache line.
[0203] According to the first offset field, the determined first cache line and the second data are merged, so as to replace the first data in the first cache line with the second data, and merged data is obtained;
[0204] Write the merged data into the first cache line in the data array.
[0205] In one implementation, the first data processing instruction is a data detection instruction, and the third data to be detected is also carried in the data detection instruction; among the P cache pipelines, there are Q2 data storage access pipelines, Q2 is a positive integer and Q2 is less than or equal to P; the first cache pipeline is any one of the Q2 data storage access pipelines; the first access address also includes a first offset field; the cache processing system further includes a detection emission unit for sending the data detection instruction;
[0206] When the chip 1302 is used to perform a data processing operation on the first data in the first cache line according to the first data processing instruction, the following steps are specifically executed:
[0207] Read M cache lines in the first cache group from the M-way data memory respectively;
[0208] Obtain the cache tag of the first cache line determined based on the tag matching result;
[0209] Determine the first cache line from the M cache lines read according to the cache tag of the first cache line;
[0210] Read the first data from the first cache line according to the first offset field;
[0211] Perform a consistency detection on the first data and the third data to obtain a detection result;
[0212] Return the detection result to the detection emission unit that sends the first data processing instruction.
[0213] In one implementation, the chip 1302 is further used to execute the following steps:
[0214] If the tag matching result indicates that there is no cache line in the first cache group whose cache tag matches the first tag field, determine the cache line to be replaced from the first cache group;
[0215] Generate a data replacement instruction for the cache line to be replaced;
[0216] Send the data replacement instruction to the second cache pipeline to schedule the second cache pipeline to write the data in the cache line to be replaced into the cache data source; the second cache pipeline is any one of the data storage access pipelines in the P cache pipelines;
[0217] Obtain a fourth data including a first data from a cache data source based on a first access address;
[0218] Backfill the fourth data into a cache line to be replaced in the data array, and determine the cache line to be replaced after the backfill process as a first cache line for storing the first data.
[0219] In one implementation, the cache processing system further includes a write-out queue; the chip 1302 is further configured to perform the following steps:
[0220] Obtain a data replacement instruction in a second cache pipeline;
[0221] Read a cache line to be replaced from the data array based on a second tag array connected to the second cache pipeline;
[0222] Write the data in the cache line to be replaced into the write-out queue;
[0223] Write the data in the cache line to be replaced in the write-out queue into the cache data source.
[0224] In one implementation, the chip 1302 is further configured to perform the following steps:
[0225] Generate a replacement cache tag based on a tag field in the first access address;
[0226] Generate a tag update signal for the cache line to be replaced, where the tag update signal includes the replacement cache tag;
[0227] Based on the tag update signal, update the cache tag of the cache line to be replaced in each tag array to the replacement cache tag.
[0228] In one implementation, the first cache pipeline is a data storage access pipeline; the cache processing system further includes a store buffer, a detection and issue unit, a preset register, and a preset scheduler; the store buffer is configured to send a data storage instruction to the preset scheduler, the detection and issue unit is configured to send a data detection instruction to the preset scheduler, and the preset register is configured to send a data replacement instruction to the preset scheduler;
[0229] The chip 1302 is further configured to perform the following steps:
[0230] Determine the sending order of S data processing instructions obtained by the preset scheduler according to priorities, where the S data processing instructions include one or more of a data storage instruction, a data replacement instruction, and a data detection instruction, and the data storage instruction, the data replacement instruction, and the data detection instruction each have their respective corresponding priorities; S is a positive integer;
[0231] Select a first data processing instruction from the S data processing instructions according to the sending order, and send the first data processing instruction to the first cache pipeline.
[0232] In one implementation, the chip 1302 is further configured to perform the following steps:
[0233] If the storage buffer obtains multiple data storage requests, and the tag fields and index fields in the access addresses carried by the multiple data storage instructions are the same, then merge the multiple data storage instructions into a new data storage instruction;
[0234] Send the new data storage instruction to a preset scheduler through the storage buffer.
[0235] In one implementation, the data array includes M data memories, each data memory includes one or more data storage banks, and each data storage bank is used to store one or more cache lines;
[0236] When the chip 1302 is used to read the first data from the first cache line, it specifically performs the following steps:
[0237] If the first data storage bank includes the first cache line and no other cache pipelines access the first data storage bank, then read the first data from the first cache line in the first data storage bank;
[0238] Wherein, the first data storage bank is any data storage bank in the data array, and the other cache pipelines are the cache pipelines except the first cache pipeline among the P cache pipelines.
[0239] Based on the same inventive concept, the principle of solving problems and the beneficial effects of the computer device provided in the embodiments of the present application are similar to the principle of solving problems and the beneficial effects of the cache processing method in the method embodiments of the present application. For details, please refer to the principle and beneficial effects of the method implementation. For the sake of brevity, they will not be elaborated here.
[0240] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of the module or unit.
[0241] In addition, it should be noted here that: The embodiments of the present application further provide a computer-readable storage medium, and a computer program is stored in the computer-readable storage medium. The computer program includes program instructions. When the chip executes the above program instructions, it can execute the methods in the corresponding embodiments described above. Therefore, details will not be repeated here. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located at one place, or executed on multiple computer devices distributed at multiple places and interconnected through a communication network.
[0242] According to one aspect of the present application, the embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The chip of the computer device reads the computer instructions from the computer-readable storage medium, and the chip executes the computer instructions, so that the computer device can execute the methods in the corresponding embodiments described above. Therefore, details will not be repeated here.
[0243] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0244] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data processing device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.
[0245] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A cache processing method, characterized in that, The method is applied to a cache processing system, which includes a data array, P cache pipelines, and P tag arrays, where P is a positive integer; one of the cache pipelines is connected to one of the tag arrays, and the P cache pipelines are all connected to the data array; The data array is used to store data in the form of cache lines; The tag array is used to store cache tags of the cache lines; the method includes: Obtain a first data processing instruction in a first cache pipeline; the first data processing instruction carries a first access address, and the first access address is used to locate the first data indicated by the first data processing instruction; Based on the first tag array connected to the first cache pipeline, determine a first cache line in the data array for storing the first data; Perform a data processing operation on the first data in the first cache line according to the first data processing instruction; Wherein, the first cache pipeline is any one of the P cache pipelines; the data processing instructions in each cache pipeline are executed in parallel based on their respective connected tag arrays.
2. The method according to claim 1, characterized in that, The data array includes M-way data memories, and each data memory includes N cache lines; the data array is divided into N cache groups, each cache group corresponds to a cache index respectively, and each cache group includes M cache lines, and the M cache lines belong to different data memories respectively; Both M and N are positive integers; the first access address includes a first index field and a first tag field; The step of determining, based on the first tag array connected to the first cache pipeline, a first cache line in the data array for storing the first data includes: Determine a first cache group in the N cache groups whose cache index matches the first index field; Read the cache tags of each cache line in the first cache group from the first tag array to obtain M cache tags; Perform a matching process on the M cache tags read and the first tag field to obtain a tag matching result; If the tag matching result indicates that there is a cache tag of a target cache line among the M cache tags read that matches the first tag field, then determine the target cache line as the first cache line for storing the first data.
3. The method according to claim 2, wherein The first data processing instruction is a data load instruction, and among the P cache pipelines, there are Q1 data load access pipelines, where Q1 is a positive integer and Q1 is less than or equal to P; the cache processing system further includes Q1 load issue units, one load issue unit is connected to one of the data load access pipelines, and each load issue unit is used to send a data load instruction to the data load access pipeline it is connected to; the first cache pipeline is any one of the Q1 data load access pipelines, and the first cache pipeline is connected to the first load issue unit among the Q1 load issue units; the first access address further includes a first offset field; Performing a data processing operation on the first data in the first cache line according to the first data processing instruction includes: Predicting a first data memory that stores the first data in the M-way data memory; Reading a second cache line from the first data memory, where the second cache line belongs to the first cache group; Obtaining the cache tag of the first cache line determined based on the tag matching result; If the cache tag of the first cache line is consistent with the cache tag of the second cache line, reading the first data from the second cache line according to the first offset field; Sending the first data to the first load and issue unit for data loading processing.
4. The method according to claim 2, wherein The first data processing instruction is a data storage instruction, and the data storage instruction also carries second data to be stored; among the P cache pipelines, there are Q2 data storage access pipelines, Q2 is a positive integer and Q2 is less than or equal to P; the first cache pipeline is any one of the Q2 data storage access pipelines; the first access address also includes a first offset field; Performing a data processing operation on the first data in the first cache line according to the first data processing instruction includes: Reading M cache lines in the first cache group from the M-way data memory respectively; Obtaining the cache tag of the first cache line determined based on the tag matching result; Determining the first cache line from the M cache lines read according to the cache tag of the first cache line; According to the first offset field, performing a merge process on the determined first cache line and the second data to replace the first data in the first cache line with the second data to obtain merged data; Writing the merged data into the first cache line in the data array.
5. The method according to claim 2, wherein The first data processing instruction is a data detection instruction, and the data detection instruction also carries third data to be detected; among the P cache pipelines, there are Q2 data storage access pipelines, Q2 is a positive integer and Q2 is less than or equal to P; the first cache pipeline is any one of the Q2 data storage access pipelines; the first access address also includes a first offset field; the cache processing system also includes a detection and issue unit for sending data detection instructions; Performing a data processing operation on the first data in the first cache line according to the first data processing instruction includes: Reading M cache lines in the first cache group from the M-way data memory respectively; Obtaining the cache tag of the first cache line determined based on the tag matching result; Determining the first cache line from the M cache lines read according to the cache tag of the first cache line; Reading the first data from the first cache line according to the first offset field; Performing a consistency check on the first data and the third data to obtain a check result; Returning the check result to the detection and issue unit that sent the first data processing instruction.
6. The method according to any one of claims 2-5, characterized in that, The method further includes: If the cache tag matching result indicates that there is no cache line in the first cache group whose cache tag matches the first tag field, determine a cache line to be replaced from the first cache group; Generate a data replacement instruction for the cache line to be replaced; Send the data replacement instruction to a second cache pipeline to schedule the second cache pipeline to write the data in the cache line to be replaced into a cache data source; the second cache pipeline is any one of the P cache pipelines for data storage access; Obtain a fourth data including the first data from the cache data source based on the first access address; Backfill the fourth data into the cache line to be replaced in the data array, and determine the cache line to be replaced after backfill processing as the first cache line for storing the first data.
7. The method according to claim 6, characterized in that, The cache processing system further includes a write-out queue; the method further includes: Obtain the data replacement instruction in the second cache pipeline; Read the cache line to be replaced from the data array based on a second tag array connected to the second cache pipeline; Write the data in the cache line to be replaced into the write-out queue; Write the data in the cache line to be replaced in the write-out queue into the cache data source.
8. The method according to claim 6, wherein The method further includes: Generate a replacement cache tag based on the tag field in the first access address; Generate a tag update signal for the cache line to be replaced, where the tag update signal includes the replacement cache tag; Based on the tag update signal, update the cache tag of the cache line to be replaced in each tag array to the replacement cache tag.
9. The method according to claim 1, characterized in that, The first cache pipeline is a data storage access pipeline; the cache processing system further includes a store buffer, a detection and emission unit, a preset register, and a preset scheduler; the store buffer is used to send a data storage instruction to the preset scheduler, the detection and emission unit is used to send a data detection instruction to the preset scheduler, and the preset register is used to send a data replacement instruction to the preset scheduler; the method further includes: Determine the sending order of S data processing instructions obtained by the preset scheduler according to priorities, where the S data processing instructions include one or more of the data storage instruction, the data replacement instruction, and the data detection instruction, and the data storage instruction, the data replacement instruction, and the data detection instruction each have their respective corresponding priorities; S is a positive integer; Select the first data processing instruction from the S data processing instructions according to the sending order, and send the first data processing instruction to the first cache pipeline.
10. The method according to claim 9, characterized in that The method further includes: If the store buffer obtains multiple data storage requests, and the tag fields and index fields in the access addresses carried by the multiple data storage instructions are the same, merge the multiple data storage instructions into a new data storage instruction; Send the new data storage instruction to the preset scheduler through the store buffer.
11. The method according to claim 1, characterized in that, The data array includes M data memories, each of the data memories includes one or more data banks, and each of the data banks is used to store one or more cache lines; Reading the first data from the first cache line includes: If the first data bank includes the first cache line and no other cache pipelines access the first data bank, reading the first data from the first cache line in the first data bank; Wherein, the first data bank is any data bank in the data array, and the other cache pipelines are cache pipelines other than the first cache pipeline among the P cache pipelines.
12. A cache processing device, characterized in that, A cache processing system is provided in the device, and the cache processing system includes a data array, P cache pipelines and P tag arrays, where P is a positive integer; one of the cache pipelines is connected to one of the tag arrays, and the P cache pipelines are all connected to the data array; The data array is used to store data in the form of cache lines; The tag array is used to store cache tags of the cache lines; the device includes: An acquisition module, configured to acquire a first data processing instruction in the first cache pipeline; the first data processing instruction carries a first access address, and the first access address is used to locate the first data indicated by the first data processing instruction; A processing module, configured to determine a first cache line for storing the first data from the data array based on the first tag array connected to the first cache pipeline; The processing module is further configured to perform a data processing operation on the first data in the first cache line according to the first data processing instruction; Wherein, the first cache pipeline is any one of the P cache pipelines; data processing instructions in each of the cache pipelines are executed in parallel based on their respective connected tag arrays.
13. The device according to claim 12, characterized in that, The data array includes M data memories, each of the data memories includes N cache lines; the data array is divided into N cache groups, each of the cache groups corresponds to a cache index, and each of the cache groups includes M cache lines, and the M cache lines belong to different data memories respectively; both M and N are positive integers; the first access address includes a first index field and a first tag field; The processing module is further configured to call a tag reading unit in the first cache pipeline to determine a first cache group whose cache index matches the first index field from the N cache groups, and read cache tags of each cache line in the first cache group from the first tag array to obtain M cache tags; The processing module is further configured to call a tag comparison unit in the first cache pipeline to perform a matching process on the read M cache tags and the first tag field to obtain a tag matching result. If the tag matching result indicates that there is a cache tag of a target cache line among the read M cache tags that matches the first tag field, the target cache line is determined as the first cache line for storing the first data.
14. The device according to claim 13, wherein The first data processing instruction is a data loading instruction. Among the P cache pipelines, there are Q1 data loading access pipelines, where Q1 is a positive integer and Q1 is less than or equal to P. The cache processing system further includes Q1 loading issue units. One loading issue unit is connected to one data loading access pipeline. Each loading issue unit is used to send a data loading instruction to the data loading access pipeline it is connected to. The first cache pipeline is any one of the Q1 data loading access pipelines, and the first cache pipeline is connected to the first loading issue unit among the Q1 loading issue units. The first access address further includes a first offset field. The processing module is further configured to call the way prediction unit in the first cache pipeline to predict the first data memory storing the first data from the M-way data memory. The processing module is further configured to call the first data reading unit in the first cache pipeline to read a second cache line from the first data memory, and the second cache line belongs to the first cache group. The processing module is further configured to call the first data selection unit in the first cache pipeline to obtain the cache tag of the first cache line determined based on the tag matching result. If the cache tag of the first cache line is the same as the cache tag of the second cache line, then according to the first offset field, read the first data from the second cache line, and send the first data to the first loading issue unit for data loading processing.
15. The device according to claim 13, characterized in that, The first data processing instruction is a data storing instruction, and the data storing instruction further carries a second data to be stored. Among the P cache pipelines, there are Q2 data storing access pipelines, where Q2 is a positive integer and Q2 is less than or equal to P. The first cache pipeline is any one of the Q2 data storing access pipelines. The first access address further includes a first offset field. The processing module is further configured to call the second data reading unit in the first cache pipeline to respectively read M cache lines in the first cache group from the M-way data memory, and send the M read cache lines to the second data selection unit in the first cache pipeline. The processing module is further configured to call the second data selection unit in the first cache pipeline to obtain the cache tag of the first cache line determined based on the tag matching result, determine the first cache line from the M read cache lines according to the cache tag of the first cache line, and according to the first offset field, perform a merging process on the determined first cache line and the second data, so as to replace the first data in the first cache line with the second data to obtain merged data. The processing module is further configured to call the data writing unit in the first cache pipeline to write the merged data into the first cache line in the data array.
16. The device according to claim 13, wherein The first data processing instruction is a data detection instruction, and the third data to be detected is also carried in the data detection instruction; among the P cache pipelines, there are Q2 data storage access pipelines, where Q2 is a positive integer and Q2 is less than or equal to P; the first cache pipeline is any one of the Q2 data storage access pipelines; the first access address further includes a first offset field; the cache processing system further includes a detection emission unit for sending data detection instructions. The processing module is further configured to call the second data reading unit in the first cache pipeline to respectively read M cache lines in the first cache group from the M-way data memory. The processing module is further configured to call the second data selection unit in the first cache pipeline to obtain the cache tag of the first cache line determined based on the tag matching result, determine the first cache line from the M read cache lines according to the cache tag of the first cache line, and read the first data from the first cache line according to the first offset field. The processing module is further configured to call the data comparison unit in the first cache pipeline to perform a consistency check on the first data and the third data to obtain a check result, and return the check result to the detection emission unit that sent the first data processing instruction.
17. The device according to any one of claims 13-16, wherein The processing module is further configured to, if the tag matching result indicates that there is no cache line in the first cache group whose cache tag matches the first tag field, call the replacement selection unit in the first cache pipeline to determine a cache line to be replaced from the first cache group. The processing module is further configured to call a preset register in the cache processing system to generate a data replacement instruction for the cache line to be replaced, send the data replacement instruction to the second cache pipeline to schedule the second cache pipeline to write the data in the cache line to be replaced into the cache data source, obtain a fourth data including the first data from the cache data source, backfill the fourth data into the cache line to be replaced in the data array, and determine the cache line to be replaced after the backfill process as the first cache line for storing the first data, where the second cache pipeline is any one of the P data storage access pipelines in the P cache pipelines.
18. A chip, characterized in that, Comprising: A processor and a cache processing device; A cache processing system is provided in the device, and the cache processing system includes a data array, P cache pipelines, and P tag arrays, where P is a positive integer; one of the cache pipelines is connected to one of the tag arrays, and all the P cache pipelines are connected to the data array; the data array is used to store data in the form of cache lines. The tag array is used to store the cache tags of the cache lines. The processor is configured to send data processing instructions to the cache processing device. The cache processing device is configured to execute the cache processing method according to any one of claims 1-11.
19. A computer device, characterized in that, The computer device includes: a computer-readable storage medium and a chip; The chip is adapted to implement a computer program; The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the chip to perform the cache processing method according to any one of claims 1-11.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the chip to perform the cache processing method according to any one of claims 1-11.