Processor pre-fetch training method, processing device, processor and computing equipment
By translating the instructions into N micro-instructions and triggering the prefetcher for training operations, the performance loss caused by CPU memory access delay is solved, and the effect of improving prefetch efficiency and accuracy is achieved.
Patent Information
- Application Number
- CN202111671282.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-31
AI Technical Summary
Modern high-performance CPUs have performance losses due to memory access latency, and existing multi-level cache architectures and prefetchers are difficult to effectively improve prefetch efficiency and accuracy.
By translating the instructions into N micro-instructions and triggering the prefetcher for training based on these micro-instructions, the prefetcher performs prefetching training based on the data width, address and attributes of the training operation.
It improves prefetching efficiency and accuracy, simplifies the write resources of the address input queue, avoids resource waste, and improves the operating efficiency of the processor.
Smart Images

Figure CN114358180B_ABST
Abstract
Description
Technical Field
[0001] Some embodiments of the present disclosure relate to the technical field of processors, and in particular, to a pre-fetch training method, a processing device, a processor, and a computing device for a processor. Background Art
[0002] In the central processing unit (CPU) architecture, program instructions and data are generally stored in memory such as dynamic random access memory (DRAM). Usually, the operating frequency of the CPU core is higher than the operating frequency of the memory. Therefore, the CPU needs to wait for hundreds of CPU clock cycles to directly obtain data from the memory, which will cause the CPU to idle due to the inability to continue to process related instructions or data, resulting in performance loss. Therefore, modern high-performance CPUs are usually equipped with a multi-level cache architecture to store recently accessed data. Furthermore, for the multi-level cache architecture, a prefetcher can also be used to identify the regularity of the CPU accessing data, and pre-fetch the data that may be accessed into one of the first-level caches of the multi-level cache architecture in advance, so that the CPU can read data from the cache more quickly. Summary of the invention
[0003] Some embodiments of the present disclosure provide a pre-fetch training method for a processor, a processing device, a processor, and a computing device, for improving pre-fetch efficiency and pre-fetch accuracy.
[0004] According to one aspect of the present disclosure, a prefetch training method for a processor is provided, comprising: according to an instruction translation rule, translating an instruction into N microinstructions, wherein N is a positive integer greater than 1; and triggering a prefetcher to perform one or more training operations based on the N microinstructions, wherein for each of the one or more training operations, the prefetcher performs prefetch training based on the data width, address, and attributes of the training operation.
[0005] According to some embodiments of the present disclosure, the parameters of the training operation include one or more of the following parameters: data width, address, and attribute.
[0006] According to some embodiments of the present disclosure, an instruction translation rule indicates that data processed by N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, wherein triggering a prefetcher based on the N microinstructions to perform one or more training operations includes: triggering a prefetcher based on the N microinstructions to perform a training operation, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the merged attribute obtained by processing the N microinstructions through the microinstruction processing pipeline.
[0007] According to some embodiments of the present disclosure, a microinstruction whose address is equal to the address of an instruction among N microinstructions is represented as a first microinstruction, and a microinstruction whose address is not equal to the address of the instruction among N microinstructions is represented as a second microinstruction. The prefetch training method also includes: for the first microinstruction, when there are attributes of the second microinstruction processed by the microinstruction processing pipeline in the address input queue for the prefetcher, the attributes of the first microinstruction processed by the microinstruction processing pipeline and the attributes of the second microinstruction processed by the microinstruction processing pipeline are merged and written into the address input queue as merged attributes; and when there are no attributes of the second microinstruction processed by the microinstruction processing pipeline in the address input queue for the prefetcher, the attributes of the first microinstruction processed by the microinstruction processing pipeline are written into the address input queue as merged attributes.
[0008] According to some embodiments of the present disclosure, the prefetch training method also includes: for the second microinstruction, when there is a merge attribute in the address input queue, merging the merge attribute with the attribute obtained by the second microinstruction through the microinstruction processing pipeline as an updated merge attribute, wherein the updated merge attribute is used for the prefetcher to perform prefetch training; and when there is no merge attribute in the address input queue, writing the attribute obtained by the second microinstruction through the microinstruction processing pipeline into an idle position in the address input queue.
[0009] According to some embodiments of the present disclosure, an instruction translation rule indicates that data processed by N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, wherein triggering a prefetcher based on the N microinstructions to perform one or more training operations includes: triggering a prefetcher based on the N microinstructions to perform a training operation, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by processing the microinstruction whose address is equal to the address of the instruction among the N microinstructions through the microinstruction processing pipeline.
[0010] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data width corresponding to each microinstruction in N microinstructions is equal to the data width corresponding to the instruction, and the address of each microinstruction is equal to the address of the instruction, wherein triggering the prefetcher to perform one or more training operations based on the N microinstructions includes: triggering the prefetcher to perform a training operation based on the N microinstructions, wherein the data width of the one training operation is equal to the data width corresponding to the instruction, the address of the one training operation is equal to the address of the instruction, and the attribute of the one training operation is equal to the attribute obtained by processing one of the N microinstructions through the microinstruction processing pipeline.
[0011] According to some embodiments of the present disclosure, an instruction translation rule indicates that the data processed by N microinstructions are discontinuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, wherein triggering a prefetcher based on the N microinstructions to perform one or more training operations includes: triggering a prefetcher based on the N microinstructions to perform N training operations, wherein the N training operations correspond one-to-one to the N microinstructions, for each of the N training operations, the data width of the training operation is equal to the data width corresponding to the corresponding microinstruction, the address of the training operation is equal to the address of the corresponding microinstruction, and the attribute of the training operation is equal to the attribute obtained by the corresponding microinstruction after being processed by the microinstruction processing pipeline.
[0012] According to some embodiments of the present disclosure, the pre-fetch training method further includes: adding tag information to N microinstructions, wherein the tag information is used to indicate a data width, an address, and an attribute of one or more training operations.
[0013] According to some embodiments of the present disclosure, the prefetch training method further includes: writing parameters of one or more training operations into an address input queue for the prefetcher, so that the prefetcher performs prefetch training based on the parameters in the address input queue.
[0014] According to another aspect of the present disclosure, a processing device for performing prefetch training is also provided, including: a translation unit, configured to translate instructions into N microinstructions according to instruction translation rules, wherein N is a positive integer greater than 1; and a processing unit, configured to trigger a prefetcher to perform one or more training operations based on the N microinstructions, wherein for each of the one or more training operations, the prefetcher performs prefetch training based on parameters of the training operation.
[0015] According to some embodiments of the present disclosure, the parameters of the training operation include one or more of the following parameters: data width, address, and attribute.
[0016] According to some embodiments of the present disclosure, an instruction translation rule indicates that data processed by N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, wherein the processing unit triggers the prefetcher based on the N microinstructions to perform one or more training operations, including: triggering the prefetcher based on the N microinstructions to perform a training operation, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the merged attribute obtained by processing the N microinstructions through the microinstruction processing pipeline.
[0017] According to some embodiments of the present disclosure, a microinstruction among N microinstructions whose address is equal to the address of an instruction is represented as a first microinstruction, and a microinstruction among N microinstructions whose address is not equal to the address of the instruction is represented as a second microinstruction, and the processing unit is further configured to: for the first microinstruction, in a case where there are attributes of the second microinstruction processed by the microinstruction processing pipeline in the address input queue for the prefetcher, merge the attributes of the first microinstruction processed by the microinstruction processing pipeline with the attributes of the second microinstruction processed by the microinstruction processing pipeline to write them into the address input queue as merged attributes; and, in a case where there are no attributes of the second microinstruction processed by the microinstruction processing pipeline in the address input queue for the prefetcher, write the attributes of the first microinstruction processed by the microinstruction processing pipeline as merged attributes into the address input queue.
[0018] According to some embodiments of the present disclosure, the processing unit is further configured to: for the second microinstruction, when there is a merge attribute in the address input queue, merge the merge attribute with the attribute obtained by the second microinstruction through the microinstruction processing pipeline as an updated merge attribute, wherein the updated merge attribute is used for the prefetcher to perform prefetch training; and when there is no merge attribute in the address input queue, write the attribute obtained by the second microinstruction through the microinstruction processing pipeline into an idle position in the address input queue.
[0019] According to some embodiments of the present disclosure, an instruction translation rule indicates that data processed by N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, wherein the processing unit triggers the prefetcher based on the N microinstructions to perform one or more training operations, including: triggering the prefetcher based on the N microinstructions to perform a training operation, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by processing the microinstruction whose address is equal to the address of the instruction among the N microinstructions through the microinstruction processing pipeline.
[0020] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data width corresponding to each microinstruction in N microinstructions is equal to the data width corresponding to the instruction, and the address of each microinstruction is equal to the address of the instruction, wherein the processing unit triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher to perform a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by processing one of the N microinstructions through the microinstruction processing pipeline.
[0021] According to some embodiments of the present disclosure, an instruction translation rule indicates that the data processed by N microinstructions are discontinuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, wherein the processing unit triggers the prefetcher based on the N microinstructions to perform one or more training operations, including: triggering the prefetcher based on the N microinstructions to perform N training operations, wherein the N training operations correspond one-to-one to the N microinstructions, for each of the N training operations, the data width of the training operation is equal to the data width corresponding to the corresponding microinstruction, the address of the training operation is equal to the address of the corresponding microinstruction, and the attribute of the training operation is equal to the attribute obtained by the corresponding microinstruction after being processed by the microinstruction processing pipeline.
[0022] According to some embodiments of the present disclosure, the translation unit is further configured to: add tag information to the N microinstructions, wherein the tag information is used to indicate the data width, address, and attributes of one or more training operations.
[0023] According to some embodiments of the present disclosure, the processing unit is further configured to: write parameters of one or more training operations into an address input queue for a prefetcher, so that the prefetcher performs prefetch training based on the parameters in the address input queue.
[0024] According to another aspect of the present disclosure, a processor is provided, comprising: a decoder configured to translate instructions into N microinstructions according to instruction translation rules, wherein N is a positive integer greater than 1; and a prefetcher configured to perform one or more training operations based on the N microinstructions, wherein for each of the one or more training operations, the prefetcher performs prefetch training based on parameters of the training operation.
[0025] According to another aspect of the present disclosure, a computing device is provided, including a processor; and a memory, wherein the memory stores a computer-readable code, and when the computer-readable code is executed by the processor, the processor pre-fetch training method as described above is executed.
[0026] Some embodiments of the present disclosure provide a prefetch training method, a processing device, a processor and a computing device for a processor. For N microinstructions obtained by instruction translation, a prefetcher is triggered to perform one or more training operations based on the N microinstructions, so that the prefetcher uses the initial instruction-level parameters for prefetch training, so as to avoid prefetch training based on multiple microinstruction parameters obtained after splitting, simplify the writing resources of the address input queue, avoid waste of resources of the address input queue, and simplify the prefetch training process of the prefetcher, which is more conducive to the prefetcher to discover data prefetching rules based on the initial instruction-level parameters, improve prefetching accuracy, and improve prefetching efficiency, thereby further improving the operating efficiency of the processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 A schematic diagram showing the processing flow of instructions in a processor;
[0029] Figure 2 A schematic diagram of microinstruction splitting is shown;
[0030] Figure 3 A schematic flow chart of a pre-fetch training method according to some embodiments of the present disclosure is shown;
[0031] Figure 4A A schematic diagram of an address input queue writing process according to some embodiments of the present disclosure is shown;
[0032] Figure 4B Another schematic diagram showing an address input queue writing process according to some embodiments of the present disclosure;
[0033] Figure 5A A flow chart of a writing mechanism of an address input queue according to some embodiments of the present disclosure is shown;
[0034] Figure 5B Another flow chart of a writing mechanism of an address input queue according to some embodiments of the present disclosure is shown;
[0035] Figure 6 A schematic block diagram of a processing device for performing pre-fetch training according to some embodiments of the present disclosure is shown;
[0036] Figure 7 A schematic block diagram of a processor according to some embodiments of the present disclosure is shown;
[0037] Figure 8 A schematic block diagram of a computing device according to some embodiments of the present disclosure is shown;
[0038] Fig. 9 A schematic diagram of a computing device architecture according to some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0040] In addition, as shown in the present disclosure and claims, unless the context clearly indicates an exception, the words "one", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The words "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Unless otherwise defined, all terms used herein have the same meaning as commonly understood by a person of ordinary skill in the art to which the present disclosure belongs.
[0041] The operating frequency of the processor (such as a CPU) core is much higher than the operating frequency of the DRAM memory (it is understood that a CPU may include one or more cores), so the processor core needs to wait for hundreds of processor clock cycles to directly obtain data and program instructions from the memory. In order to avoid the time delay caused by directly accessing data from the memory, the processor is usually configured with multiple levels of cache. Cache can refer to a high-speed memory with a faster data access speed, which exchanges data with the processor before the memory, and the cache setting enables the computer system to exert higher performance. It is understandable that the processor can be the above-mentioned CPU, or it can be other types of processors, such as a graphics processing unit (GPU), etc., which is not limited in this article.
[0042] In general, the order in which the processor reads data is first from the cache and then from the memory. When the processor needs to access a certain data, it first searches in the cache. If the data exists in the cache, it is immediately read and sent to the processor for processing. If the data does not exist in the cache, it is read from the memory with a relatively slow access rate and sent to the processor for processing. At the same time, the data block containing the data is transferred into the cache, so that in the subsequent stages, the processor reads the entire block of data from the cache without having to call the memory again. This reading mechanism makes the hit rate of the processor reading the cache relatively high, that is, the data that the processor wants to read next time is more likely to exist in the cache, and only a small part needs to be read from the memory. This greatly saves the time for the processor to read the memory directly, and also makes it basically unnecessary for the processor to wait when reading data.
[0043] Data prefetching is one of the key technologies to further improve the operating efficiency of the processor. Since the cache can only store the most recently accessed data, when the processor needs to read data that has never been accessed or data that has been replaced from the cache due to cache size limitations, the processor still needs to wait for dozens or even hundreds of clock cycles to read the data from the memory, which will cause performance loss. By analyzing the past (or historical) access patterns, the prefetcher can generate a suitable prefetch request address for the request address that triggers the prefetch, so as to prefetch the data that may be used into the cache in advance, thereby reducing the clock cycle of the processor waiting for data and improving the overall performance of the processor.
[0044] Figure 1 A schematic diagram of the processing flow of instructions in a processor is shown. Among them, the steps described about the processor can be understood as the steps involved in the core of the processor. An instruction can be a command for instructing the processor to perform a certain operation, which consists of a string of binary numbers. For example, the instruction can be a memory access (load / store) instruction, that is, an instruction for accessing or storing data. Hereinafter, the instruction will be described as a memory access instruction as a specific example according to some implementations of the present disclosure, and it can be understood that other types of instructions can also be used.
[0045] like Figure 1 As shown, the processor first fetches instructions and passes the fetched instructions to the decoding unit. As an example, the microinstructions herein may refer to microinstructions with memory access operations. Then, the decoding unit translates the instructions into microinstructions executed inside the processor. For example, when the instruction is an instruction with memory access operations, the microinstructions translated from the instruction may be microinstructions with memory access operations. Specifically, in this translation process, the instructions and microinstructions may be one-to-one or one-to-many. In other words, the decoding unit may translate an instruction into one microinstruction, or may translate one microinstruction into two or more microinstructions.
[0046] Then, the translated microinstructions will enter the processing queue after address calculation. The microinstructions in the processing queue are selected and enter the microinstruction processing pipeline to access data based on the cache. In order to further improve the data access rate of the processor and shorten the data processing time, it is proposed to use a prefetcher to find patterns from the address sequence of historical accesses, and based on this pattern, perform data prefetching, that is, predict the possible access address in the future in advance, and cache data based on the address. It can be understood that Figure 1 The processing queue shown in may refer to a memory access microinstruction processing queue, and Figure 1 The microinstruction processing pipeline shown in may be a memory access microinstruction processing pipeline.
[0047] The prefetcher is a structure independent of the processing queue and pipeline, and its input comes from the training information of the microinstructions being processed in the pipeline. The training information of the microinstructions will first be written into the address input queue, and then the prefetcher will take out the relevant information from the address input queue. On the one hand, it is used to train the prefetcher, that is, to discover the data access rules in the address sequence, and on the other hand, it is used to generate predicted addresses based on the discovered data access rules. The generated predicted addresses can be used to move data from the next level cache or memory to the current level data cache in advance. The process of the prefetcher discovering the data access rules in the address sequence can be called a training process (or a testing process), and the process of prefetching data based on the rules generated by the training process can be called a prefetch process. In the running stage of the processor, the training process and the prefetch process are carried out continuously and in parallel. Specifically, the above-mentioned training information may include the address of the processed data, the data width, and the cache hit feature. Specifically, the cache hit feature may be the result obtained after execution in the pipeline. As an example, the cache hit feature indicates whether the training hits the cache. For example, if the data corresponding to the request exists in the cache, it is represented as a hit (Hit), and if it does not exist in the cache, it is represented as a miss (Miss). Therefore, the cache hit feature can also be called a dynamic attribute.
[0048] For a certain instruction fetched by the instruction fetch unit, for example, a memory access instruction for accessing 512b (bit) data, the instruction may be divided into multiple transactions during the actual execution of the processor. In the decoding stage of the processor, the corresponding memory access instruction with a wider data width can be translated into multiple corresponding microinstructions with a narrower data width according to the actual execution width of the processor. This can also be called splitting the instruction into multiple microinstructions. This splitting process can also be expressed as microinstruction-level splitting. The actual execution width of the above-mentioned processor can refer to the data width that can be processed by the microinstruction processing pipeline.
[0049] As an example, the memory access instruction for accessing 512b data can be translated into two microinstructions for accessing 256b data respectively. That is, one instruction is translated into two microinstructions in the decoding stage. The two microinstructions obtained by translation will be written into the processing queue respectively, that is, occupying two items in the processing queue. In addition, in the process of the two microinstructions being selected from the processing queue and entering the microinstruction processing pipeline, due to the out-of-order execution mechanism of the processor (that is, the selection process is not performed in a fixed order), the timing of the selection between the two microinstructions is not related, and they may be selected to enter the pipeline continuously, or they may be selected at intervals, and then enter the pipeline for processing according to the order of selection.
[0050] Figure 2 Schematic diagram of microinstruction splitting is shown. Specifically, Figure 2 The figure shows the situation of splitting the instruction into two microinstructions and performing out-of-order execution between the microinstructions. Figure 2 In the example, the instruction with address 0x40 (hexadecimal number) corresponds to a data width of 512b (bit). If the microinstruction is a memory access microinstruction, it means that the 512b data starting from the storage location 0x40 needs to be accessed. Figure 2 In the example, in the decoding stage, the instruction is split into two microinstructions: microinstruction 1 with a first address and microinstruction 2 with a second address. Among them, the first address of microinstruction 1 is 0x40 and the corresponding data width is 256b, and the second address of microinstruction 2 is 0x60 and the corresponding data width is 256b. Similarly, the first address 0x40 and the second address 0x60 represent the starting address of the data to be accessed. Figure 2 The instructions at addresses 0x140, 0x240, and 0x340 shown in FIG. 1 may be split similarly to the instruction at address 0x40. Further, the two microinstructions split from the instruction at address 0x40 occupy two items in the processing queue, and the order in which they are picked may be out of order, such as Figure 2 As shown, due to the out-of-order execution mechanism, the order in which microinstructions 1 and 2 enter the pipeline is changed to microinstructions 2 and microinstructions 1. Figure 2 As shown in the figure, the microinstructions at addresses 0x40 / 0x60 and 0x240 / 0x260 are out of order. Then, after entering the pipeline, each microinstruction obtained by splitting will be written into the address input queue of the prefetcher, triggering the training and prediction of the prefetcher.
[0051] Due to the above microinstruction level splitting process, the original regular address sequence at the instruction level is disrupted after the splitting. Figure 2Taking the splitting scenario shown as an example, the initial instruction access addresses are: 0x40, 0x140, 0x240, 0x340, that is, the address distance between any two adjacent instructions is 0x100, which is a relatively simple fixed-step access rule. However, after the microinstruction level splitting, the address distance between the split microinstructions becomes: 0x20, 0xe0, 0x20, 0xe0, 02x20 and 0xe0, and the access rule becomes more complicated than the initial instruction access address rule. Furthermore, after the processor executes out of order, the address distance between the microinstructions that actually enter the pipeline becomes 0x20, 0x100, 0x20, 0x100, 0x20, 0x100, 0x20, which makes the access rule more complicated. Finally, in the related technology, each split microinstruction triggers pre-fetch training without considering the original instruction-level access rule. In other words, the above Figure 2 The information of the eight microinstructions with address intervals of 0x20, 0x100, 0x20, 0x100, 0x20, 0x100, and 0x20 shown in the figure will be written into the address input queue respectively and trigger the prefetcher to perform eight prefetch training processes.
[0052] The existence of the above-mentioned microinstruction-level splitting and out-of-order execution is not conducive to the prefetcher to discover the data access rules for accurate prefetching, which reduces the accuracy of the prefetcher. In addition, the information of multiple microinstructions obtained after splitting will be written into the address input queue respectively. For example, an initial instruction will occupy one item in the address input queue, and when it is split into two microinstructions, it will occupy two items in the address input queue, which wastes the resources of the address input queue.
[0053] Based on the above, some embodiments of the present disclosure provide a pre-fetch training method for a processor, which is used to improve pre-fetch efficiency and pre-fetch accuracy, avoid resource waste of the address input queue, and further improve the operating efficiency of the processor.
[0054] Figure 3 A schematic flow chart of a pre-fetch training method according to some embodiments of the present disclosure is shown below. Figure 3 To describe the implementation process of the pre-fetch training method according to an embodiment of the present disclosure.
[0055] First, in step S101, according to the instruction translation rule, the instruction is translated into N microinstructions. N is a positive integer greater than 1. For example, N can be equal to 2, that is, the instruction is split into 2 microinstructions, which corresponds to the above combination. Figure 2In the splitting situation described above, for example, N can also be greater than 2, that is, the instruction is split into more than 2 microinstructions. In the specific description below, N equals 2 as an example. It can be understood that the instruction splitting situation can be caused by the actual execution width of the above-mentioned processor, or it can be caused by other reasons, which is not limited here.
[0056] For this step S101, it can occur Figure 1 The instruction may be any type of instruction, and as an example, may be the above-mentioned instruction with memory access operation, and thus, the N microinstructions obtained by splitting the instruction may be microinstructions with memory access operation. The above-mentioned instruction translation rules will be described below in conjunction with a specific implementation method.
[0057] Then as Figure 3 As shown, in step S102, the prefetcher is triggered to perform one or more training operations based on N microinstructions, wherein for each of the one or more training operations, the prefetcher performs prefetch training based on the parameters of the training operation. Wherein, the multiple training operations here refer to two or more than two, for example, it can be N training operations. The number of triggered training operations will be described below in conjunction with examples. According to some embodiments of the present disclosure, the parameters of the training operation may include one or more of the following parameters: data width, address, and attribute. As some examples, the parameters of the training operation may include data width, address, and attribute. As other examples, the parameters of the training operation may also be other parameters, which are not limited here, and the category of the parameter can be determined according to, for example, the needs or scenarios for the prefetcher to perform prefetch training. The data width, address, and attributes of the training operation used for the prefetcher to perform prefetch training can be collectively referred to as training information, for example, and the training information is first written into the address input queue before being used for prefetch training, and then the prefetcher performs prefetch training based on the training information in the address input queue.
[0058] According to some embodiments of the present disclosure, the prefetch training method may further include: writing parameters of one or more training operations, such as data width, address, and attributes, into an address input queue for a prefetcher, so that the prefetcher performs prefetch training based on the data width, address, and attributes in the address input queue.
[0059] Specifically, in the pre-fetch training method according to some embodiments of the present disclosure, for the N microinstructions obtained by instruction splitting, the pre-fetcher will be triggered based on the N microinstructions to perform one or more training operations, and the training information corresponding to the training operation will be written into the address input queue for the pre-fetcher, so that the pre-fetcher performs pre-fetch training based on the above training information, thereby avoiding directly writing the training information of the microinstructions obtained after splitting into the address input queue respectively. It can be understood that in the embodiments according to the present disclosure, since the pre-fetcher is triggered to perform one or more training operations based on the information of the N microinstructions is comprehensively considered, the pre-fetch training can be performed with reference to the original instruction-level training information. Therefore, some embodiments of the present disclosure can restore the initial instruction-level training information, reduce the resource occupation of the address input queue, avoid the waste of resources of the address input queue, and simplify the pre-fetch training process of the pre-fetcher, which is more conducive to the pre-fetcher to discover the data pre-fetching law based on the initial instruction-level training information, improve the pre-fetching accuracy, and can also improve the pre-fetching efficiency and further improve the operating efficiency of the processor. Furthermore, the prefetch training method provided according to some embodiments of the present disclosure does not affect the existing prefetcher and processor architecture, and is convenient for wide application and implementation in existing processors.
[0060] The pre-fetch training method according to some embodiments of the present disclosure will be described below in conjunction with specific implementations.
[0061] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data processed by the above-mentioned N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction. In these embodiments, triggering the prefetcher to perform one or more training operations based on N microinstructions includes: triggering the prefetcher to perform a training operation based on N microinstructions. That is, the N microinstructions generated by the translation will only trigger one training operation instead of triggering N training operations, which is conducive to avoiding resource waste of the address input queue and simplifying the prefetch training process of the prefetcher. Specifically, the data width of the above-mentioned one training operation is equal to the data width corresponding to the instruction, the address of the above-mentioned one training operation is equal to the address of the instruction, and the attribute of the above-mentioned one training operation is equal to the merged attribute obtained by processing the N microinstructions through the microinstruction processing pipeline.
[0062] Specifically, in the above embodiment, an instruction is split into N microinstructions, each of the N microinstructions is used to access part of the data corresponding to the instruction, and the data processed by the N microinstructions are continuous, and the N microinstructions are put together to complete the function of the entire instruction. As an example, Figure 2The instruction with address 0x40 is split into two microinstructions, namely microinstruction 1 with address 0x40 and microinstruction 2 with address 0x60, wherein microinstruction 1 is used to access the first 256b of the 512b data corresponding to the original instruction (starting address 0x40), and microinstruction 2 is used to access the last 256b of the 512b data corresponding to the original instruction (starting address 0x60). In this case, the data processed by microinstruction 1 and microinstruction 2 are continuous (it can also be expressed as the addresses of microinstruction 1 and microinstruction 2 are continuous), and the sum of the data widths corresponding to microinstruction 1 and microinstruction 2 is equal to the data width corresponding to the instruction. That is, microinstruction 1 and microinstruction 2 together complete the function of the entire instruction.
[0063] In the above embodiment, the data width (512b) corresponding to the instruction, the address (0x40) of the instruction, and the merge attributes of N microinstructions processed by the microinstruction processing pipeline can be written into the address input queue for pre-fetch training.
[0064] As an example, the attribute obtained by the microinstruction processing pipeline after the microinstruction is processed may be the above-mentioned feature for indicating whether the cache is hit, for example, called a cache hit feature. Figure 1 As shown, the microinstructions in the processing queue are selected and enter the microinstruction processing pipeline. After pipeline processing, it can be determined whether the microinstruction being processed hits the cache. If the data corresponding to the microinstruction exists in the cache, it is represented as a hit (Hit), and if it does not exist in the cache, it is represented as a miss (Miss). In this example, since this cache hit feature is a feature obtained after pipeline processing, this attribute can also be called a dynamic attribute. Based on this, for Figure 2 In the example shown in , the merged attribute obtained by processing N microinstructions through the microinstruction processing pipeline may refer to the merged attribute obtained by merging the attribute obtained by processing microinstruction 1 (whose address is 0x40) through the microinstruction processing pipeline and the attribute obtained by processing microinstruction 2 (whose address is 0x60) through the microinstruction processing pipeline. For example, in the case where the attribute is a cache hit feature, the merged attribute is the merged cache hit feature obtained by merging the cache hit feature obtained by processing microinstruction 1 through the microinstruction processing pipeline and the cache hit feature obtained by processing microinstruction 2 through the microinstruction processing pipeline. According to an embodiment of the present disclosure, after pipeline processing, the above training information will be written into the address input queue for pre-fetch training.
[0065] Specifically, when the cache hit feature of at least one of the N microinstructions is a cache miss, the merge attribute is a cache miss, and when the cache hit feature of each of the N microinstructions is a cache hit, the merge attribute is a cache hit. For example, assuming that the cache hit feature of the above-mentioned microinstruction 1 is a hit and the cache hit feature of microinstruction 2 is a miss, it can be understood that the merged attribute after merging is a cache miss. Thus, the merge attribute of the instruction level before the split can be used for pre-fetch training, that is, the data access characteristics of the instructions before the split are retained.
[0066] As another example, the attribute obtained by the microinstruction after the microinstruction processing pipeline can also be an attribute for indicating whether to merge the Miss feature, for example, it is called the Miss feature attribute. This is because the granularity of each transmission between caches and between cache / DRAM is cacheline, for example, cacheline can have 64 bytes. If the addresses of multiple requests fall within the same 64-byte range, then the features of the subsequent Miss requests can be merged into the features of the first Miss request, and the first Miss request is responsible for retrieving this cacheline. Based on this, in the case where the attribute is the Miss feature attribute, the merged attribute is the Miss feature attribute obtained by microinstruction 1 after the microinstruction processing pipeline and the Miss feature attribute obtained by microinstruction 2 after the microinstruction processing pipeline. After merging, the merged Miss feature attribute is obtained. As another example, the above-mentioned attributes may also include both the above-mentioned cache hit feature and the Miss feature attribute. As other examples, the above-mentioned attributes may also be other information used for pre-fetching training, which will not be listed one by one here.
[0067] According to some embodiments of the present disclosure, in order to indicate training information such as data width, address, and attributes of one or more training operations, the pre-fetch training method may further include, for example, adding tag information to N microinstructions during the translation phase to indicate the above training information, thereby determining a writing mechanism for the address input queue based on the tag information. The implementation of the tag information will be described in detail below.
[0068] Figure 4A and Figure 4B FIG. 2 shows a schematic diagram of an address input queue writing process according to some embodiments of the present disclosure. Figure 4A , Figure 4B In the example shown, the microinstruction processing pipeline can continuously process microinstructions entering the pipeline.
[0069] The following combination Figure 4A and Figure 4B ,by Figure 2 The instruction with address 0x40 in is used as an example to describe the processing flow of the pipeline. Figure 2 In the example, the instruction at address 0x40 is first split into microinstruction 1 at address 0x40 and microinstruction 2 at address 0x60, and written into the processing queue respectively. Then, in the processing queue, due to the out-of-order execution mechanism, microinstruction 2 is first selected to enter the pipeline for processing, such as Figure 4A As shown, the mth stage pipeline performs data access based on microinstruction 2 and obtains attribute 2. After obtaining attribute 2, the writing mechanism of the address input queue can be determined according to the tag information related to microinstruction 2 to write the corresponding training information into the address input queue.
[0070] Then, if Figure 4B As shown, the processing of microinstruction 2 flows into the m+1th stage pipeline, and the mth stage pipeline performs data access based on microinstruction 1 and obtains attribute 1. Similarly, after obtaining attribute 1, the writing mechanism of the address input queue can be determined according to the tag information related to microinstruction 1 to write the corresponding training information into the address input queue. Figure 4B Schematically shows that microinstruction 1 is in the mth stage pipeline and microinstruction 2 is in the m+1th stage pipeline. It can be understood that in other cases, when microinstruction 1 is in the mth stage pipeline, microinstruction 2 may have flowed into the m+nth stage pipeline or has completed the pipeline processing process, that is, microinstruction 1 and microinstruction 2 are not adjacent in the pipeline. This article will take microinstruction 1 in the mth stage pipeline and microinstruction 2 in the m+1th stage pipeline as a specific example for description.
[0071] When the instruction translation rules indicate that the data processed by the above-mentioned N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, the microinstruction among the N microinstructions whose address is equal to the address of the instruction can be called the first microinstruction, and the microinstructions among the N microinstructions except the first microinstruction can be called the second microinstruction.
[0072] According to some embodiments of the present disclosure, the prefetch training method may further include: for a first microinstruction, when there are attributes of a second microinstruction processed by a microinstruction processing pipeline in the address input queue for the prefetcher, merging the attributes of the first microinstruction processed by the microinstruction processing pipeline with the attributes of the second microinstruction processed by the microinstruction processing pipeline to write them into the address input queue as merged attributes; and when there are no attributes of the second microinstruction processed by the microinstruction processing pipeline in the address input queue for the prefetcher, writing the attributes of the first microinstruction processed by the microinstruction processing pipeline as merged attributes into the address input queue.
[0073] Further, according to some embodiments of the present disclosure, the prefetch training method may also include: for the second microinstruction, when there is a merge attribute in the address input queue, merging the merge attribute with the attribute obtained by the second microinstruction through the microinstruction processing pipeline as an updated merge attribute, wherein the updated merge attribute is used for the prefetcher to perform prefetch training; and when there is no merge attribute in the address input queue, writing the attribute obtained by the microinstruction processing pipeline processing of the second microinstruction into an idle position in the address input queue.
[0074] According to some embodiments of the present disclosure, tag information can be added to the translated microinstructions during the decoding stage to indicate the training information for prefetch training. Specifically, the tag information can include method features for indicating the training method and data features for indicating the data width.
[0075] As an example, regarding the marking information in the above embodiment, for the first microinstruction, its method feature can be marked as "splicing training", indicating that the information of this microinstruction after pipeline processing is used to trigger pre-fetch training and can be spliced with information such as the second microinstruction. In addition, the data feature of the first microinstruction can be marked as equal to the data width of the initial instruction (for example, 512b). This allows training based on the initial instruction address and the data width of the instruction during training based on this first microinstruction. In the above example of splitting into microinstruction 1 and microinstruction 2, the address of microinstruction 1 is equal to the address of the instruction, that is, as the first microinstruction, thus, microinstruction 1 can be marked as "splicing training", and its data feature is equal to the data width of the initial instruction (512b).
[0076] For the second microinstruction, its method feature can be marked as "spliced training", indicating that the information of these microinstructions after pipeline processing does not trigger training operations but can be used to merge with the information of the first microinstruction. Since the second microinstruction does not trigger pre-fetch training, its data feature can be marked as an arbitrary value or unmarked. In the above example of splitting into microinstruction 1 and microinstruction 2, microinstruction 2 can be used as the above-mentioned second microinstruction. Therefore, microinstruction 2 can be marked as "spliced training" and its data feature is an arbitrary value. Regarding the above implementation method of distinguishing and marking method features as "splicing training" and "spliced training", it will be combined below. Figure 5A Give a description.
[0077] Figure 5A A flow chart of a writing mechanism of an address input queue according to some embodiments of the present disclosure is shown. Figure 5A The writing mechanism of the address input queue shown corresponds to the case of the tag information described above.
[0078] After being processed by the pipeline, the properties of the microinstruction will be obtained. Then, the method characteristics in the tag information can be first determined. For example, for the microinstruction 2 that is first processed by the pipeline, its properties can be obtained, and its method characteristics can be determined to be "spliced training". Next, it can be determined whether there is a first microinstruction belonging to the same instruction as the current microinstruction in the address input queue, that is, for the microinstruction 2 marked as being spliced training, it can be determined whether there is training information of microinstruction 1 belonging to the same instruction as microinstruction 2 in the address input queue. If it is determined that there is, the properties of the current microinstruction can be merged with the properties of the first microinstruction and used as the updated merged properties. That is to say, first merge the properties of microinstruction 2 with the properties of microinstruction 1, and then use the merged properties as the updated training information. If it is determined that there is no training information of microinstruction 1 belonging to the same instruction as microinstruction 2 in the address input queue, the information of microinstruction 2 can be written into the idle position of the address input queue. It can be understood that although the information of this microinstruction 2 is written into the queue, it does not trigger the prefetch training process of the prefetcher, because its method characteristics are marked as "spliced training". In addition, if there is no free position in the address input queue, the information of microinstruction 2 may not be written into the address input queue to avoid occupying queue resources. Meanwhile, when the address input queue is full, the table entries marked as spliced training may be overwritten first.
[0079] Next, refer to Figure 5A , for the microinstruction 1 processed by the pipeline, its attributes can be obtained, and its method feature can be determined to be "splicing training". Next, it can be determined whether there is a second microinstruction belonging to the same instruction as the current microinstruction in the address input queue, that is, for the microinstruction 1 marked as splicing training, it can be determined whether there is information about microinstruction 2 belonging to the same instruction as microinstruction 1 in the address input queue. If it is determined to exist, the attributes of the current microinstruction can be merged with the attributes of the second microinstruction, and the merged attributes can be written, that is, the attributes of microinstruction 1 are first merged with the attributes of microinstruction 2, and then the merged attributes are used as the updated training information. In addition, after the attribute merging is performed, the information related to microinstruction 2 can also be deleted from the address input queue. If it is determined that there is no training information of microinstruction 2 belonging to the same instruction as microinstruction 1 in the address input queue, the information of microinstruction 1 can be directly written into the address input queue.
[0080] Based on the above combination Figure 5AAccording to the described implementation method, some embodiments of the present disclosure can add tag information to the split microinstructions in the decoding stage, so that after pipeline processing, the writing mechanism of the address input queue can be determined according to the tag information to write the instruction-level training information into the address input queue, thereby realizing pre-fetch training with the instruction-level training information, which helps to reduce the resource waste of the address input queue and improve the pre-fetch accuracy.
[0081] According to some embodiments of the present disclosure, when the instruction translation rule indicates that the data processed by N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, triggering the prefetcher to perform one or more training operations based on the N microinstructions includes: triggering the prefetcher to perform a training operation based on the N microinstructions. Specifically, the data width of the above-mentioned training operation is equal to the data width corresponding to the instruction, the address of the above-mentioned training operation is equal to the address of the instruction, and the attribute of the above-mentioned training operation is equal to the attribute obtained by the microinstruction whose address is equal to the address of the instruction among the N microinstructions after being processed by the microinstruction processing pipeline.
[0082] Combined with the above Figure 5A In the described embodiment, the attribute of the training operation is equal to the merged attribute obtained by processing N microinstructions through the microinstruction processing pipeline, and thus the method characteristics in the tag information are distinguished into "splicing training" and "splicing training" to obtain the merged attribute in the process of writing the address input queue. In comparison, in this part of the embodiment, the attribute of the training operation is equal to the attribute obtained by processing the microinstruction whose address is equal to the address of the instruction through the microinstruction processing pipeline among the N microinstructions, that is, the attribute obtained by processing the first microinstruction through the microinstruction processing pipeline, and there is no need to obtain the merged attribute. Based on this, in this part of the embodiment, the method characteristics in the tag information can be distinguished into "normal training" and "no training".
[0083] As an example, regarding the marking information in the above embodiment, for the first microinstruction (i.e., the microinstruction whose address is equal to the address of the instruction), its method feature can be marked as "normal training", indicating that the information of this microinstruction after pipeline processing can be directly used to trigger prefetch training. In addition, the data feature of the first microinstruction can be marked as equal to the data width of the initial instruction (e.g., 512b). This allows training to be performed based on the initial instruction address and the data width of the instruction during training based on this first microinstruction. In the above example of splitting into microinstruction 1 and microinstruction 2, the address of microinstruction 1 is equal to the address of the instruction, that is, as the first microinstruction, and thus, microinstruction 1 can be marked as "normal training" and its data feature is equal to the data width of the initial instruction (512b).
[0084] For the second microinstruction (i.e., the microinstruction whose address is not equal to the address of the instruction), its method feature can be marked as "no training", indicating that the information of these microinstructions after pipeline processing does not trigger training operations. Since the second microinstruction does not trigger pre-fetch training, its data feature can be marked as an arbitrary value or unmarked. In the above example of splitting into microinstruction 1 and microinstruction 2, microinstruction 2 can be used as the above-mentioned second microinstruction. Therefore, microinstruction 2 can be marked as "no training" and its data feature is an arbitrary value. The implementation method of distinguishing the above method features and marking them as "normal training" and "no training" will be discussed below in conjunction with Figure 5B Give a description.
[0085] Figure 5B A flow chart of a writing mechanism of an address input queue according to some embodiments of the present disclosure is shown. Figure 5B The writing mechanism of the address input queue shown corresponds to the implementation mode of distinguishing the method characteristics into "normal training" and "no training". Figure 5A As shown, after pipeline processing, the properties of the microinstructions will be obtained, and then the method features in the tag information can be first determined. For example, the properties of the microinstruction 2 first processed by the pipeline can be obtained, and then, it is determined that the method feature marked is "not trained", and the training information related to the microinstruction 2 is not written into the address input queue, thereby saving the writing resources of the address input queue.
[0086] For another example, for microinstruction 1 that enters the pipeline processing after microinstruction 2, after obtaining its attributes, it can be determined that its marked method feature is "normal training", then the data feature (512b), address (0x40) and attributes (for example, a cache hit feature indicating whether the cache is hit) of this microinstruction 1 are written into the address input queue as training information.
[0087] Based on the above implementation method, some embodiments of the present disclosure can add tag information to the split microinstructions in the decoding stage, so that after pipeline processing, the writing mechanism of the address input queue can be determined according to the tag information to write the instruction-level training information into the address input queue, thereby realizing pre-fetch training with the instruction-level training information, which helps to reduce the resource waste of the address input queue (because the information of microinstruction 2 is no longer written into the address input queue) and improve the pre-fetch accuracy.
[0088] According to some embodiments of the present disclosure, the instruction translation rule may also indicate that the data width corresponding to each microinstruction in the N microinstructions is equal to the data width corresponding to the instruction and the address of each microinstruction is equal to the address of the instruction. In these embodiments, triggering the prefetcher to perform one or more training operations based on the N microinstructions includes: triggering the prefetcher to perform a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by processing one of the N microinstructions through the microinstruction processing pipeline.
[0089] Specifically, in the above embodiment, an instruction is split into N microinstructions, each of the N microinstructions is used to access all the data corresponding to the instruction, and the addresses of the N microinstructions are all the addresses of the instruction. As an example, the instruction can be an instruction for data expansion, for example, the address of the instruction is 0x40, the corresponding data width is 8b, and it is used to expand 64bit data based on the 8b data. In this case, the instruction may be translated into two microinstructions in the decoding stage, each of which has an address of 0x40 and accesses the 8b data. After obtaining the data, it is expanded to 32b data respectively, so as to splice the two expanded 32b data to obtain the target 64b data. In such a case, since the address and data width accessed by each microinstruction are the same, it can be foreseen that its properties may also be the same.
[0090] In order to implement pre-fetch training based on the data width (8b) corresponding to the instruction, the address (0x40) of the instruction, and the attribute obtained by processing one of the N microinstructions through the microinstruction processing pipeline as training information through marking information, marking information can be added to the multiple microinstructions obtained by splitting. As an implementation method, the method feature of the first microinstruction obtained by splitting can be marked as "normal training", and the data feature of the microinstruction can be marked as the data width (8b) corresponding to the instruction. In addition, the method features of the other microinstructions in the N microinstructions except the first microinstruction can also be marked as "no training", indicating that the information of this microinstruction after pipeline processing is not used to trigger pre-fetch training, and its data feature can be marked as any value or not marked.
[0091] According to the above embodiment of the present disclosure, after pipeline processing, the above marking information is used to implement pre-fetch training based on instruction-level training information. Specifically, the processing flow of "normal training" and "no training" can be referred to. Figure 5B The described implementation method will not be repeated here.
[0092] According to some embodiments of the present disclosure, the instruction translation rules may also indicate that the data processed by N microinstructions are discontinuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction. In these embodiments, triggering the prefetcher to perform one or more training operations based on N microinstructions includes: triggering the prefetcher to perform N training operations based on N microinstructions, wherein the N training operations correspond one to one with the N microinstructions. That is, in this case, each microinstruction will be used to trigger a training operation respectively. For each of the N training operations, the data width of the training operation is equal to the data width corresponding to the corresponding microinstruction, the address of the training operation is equal to the address of the corresponding microinstruction, and the attribute of the training operation is equal to the attribute obtained by the corresponding microinstruction after being processed by the microinstruction processing pipeline.
[0093] Specifically, in the above embodiment, an instruction is split into N microinstructions, each of the N microinstructions is used to access part of the data corresponding to the instruction, and the addresses of the N microinstructions are not related, and the N microinstructions jointly complete the function of the entire instruction.
[0094] As an example, the instruction may be an instruction for data collection, for example, the instruction is used to access data from multiple addresses and aggregate the accessed data. Thus, in the decoding stage, the instruction may be translated into microinstructions for accessing data from each address respectively. In this case, since the address of each microinstruction and the accessed data are discontinuous, pre-fetch training may be performed for each microinstruction respectively. In other words, the training information of each microinstruction triggers the pre-fetcher.
[0095] In order to indicate the above training information through tag information, tag information can be added to the multiple microinstructions obtained by splitting. As an implementation method, the method characteristics of each microinstruction obtained by splitting can be marked as "normal training", and the data characteristics of the microinstruction can be marked as the data width corresponding to the microinstruction. Based on the above tag information, after one of the microinstructions is processed by the pipeline, it will be written into the address input queue based on the tag information and attributes to trigger pre-fetch training.
[0096] The scheme of prefetch training by considering the instruction-level training information before splitting can reduce the resource occupation of the address input queue, avoid waste of resources of the address input queue, and simplify the prefetch training process of the prefetcher, which is more conducive to the prefetcher to discover data prefetching rules based on the initial instruction-level parameters, improve prefetching accuracy, improve prefetching efficiency and further improve the operating efficiency of the processor.
[0097] By using the pre-fetch training method according to the embodiment of the present disclosure, the accuracy of pre-fetching can be improved, and the more accurate the data pre-fetching is, the more likely the data required by the processor can be included in the cache as much as possible, thereby avoiding the excessive waiting time required for direct access to the memory and avoiding the processor from idling. Therefore, by using the pre-fetch training method according to the embodiment of the present disclosure, the operating efficiency of the processor can be improved.
[0098] According to another aspect of the present disclosure, a processing device for performing pre-fetch training is also provided. Figure 6 A schematic block diagram of a processing device for performing prefetch training according to some embodiments of the present disclosure is shown. Using the processing device for performing prefetch training according to some embodiments of the present disclosure, the prefetch efficiency and prefetch accuracy of the prefetcher can be improved, so that the processor's demand for data prefetching can be flexibly responded to, and the operating efficiency of the processor can be further improved.
[0099] like Figure 6 As shown, the processing device 1000 for performing pre-fetch training may include a translation unit 1010 and a processing unit 1020. According to some embodiments of the present disclosure, the translation unit 1010 may be configured to: translate the instruction into N microinstructions according to the instruction translation rule. The above instruction may be any type of instruction, such as an instruction with a memory access operation. N is a positive integer greater than 1, for example, N may be equal to 2, that is, the instruction is translated into 2 microinstructions, which corresponds to the above combination Figure 2 In the described situation, for example, N can also be greater than 2, that is, the instruction is translated into more than 2 microinstructions. According to the processing device of some embodiments of the present disclosure, the translation unit 1010 translates the instruction into two or more microinstructions due to factors such as the actual execution width of the processor, and the improvement made to the training operation of the prefetcher, thus, N is defined here as an integer greater than 1. It can be understood that the translation unit can also translate the instruction into 1 microinstruction (for example, because the data width corresponding to the instruction is smaller than the actual execution width of the processor and is not split), in which case, the 1 microinstruction can be directly used to train the prefetcher.
[0100] The processing unit 1020 can be configured to trigger the prefetcher to perform one or more training operations based on N microinstructions, wherein, for each of the one or more training operations, the prefetcher performs prefetch training based on the parameters of the training operation. According to some embodiments of the present disclosure, the parameters of the training operation may include one or more of the following parameters: data width, address, and attribute. As some examples, the parameters of the training operation may include data width, address, and attribute. As other examples, the parameters of the training operation may also include other parameters, which are not limited here, and the category of the parameter can be determined according to, for example, the need or scenario for the prefetcher to perform prefetch training. In the processing device according to the embodiment of the present disclosure, the above instruction-level training information can be first written into the address input queue, and then the prefetcher can perform prefetch training based on the training information in the address input queue to avoid prefetch training based on multiple microinstruction information obtained after splitting, simplify the writing resources of the address input queue, avoid waste of resources of the address input queue, and simplify the prefetch training process of the prefetcher, which is more conducive to the prefetcher to discover data prefetching rules based on the initial instruction-level parameters, improve prefetching accuracy, and improve prefetching efficiency, thereby further improving the operating efficiency of the processing device. As an example, the above processing device can be implemented as a processor.
[0101] According to some embodiments of the present disclosure, the translation unit 1010 may also be configured to: add tag information to N microinstructions, wherein the tag information is used to indicate the data width, address, and attributes of one or more training operations.
[0102] According to some embodiments of the present disclosure, the processing unit 1020 may also be configured to: write parameters of one or more training operations, such as data width, address, and attributes, into an address input queue for a prefetcher, so that the prefetcher performs prefetch training based on the data width, address, and attributes in the address input queue.
[0103] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data processed by the N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction. The processing unit 1020 triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher to perform a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the merged attribute obtained by the N microinstructions processed by the microinstruction processing pipeline.
[0104] According to some embodiments of the present disclosure, a microinstruction whose address is equal to the address of the instruction among the N microinstructions is represented as a first microinstruction, and a microinstruction whose address is not equal to the address of the instruction among the N microinstructions is represented as a second microinstruction. The processing unit 1020 is also configured to: for the first microinstruction, in the case where there is an attribute obtained by the second microinstruction through the microinstruction processing pipeline in the address input queue for the prefetcher, merge the attribute obtained by the first microinstruction through the microinstruction processing pipeline with the attribute obtained by the second microinstruction through the microinstruction processing pipeline to write into the address input queue as a merged attribute; and in the case where there is no attribute obtained by the second microinstruction through the microinstruction processing pipeline in the address input queue for the prefetcher, write the attribute obtained by the first microinstruction through the microinstruction processing pipeline as a merged attribute into the address input queue.
[0105] According to some embodiments of the present disclosure, the processing unit 1020 is further configured to: for the second microinstruction, when there is a merge attribute in the address input queue, merge the merge attribute with the attribute obtained by the second microinstruction through the microinstruction processing pipeline as an updated merge attribute, wherein the updated merge attribute is used for the prefetcher to perform prefetch training; and when there is no merge attribute in the address input queue, write the attribute obtained by the second microinstruction through the microinstruction processing pipeline into an idle position in the address input queue.
[0106] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data processed by the N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction. The processing unit 1020 triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher to perform a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by the microinstruction processing pipeline processing of the microinstruction whose address is equal to the address of the instruction in the N microinstructions.
[0107] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data width corresponding to each microinstruction in the N microinstructions is equal to the data width corresponding to the instruction, and the address of each microinstruction is equal to the address of the instruction. The processing unit 1020 triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher to perform a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by processing one of the N microinstructions through the microinstruction processing pipeline.
[0108] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data processed by the N microinstructions are discontinuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction. The processing unit 1020 triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher to perform N training operations based on the N microinstructions, wherein the N training operations correspond to the N microinstructions one by one, for each of the N training operations, the data width of the training operation is equal to the data width corresponding to the corresponding microinstruction, the address of the training operation is equal to the address of the corresponding microinstruction, and the attribute of the training operation is equal to the attribute obtained by the corresponding microinstruction after being processed by the microinstruction processing pipeline.
[0109] The specific implementation process of the steps executed by the processing device 1000 for performing pre-fetch training can refer to the pre-fetch training method according to some embodiments of the present disclosure described above in combination with the drawings, and will not be repeated here.
[0110] By utilizing a processing device that performs prefetch training according to some embodiments of the present disclosure, it is possible to restore initial instruction-level training information, simplify the writing resources of the address input queue, avoid wasting resources of the address input queue, and simplify the prefetch training process of the prefetcher, which is more conducive to the prefetcher discovering data prefetching rules based on the initial instruction-level training information, improving prefetching accuracy, and also improving prefetching efficiency and further improving the operating efficiency of the processing device.
[0111] According to another aspect of the present disclosure, a processor is provided. Figure 7 A schematic block diagram of a processor according to some embodiments of the present disclosure is shown.
[0112] like Figure 7 As shown, the processor 2000 may include a decoder 2010 and a prefetcher 2020. According to some embodiments of the present disclosure, the decoder 2010 may be configured to translate the instruction into N microinstructions according to the instruction translation rule, wherein N is a positive integer greater than 1. The above instruction may be any type of instruction, such as an instruction with a memory access operation. N is a positive integer greater than 1, for example, N may be equal to 2, that is, the instruction is split into 2 microinstructions, which corresponds to the above combination Figure 2In the splitting situation described, for example, N can also be greater than 2, that is, the instruction is split into more than 2 microinstructions. According to the processor of some embodiments of the present disclosure, the decoder 2010 translates the instruction into two or more microinstructions due to factors such as the actual execution width of the processor, and the improvement made to the training operation of the prefetcher 2020, thus, N is defined here as an integer greater than 1. It can be understood that the decoder 2010 can also translate the instruction into 1 microinstruction (for example, because the data width corresponding to the instruction is less than the actual execution width of the processor and the split is not performed), in which case, the 1 microinstruction can be directly used to train the prefetcher 2020.
[0113] The prefetcher 2020 can be configured to perform one or more training operations based on N microinstructions, wherein, for each of the one or more training operations, the prefetcher performs prefetch training based on the parameters of the training operation. According to some embodiments of the present disclosure, the parameters of the training operation may include one or more of the following parameters: data width, address, and attribute. As some examples, the parameters of the training operation may include data width, address, and attribute. As other examples, the parameters of the training operation may also include other parameters, which are not limited here, and the category of the parameter can be determined according to, for example, the need or scenario for the prefetcher to perform prefetch training. It can be understood that the processor 2000 may also include other devices not shown, such as a memory, which are not listed here.
[0114] In the processor according to the embodiment of the present disclosure, the above-mentioned instruction-level training information can be first written into the address input queue, and then the prefetcher can perform prefetch training based on the training information in the address input queue to avoid prefetch training based on multiple microinstruction information obtained after splitting, simplify the writing resources of the address input queue, avoid waste of resources of the address input queue, and simplify the prefetch training process of the prefetcher, which is more conducive to the prefetcher to discover data prefetching rules based on the initial instruction-level parameters, improve prefetching accuracy, and improve prefetching efficiency, thereby further improving the operating efficiency of the processor.
[0115] According to some embodiments of the present disclosure, the decoder 2010 may also be configured to: add tag information to N microinstructions, wherein the tag information is used to indicate the data width, address, and attribute of one or more training operations. That is, the decoder 2010 in the embodiments of the present disclosure may be extended to implement adding tag information to N microinstructions. It is understandable that this function may also be implemented by other components in the processor, which is not limited here.
[0116] According to some embodiments of the present disclosure, the prefetcher 2020 may also be configured to: write the data width, address, and attributes of one or more training operations into the address input queue, so that the prefetcher performs prefetch training based on the data width, address, and attributes in the address input queue. That is, the prefetcher 2020 in the embodiments of the present disclosure may be extended to implement writing training information into the address input queue, and it is understandable that this function may also be implemented by other components in the processor, which is not limited here.
[0117] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data processed by N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction. Specifically, the prefetcher 2020 performs one or more training operations based on the N microinstructions, including: performing a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the merged attribute obtained by the N microinstructions processed by the microinstruction processing pipeline.
[0118] According to some embodiments of the present disclosure, a microinstruction whose address is equal to the address of the instruction among N microinstructions is represented as a first microinstruction, and a microinstruction whose address is not equal to the address of the instruction among N microinstructions is represented as a second microinstruction. The prefetcher 2020 can also be configured to: for the first microinstruction, in the case where there is an attribute obtained by the second microinstruction processed by the microinstruction processing pipeline in the address input queue for the prefetcher, the attribute obtained by the first microinstruction processed by the microinstruction processing pipeline and the attribute obtained by the second microinstruction processed by the microinstruction processing pipeline are merged to be written into the address input queue as a merged attribute; and in the case where there is no attribute obtained by the second microinstruction processed by the microinstruction processing pipeline in the address input queue for the prefetcher, the attribute obtained by the first microinstruction processed by the microinstruction processing pipeline is written into the address input queue as a merged attribute.
[0119] According to some embodiments of the present disclosure, the prefetcher 2020 is also configured to: for the second microinstruction, when there is a merge attribute in the address input queue, merge the merge attribute with the attribute obtained by the microinstruction processing pipeline of the second microinstruction as an updated merge attribute, wherein the updated merge attribute is used for prefetch training; and when there is no merge attribute in the address input queue, write the attribute obtained by the microinstruction processing pipeline of the second microinstruction into an idle position in the address input queue.
[0120] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data processed by the N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction. The prefetcher 2020 performs one or more training operations based on the N microinstructions, including: performing a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by the microinstruction processing pipeline processing of the microinstruction whose address is equal to the address of the instruction in the N microinstructions.
[0121] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data width corresponding to each microinstruction in the N microinstructions is equal to the data width corresponding to the instruction, and the address of each microinstruction is equal to the address of the instruction. The prefetcher 2020 performs one or more training operations based on the N microinstructions, including: performing a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by processing one of the N microinstructions through the microinstruction processing pipeline.
[0122] According to some embodiments of the present disclosure, the instruction translation rule indicates that the data processed by the N microinstructions is discontinuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction. The prefetcher 2020 performs one or more training operations based on the N microinstructions, including: performing N training operations based on the N microinstructions, wherein the N training operations correspond one-to-one to the N microinstructions, for each of the N training operations, the data width of the training operation is equal to the data width corresponding to the corresponding microinstruction, the address of the training operation is equal to the address of the corresponding microinstruction, and the attribute of the training operation is equal to the attribute obtained by the corresponding microinstruction after being processed by the microinstruction processing pipeline.
[0123] The specific implementation process of the steps executed by the processing device 1000 for performing pre-fetch training can refer to the pre-fetch training method according to some embodiments of the present disclosure described above in combination with the drawings, and will not be repeated here.
[0124] By utilizing the processor according to some embodiments of the present disclosure, it is possible to restore the initial instruction-level training information, simplify the writing resources of the address input queue, avoid wasting resources of the address input queue, and simplify the prefetch training process of the prefetcher, which is more conducive to the prefetcher discovering data prefetching rules based on the initial instruction-level training information, improving prefetching accuracy, and also improving prefetching efficiency and further improving the operating efficiency of the processor.
[0125] The processor 2000 according to the embodiment of the present disclosure can realize the combination of Figure 3 The described pre-fetch training method can thus improve pre-fetch efficiency and pre-fetch accuracy. The specific implementation process can be referred to above and will not be repeated here.
[0126] According to yet another aspect of the present disclosure, a computing device is provided. Figure 8 A schematic block diagram of a computing device according to some embodiments of the present disclosure is shown.
[0127] like Figure 8 As shown, the computing device 3000 may include a processor 3010 and a memory 3020. According to an embodiment of the present disclosure, the memory 3020 stores a computer-readable code, and when the computer-readable code is executed by the processor 3010, the pre-fetch training method described above may be executed.
[0128] The processor 3010 can perform various actions and processes according to the program stored in the memory 3020. Specifically, the processor 3010 can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be an X86 architecture or an ARM architecture, etc. The processor 3010 can implement or execute various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure.
[0129] The memory 3020 stores a computer executable instruction code, which is used to implement the pre-fetch training method according to an embodiment of the present disclosure when executed by the processor 3010. The memory 3020 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous connection dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0130] By executing the prefetch training method according to some embodiments of the present disclosure, the processor 3010 can use the initial instruction-level parameters to perform prefetch training on the prefetcher for the N microinstructions obtained by instruction translation, so as to avoid performing prefetch training based on the multiple microinstruction information obtained after splitting, simplify the writing resources of the address input queue, avoid waste of resources of the address input queue, and simplify the prefetch training process of the prefetcher, which is more conducive to the prefetcher to discover data prefetching rules based on the initial instruction-level parameters, improve prefetching accuracy, and improve prefetching efficiency, thereby further improving the operating efficiency of the processor.
[0131] The pre-fetch training method according to the embodiment of the present disclosure can also be used by Fig. 9 The computing device architecture 4000 shown in FIG. Fig. 9 As shown, the computing device 4000 may include a bus 4010, one or more CPUs (cores) 4020, a read-only memory (ROM) 4030, a random access memory (RAM) 4040, a communication port 4050 connected to a network, an input / output 4060, a hard disk 4070, etc. The storage device in the computing device architecture 4000, such as the ROM 4030 or the hard disk 4070, may store various data or files used for processing and / or communication of the pre-fetch training method provided in the present disclosure, as well as program instructions executed by the CPU. The computing device architecture 4000 may also include a user interface 4080 for receiving user input information and instructions. Of course, Fig. 9 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Fig. 9 One or more components of a computing device are shown.
[0132] In addition, according to other aspects of the present disclosure, a non-transitory computer-readable storage medium may be provided. Instructions may be stored on the computer-readable storage medium, and the instructions are, for example, computer-readable instructions. When the computer-readable instructions are executed by the processor, the pre-fetch training method described with reference to the above figures may be executed. Computer-readable storage media include, but are not limited to, for example, volatile memory and / or non-volatile memory. According to other aspects of the present disclosure, a computer program product may be provided, the computer program product including computer-readable instructions, which are stored in a computer-readable storage medium. The processor may read the computer-readable instructions from the computer-readable storage medium, and the processor executes the computer-readable instructions so as to execute the pre-fetch training method described in the above-mentioned various embodiments.
[0133] Those skilled in the art will appreciate that the contents disclosed in this disclosure may be subject to various modifications and improvements. For example, the various devices or components described above may be implemented by hardware, or by software, firmware, or a combination of some or all of the three.
[0134] In addition, although the present disclosure makes various references to certain units in the system according to embodiments of the present disclosure, any number of different units can be used and run on the client and / or server. The units are only illustrative, and different aspects of the system and method can use different units.
[0135] Flowcharts are used in this disclosure to illustrate the steps of the method according to the embodiments of the present disclosure. It should be understood that the preceding or following steps are not necessarily performed precisely in order. On the contrary, various steps may be processed in reverse order or simultaneously. At the same time, other operations may also be added to these processes.
[0136] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of a software functional module. The present disclosure is not limited to any particular form of combination of hardware and software.
[0137] Unless otherwise defined, all terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present disclosure belongs. It should also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an idealized or highly formal sense, unless explicitly defined as such herein.
[0138] The above is an explanation of the present disclosure and should not be considered as a limitation thereof. Although several exemplary embodiments of the present disclosure are described, it will be readily understood by those skilled in the art that many modifications may be made to the exemplary embodiments without departing from the novel teachings and advantages of the present disclosure. Therefore, all such modifications are intended to be included within the scope of the present disclosure as defined in the claims. It should be understood that the above is an explanation of the present disclosure and should not be considered to be limited to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The present disclosure is defined by the claims and their equivalents.
Claims
1. A pre-fetch training method for a processor, comprising: According to the instruction translation rule, the instruction is translated into N microinstructions, where N is a positive integer greater than 1; and Triggering a prefetcher to perform one or more training operations based on the N microinstructions, wherein for each of the one or more training operations, the prefetcher performs prefetch training based on parameters of the training operation, wherein the parameters of the training operation are used to characterize initial instruction-level training information, wherein, In the case where the instruction translation rule indicates that the data processed by the N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, the triggering of the prefetcher based on the N microinstructions to perform one or more training operations includes: triggering the prefetcher based on the N microinstructions to perform a training operation, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, the attribute of the training operation is equal to the merged attribute of the N microinstructions processed by the microinstruction processing pipeline, or the attribute of the training operation is equal to the attribute of the microinstruction whose address is equal to the address of the instruction among the N microinstructions processed by the microinstruction processing pipeline, or, In the case where the instruction translation rule indicates that the data width corresponding to each microinstruction in the N microinstructions is equal to the data width corresponding to the instruction, and the address of each microinstruction is equal to the address of the instruction, the triggering of the prefetcher based on the N microinstructions to perform one or more training operations includes: triggering the prefetcher based on the N microinstructions to perform a training operation, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by processing one of the N microinstructions through the microinstruction processing pipeline, or, When the instruction translation rule indicates that the data processed by the N microinstructions are discontinuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, the triggering of the prefetcher based on the N microinstructions to perform one or more training operations includes: triggering the prefetcher based on the N microinstructions to perform N training operations, wherein the N training operations correspond one-to-one to the N microinstructions, for each of the N training operations, the data width of the training operation is equal to the data width corresponding to the corresponding microinstruction, the address of the training operation is equal to the address of the corresponding microinstruction, and the attribute of the training operation is equal to the attribute obtained by the corresponding microinstruction after processing by the microinstruction processing pipeline.
2. The method according to claim 1, characterized in that The parameters of the training operation include one or more of the following parameters: data width, address, and attribute.
3. The method according to claim 1, characterized in that In a case where the attribute of the one training operation is equal to the merged attribute of the N microinstructions processed by the microinstruction processing pipeline, a microinstruction whose address is equal to the address of the instruction among the N microinstructions is represented as a first microinstruction, and a microinstruction whose address is not equal to the address of the instruction among the N microinstructions is represented as a second microinstruction, and the method further includes: For the first microinstruction, when the address input queue for the prefetcher has the attribute of the second microinstruction processed by the microinstruction processing pipeline, the attribute of the first microinstruction processed by the microinstruction processing pipeline is merged with the attribute of the second microinstruction processed by the microinstruction processing pipeline to write the merged attribute into the address input queue; and When the attribute of the second microinstruction processed by the microinstruction processing pipeline does not exist in the address input queue for the prefetcher, the attribute of the first microinstruction processed by the microinstruction processing pipeline is written into the address input queue as the merged attribute.
4. The method according to claim 3, characterized in that The method further comprises: For the second microinstruction, if the merge attribute exists in the address input queue, merge the merge attribute with the attribute of the second microinstruction obtained by the microinstruction processing pipeline to serve as an updated merge attribute, wherein the updated merge attribute is used for the prefetcher to perform prefetch training; and In the case that the merge attribute does not exist in the address input queue, the attribute obtained by processing the second microinstruction through the microinstruction processing pipeline is written into an idle position in the address input queue.
5. The method according to claim 2, characterized in that The method further comprises: Tag information is added to the N microinstructions, wherein the tag information is used to indicate data width, address, and attributes of the one or more training operations.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: The parameters of the one or more training operations are written into an address input queue for the prefetcher, so that the prefetcher performs prefetch training based on the parameters in the address input queue.
7. A processing device for performing pre-fetch training, comprising: A translation unit configured to translate the instruction into N microinstructions according to an instruction translation rule, wherein N is a positive integer greater than 1; and A processing unit is configured to trigger a prefetcher to perform one or more training operations based on the N microinstructions, wherein for each of the one or more training operations, the prefetcher performs prefetch training based on parameters of the training operation, wherein the parameters of the training operation are used to characterize initial instruction-level training information, wherein, In the case where the instruction translation rule indicates that the data processed by the N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, the processing unit triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher based on the N microinstructions to perform a training operation, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, the attribute of the training operation is equal to the merged attribute of the N microinstructions processed by the microinstruction processing pipeline, or the attribute of the training operation is equal to the attribute of the microinstruction whose address is equal to the address of the instruction among the N microinstructions processed by the microinstruction processing pipeline, or, In the case where the instruction translation rule indicates that the data width corresponding to each microinstruction in the N microinstructions is equal to the data width corresponding to the instruction, and the address of each microinstruction is equal to the address of the instruction, the processing unit triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher to perform a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by processing one of the N microinstructions through the microinstruction processing pipeline, or, When the instruction translation rule indicates that the data processed by the N microinstructions are discontinuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instructions, the processing unit triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: performing N training operations based on the N microinstructions triggering the prefetcher, wherein the N training operations correspond one-to-one to the N microinstructions, for each of the N training operations, the data width of the training operation is equal to the data width corresponding to the corresponding microinstruction, the address of the training operation is equal to the address of the corresponding microinstruction, and the attribute of the training operation is equal to the attribute obtained by the corresponding microinstruction after being processed by the microinstruction processing pipeline.
8. The processing device according to claim 7, characterized in that The parameters of the training operation include one or more of the following parameters: data width, address, and attribute.
9. The processing device according to claim 7, characterized in that In a case where the attribute of the one training operation is equal to the merged attribute of the N microinstructions processed by the microinstruction processing pipeline, a microinstruction whose address is equal to the address of the instruction among the N microinstructions is represented as a first microinstruction, and a microinstruction whose address is not equal to the address of the instruction among the N microinstructions is represented as a second microinstruction, and the processing unit is further configured to: For the first microinstruction, when the attribute of the second microinstruction processed by the microinstruction processing pipeline exists in the address input queue for the prefetcher, the attribute of the first microinstruction processed by the microinstruction processing pipeline is merged with the attribute of the second microinstruction processed by the microinstruction processing pipeline to be written into the address input queue as the merged attribute; as well as When the attribute of the second microinstruction processed by the microinstruction processing pipeline does not exist in the address input queue for the prefetcher, the attribute of the first microinstruction processed by the microinstruction processing pipeline is written into the address input queue as the merged attribute.
10. The processing device according to claim 9, characterized in that The processing unit is further configured to: For the second microinstruction, if the merge attribute exists in the address input queue, merge the merge attribute with the attribute of the second microinstruction obtained by the microinstruction processing pipeline to serve as an updated merge attribute, wherein the updated merge attribute is used for the prefetcher to perform prefetch training; and In the case that the merge attribute does not exist in the address input queue, the attribute obtained by processing the second microinstruction through the microinstruction processing pipeline is written into an idle position in the address input queue.
11. The processing device according to claim 8, characterized in that The translation unit is also configured to: Tag information is added to the N microinstructions, wherein the tag information is used to indicate data width, address, and attributes of the one or more training operations.
12. The processing device according to any one of claims 7 to 11, characterized in that: The processing unit is further configured to: The parameters of the one or more training operations are written into an address input queue for the prefetcher, so that the prefetcher performs prefetch training based on the parameters in the address input queue.
13. A processor, comprising: A decoder configured to translate the instruction into N microinstructions according to an instruction translation rule, wherein N is a positive integer greater than 1; and A prefetcher configured to perform one or more training operations based on the N microinstructions, wherein for each of the one or more training operations, the prefetcher performs prefetch training based on parameters of the training operation, wherein the parameters of the training operation are used to characterize initial instruction-level training information, wherein In the case where the instruction translation rule indicates that the data processed by the N microinstructions are continuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, the prefetcher triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher to perform a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, the attribute of the training operation is equal to the merged attribute of the N microinstructions processed by the microinstruction processing pipeline, or the attribute of the training operation is equal to the attribute of the microinstruction whose address is equal to the address of the instruction among the N microinstructions processed by the microinstruction processing pipeline, or, In a case where the instruction translation rule indicates that the data width corresponding to each microinstruction in the N microinstructions is equal to the data width corresponding to the instruction, and the address of each microinstruction is equal to the address of the instruction, the prefetcher triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher to perform a training operation based on the N microinstructions, wherein the data width of the training operation is equal to the data width corresponding to the instruction, the address of the training operation is equal to the address of the instruction, and the attribute of the training operation is equal to the attribute obtained by processing one of the N microinstructions through the microinstruction processing pipeline, or, When the instruction translation rule indicates that the data processed by the N microinstructions are discontinuous and the sum of the data widths corresponding to the N microinstructions is equal to the data width corresponding to the instruction, the prefetcher triggers the prefetcher to perform one or more training operations based on the N microinstructions, including: triggering the prefetcher to perform N training operations based on the N microinstructions, wherein the N training operations correspond one-to-one to the N microinstructions, for each of the N training operations, the data width of the training operation is equal to the data width corresponding to the corresponding microinstruction, the address of the training operation is equal to the address of the corresponding microinstruction, and the attribute of the training operation is equal to the attribute obtained by the corresponding microinstruction after processing by the microinstruction processing pipeline.
14. A computing device comprising: processor; and A memory, wherein the memory stores a computer-readable code, and when the computer-readable code is executed by the processor, the pre-fetch training method of the processor according to any one of claims 1 to 6 is executed.
Citation Information
Patent Citations
Instruction execution method and instruction execution device
CN111857826A
Data prefetching method and data processing device
CN112527395A