A CPU instruction processing method and related device
By updating the storage sequence in the superscalar CPU, filtering the storage instructions using the serial number and access address of the storage instructions, the speculation error problem caused by index limitation is solved, and the performance and efficiency of the CPU are improved.
Patent Information
- Application Number
- CN202410927066.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-11
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-07-11
AI Technical Summary
In hyperscalar CPUs, due to hardware limitations, the number and form of indexes are limited, resulting in inferred loading errors, resulting in performance losses.
By updating the storage sequence of the loading instruction based on the sequence number and access address of the storage instruction, filtering the storage instructions entering the storage sequence, ensuring the validity of the storage sequence and reducing invalid conflict relationships.
Improves the accuracy of CPU speculative loading, reduces performance losses caused by speculative errors, and reduces circuit area and power consumption.
Smart Images

Figure CN118939319B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and particularly to a CPU instruction processing method and related devices. Background Art
[0002] In the design of superscalar Central Processing Units (CPUs), an out-of-order execution method is often adopted, that is, the actual instruction dispatch and issue order do not follow the program order.
[0003] This is because if the CPU strictly follows the program order to execute instructions, then a load instruction that appears after several store instructions can only be issued after these several store instructions are all completed, which will cause excessive latency. The out-of-order execution method needs to ensure as much as possible that the execution order of load instructions and store instructions accessing the same memory address is correct to avoid the CPU reading incorrect data.
[0004] In some CPUs, this problem is solved by means of speculative load, that is, the load instruction is speculatively distributed to the memory. If a previous store instruction is later determined to write to the same memory address as this load instruction, then this load instruction is rolled back and the correct data is used. However, the cost of rolling back is complex and expensive and should be avoided as much as possible.
[0005] That is to say, during the out-of-order execution process, the CPU not only needs to issue load instructions as early as possible, but also needs to ensure as much as possible that the execution order of this load instruction and the store instruction at the same memory location conforms to the read after write (RAW) constraint to avoid performance loss caused by RAW conflicts.
[0006] In current design solutions, the CPU can predict whether they will conflict (for example, the store instruction will write to the same memory address as the load instruction) based on whether the load instruction conflicts with previous store instructions in the past, so as to reduce the number of rollbacks. Among them, the program counter (PC) can be used to calculate the indexes of different instructions so that the CPU can determine whether the load instructions and store instructions in the table or set conflicted in the past. However, due to hardware limitations, the number and form of indexes are limited, and the indexes of different instructions may be the same, resulting in speculative errors and causing performance loss of the CPU. Summary of the Invention
[0007] Embodiments of the present invention provide a CPU instruction processing method and related devices, which can improve the accuracy of CPU speculative load and reduce the performance loss caused by CPU speculative errors.
[0008] In the first aspect of the embodiments of the present application, a method for processing (Central Processing Unit, CPU) instructions is provided, which is applied to the CPU. The method includes: obtaining a first load instruction and a first store instruction with a read after write (RAW) conflict; updating the storage sequence of the first load instruction based on the sequence number and / or access address of the first store instruction, where the storage sequence includes store instructions that have a RAW conflict with the first load instruction; and distributing the first load instruction based on the storage sequence.
[0009] In the embodiments of the present application, by using the sequence number of the first store instruction as the basis for updating the storage sequence, the number of store instructions in the storage sequence can be streamlined; by using the access address of the first store instruction as the basis for updating the storage sequence, it can better ensure that there is a RAW conflict between the store instructions in the storage sequence and the first load instruction. Therefore, through the method provided by the embodiments of the present application, it is possible to improve the quality of the storage sequence, improve the accuracy of speculative loading of the CPU, and reduce the performance loss caused by incorrect speculation of the CPU when the number and form of indexes are limited.
[0010] In a possible implementation, updating the storage sequence of the first load instruction based on the sequence number and / or access address of the first store instruction includes: if the sequence number of the first store instruction does not exist in a preset memory access dependency mapping table, adding the first store instruction to the storage sequence; if the sequence number of the first store instruction exists in the preset memory access dependency mapping table, keeping the storage sequence unchanged.
[0011] In the embodiments of the present application, through the sequence number of the first store instruction and the memory access dependency mapping table, it is possible to avoid repeatedly recording the dependency relationship between the first store instruction and the first load instruction in the storage sequence and the memory access dependency table, thereby saving circuit area; at the same time, reducing the number of times of updating the storage sequence can also make the power consumption of the CPU lower.
[0012] In a possible implementation, updating the storage sequence of the first load instruction based on the sequence number and / or access address of the first store instruction includes: if the sequence number of the first store instruction does not exist in a preset memory access dependency mapping table and the access address of the first store instruction is the same as the access address of the first load instruction, adding the first store instruction to the storage sequence; if the sequence number of the first store instruction exists in the preset memory access dependency mapping table, or the access address of the first store instruction is different from the access address of the first load instruction, keeping the storage sequence unchanged.
[0013] In the embodiment of the present application, the update logic of the store sequence is activated only when the first load instruction and the first store instruction access the same memory address. This ensures that the RAW conflicts maintained in the store sequence are truly valid conflicts, avoids invalid conflicts caused by the same PC index calculated by different store instructions, and thus reduces the performance loss caused by invalid conflicts.
[0014] In one possible implementation, after updating the storage sequence of the first load instruction, the method further includes: when adding the first storage instruction to the storage sequence, adding a counter corresponding to the first storage instruction, the initial count of the counter is zero; whenever a RAW conflict is detected between the first storage instruction and the first load instruction, adding one to the counter; if the counter is greater than or equal to a preset threshold, retaining the first storage instruction in the storage sequence until the storage sequence is cleared.
[0015] In an embodiment of the present application, by counting each storage instruction in the storage sequence and retaining the first storage instruction in the storage sequence of the first load instruction when a RAW conflict between the first load instruction and the first storage instruction is repeatedly detected, the storage instruction oscillation can be stopped as quickly as possible when the storage instruction oscillates in different storage sequences, thereby reducing the performance loss caused by the storage instruction oscillation.
[0016] In one possible implementation, the first load instruction comprises the oldest load instruction in the global history register.
[0017] In the embodiment of the present application, by giving priority to processing the oldest instruction in the global history register, instructions can be retired and abnormal situations can be handled as quickly as possible.
[0018] In one possible implementation, the processor includes a store set identifier table (SSIT) and a store instruction list (store list); the method further includes: when adding the first store instruction to the store sequence, writing the store set identifier (SSID) of the first load instruction into the SSIT; and writing the serial number and SSID of the first store instruction into the store instruction list.
[0019] In one possible implementation, the processor includes an SSIT, which includes multiple blocks, and the identifier of each block corresponds to the low-bit information of the program counter PC; after obtaining the first load instruction and the first storage instruction with a RAW conflict, the method also includes: based on the first low-bit information of the PC of the first load instruction, writing the storage sequence identification number SSID of the first load instruction into the block corresponding to the first low-bit information.
[0020] In the embodiments of the present application, by separately storing the information of the first load instruction with the low-order information of the PC, the invalid conflict relationship caused by the same index calculated by different load instructions can be avoided, thereby reducing the performance loss caused by the invalid conflict relationship.
[0021] In the second aspect of the embodiments of the present application, a processor is provided, and the processor is used to execute the CPU instruction processing method described in any possible implementation in the first aspect.
[0022] In the third aspect of the embodiments of the present application, an electronic device is provided, which is characterized in that the electronic device includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program instructions; the processor is used to call the computer program instructions to execute the CPU instruction processing method described in any possible implementation in the first aspect.
[0023] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which is characterized in that computer-executable instructions are stored in the computer-readable storage medium; wherein, when the computer-executable instructions are executed by the processor, the CPU instruction processing method described in any possible implementation in the first aspect is implemented.
[0024] It should be understood that the beneficial effects of the above aspects can be referred to each other. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a schematic flowchart of a CPU instruction processing method provided by the embodiments of the present application;
[0027] Figure 2 It is a schematic flowchart of the preprocessing before adding a storage instruction to a storage sequence provided by the embodiments of the present application;
[0028] Figure 3 It is a schematic logical diagram of instruction processing provided by the embodiments of the present application;
[0029] Figure 4 It is a structural block diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The embodiments of the present application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0031] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that shown or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] In the traditional speculative loading method, the Central Processing Unit (CPU) speculates based on the store set corresponding to each load instruction to determine whether to distribute the load instruction to the corresponding execution unit. Among them, the store set is used to record the store instructions that have had a read after write (RAW) conflict with the corresponding load instruction, and the CPU can speculate whether the execution of the current load instruction will cause a rollback based on historical conflict data.
[0033] Therefore, in order to improve the accuracy of speculative loading and reduce the performance loss caused by the rollback of load instructions, it is necessary to ensure the validity of the store instructions in the store set as much as possible, that is, the store instructions and the corresponding load instructions have actually had a RAW conflict.
[0034] Currently, the CPU manages the store set through two tables: the store set identifier table (SSIT) and the last fetched store table (LFST) of the most recently fetched store instruction.
[0035] Among them, the entry index of the SSIT includes the values calculated from the program counters (PCs) of the store instruction and the load instruction with RAW conflicts respectively, and the store instruction and the load instruction correspond to their respective indexes; the entry content includes the value calculated from the PC of the store instruction or the load instruction, and this value is used as the shared tag of the store instruction and the load instruction, which is called the store set identifier (SSID).
[0036] Among them, the entry index of the LFST is the SSID, and the entry content is the serial number (SN) of the store instruction in the corresponding store instruction-load instruction conflict pair.
[0037] Based on the above two tables, before distributing the load instruction, the CPU can find the corresponding SSID in the SSIT through the PC of the load instruction; if there is a corresponding SSID in the SSIT, it means that there is a store instruction conflicting with the load instruction, and the CPU needs to first complete the corresponding write-back operation based on the serial number of the store instruction, and wait for the entry of the store instruction to be removed from the SSIT before distributing the load instruction; if not in the SSID, the CPU can directly distribute the load instruction.
[0038] However, in the above traditional speculative load method, due to hardware limitations, the number of entries in the SSIT and LFST is limited, such as 64 or 128 entries. At this time, the index calculated by the CPU needs to fall within this number range; therefore, the indexes calculated by different PCs may be the same, and the wrong index will lead to the wrong RAW conflict relationship, resulting in instruction write-back and causing performance loss of the CPU.
[0039] To solve the above technical problems, the embodiments of the present application provide a CPU instruction processing method, which can improve the quality of the storage sequence, improve the accuracy of CPU speculative loading, and reduce the performance loss caused by speculative errors of the CPU when the quantity and form of the indexes are limited.
[0040] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the CPU instruction processing method provided by the embodiments of the present application. This method is applied to the CPU and specifically includes the following steps:
[0041] Step 102, the CPU obtains a first load instruction and a first store instruction with read after write (RAW) conflicts.
[0042] Among them, when the execution of a load instruction requires the execution result of a store instruction, there is a dependency relationship between the load instruction and the store instruction; when the load instruction is executed before the store instruction, the CPU will read incorrect data, resulting in a RAW conflict.
[0043] Among them, there is a RAW conflict between the first load instruction and the first store instruction. Specifically, the quantitative relationship between the first load instruction and the second store instruction can be one-to-one, one-to-many, many-to-one, or many-to-many. It is understood that instructions that do not have a RAW conflict with other instructions can be named using the second, third, or other words, such as the second load instruction, the third store instruction, etc., and this embodiment of the application is not specifically limited to this.
[0044] When the CPU determines that the access address of the skipped store instruction is the same as the access address of the pre-executed load instruction, that is, when the CPU determines that the pre-executed load instruction reads erroneous data, the CPU may determine that a RAW conflict exists between the load instruction and the store instruction. In this case, the CPU may obtain the load instruction as the first load instruction and the store instruction as the first store instruction.
[0045] Step 104: The CPU updates the storage sequence of the first load instruction based on the sequence number and / or access address of the first storage instruction.
[0046] The storage sequence of the first load instruction refers to a set of storage instructions that have a RAW conflict with the first load instruction. It can be understood that each load instruction has its own storage sequence.
[0047] To improve the prediction accuracy of speculative loads, it is necessary to ensure the validity of the store sequence. In other words, it is necessary to ensure that each store instruction in the store sequence has a RAW conflict with the corresponding load. Specifically, the access address can be used to determine whether the first load instruction and the first store instruction have a RAW conflict, avoiding false RAW conflict relationships caused by CPU misjudgment.
[0048] In one possible implementation, after obtaining the first load instruction and the first store instruction, if the access address of the first store instruction is the same as the access address of the first load instruction, then the CPU can add the first store instruction to the storage sequence of the first load instruction; if the access address of the first store instruction is different from the access address of the first load instruction, then the CPU can maintain the storage sequence unchanged.
[0049] In another possible implementation, if the sequence number of the first store instruction does not exist in the preset memory access dependency mapping table, the CPU can add the first store instruction to the store sequence; if the sequence number of the first store instruction exists in the preset memory access dependency mapping table, the CPU can keep the store sequence unchanged.
[0050] Among them, the CPU presets a memory access dependency mapping table for maintaining the correct execution order of memory access instructions. Specifically, the mapping table includes the read-write dependency relationships among all in-flight memory access instructions. For example, the execution of the read operation of instruction A depends on the execution result of the write operation of instruction B. It can be understood that the CPU can also perform speculative loading based on the dependency relationships recorded in the mapping table. That is to say, the dependency relationships recorded in the access dependency mapping table are similar in function to the RAW conflict relationships represented by the store sequence.
[0051] Therefore, before the CPU adds the first store instruction to the store sequence, it can traverse the sequence number of the first store instruction in the memory access dependency mapping table; when the sequence number of the first store instruction exists in the mapping table, it means that the mapping relationship of the first store instruction has been recorded. At this time, there is no need to repeatedly write the RAW conflict relationship with similar functions into the store sequence, thus saving circuit area; at the same time, reducing the number of times of updating the store sequence can also make the power consumption of the CPU lower.
[0052] On the contrary, when the sequence number of the first store instruction does not exist in the mapping table, the CPU can add the first store instruction to the store sequence, expand the sample database of speculative loading, and improve the prediction accuracy.
[0053] It can be understood that the above two possible implementations can also be combined with each other. In one possible implementation, the CPU can determine whether the sequence number of the first store instruction exists in the access dependency mapping table, and determine whether the access address of the first store instruction is the same as the access address of the first load instruction; if the sequence number of the first store instruction does not exist in the memory access dependency mapping table, and the access address of the first store instruction is the same as the access address of the first load instruction, then the CPU can add the first store instruction to the store sequence; if the sequence number of the first store instruction exists in the memory access dependency mapping table, or the access address of the first store instruction is different from the access address of the first load instruction, the store sequence remains unchanged.
[0054] Specifically, reference can be made to Figure 2 , Figure 2 which is a schematic diagram of a preprocessing process before adding a store instruction to a store sequence provided by an embodiment of this application.
[0055] After the CPU executes the first load instruction and finds the first store instruction that has a RAW conflict with the first load instruction among the instructions to be executed, the CPU can execute the following Figure 2 The pre-processing flow shown is to determine whether to add the first storage instruction to the storage sequence or keep the storage sequence unchanged.
[0056] exist Figure 2 In the specific example, the CPU first obtains the serial number of the first storage instruction and performs a hash calculation to obtain a serial number hash value A; then traverses the serial number hash value A in the access dependency mapping table and determines whether it hits, that is, whether the hash value A exists in the mapping table.
[0057] If the hit is found, the CPU can maintain the storage sequence unchanged. If the hit is not found, the CPU can further obtain the access addresses of the first storage instruction and the first load instruction and determine whether the access addresses are the same. If the access addresses are different, the CPU can maintain the storage sequence unchanged. If the access addresses are the same, the CPU can add the first storage instruction to the storage sequence.
[0058] In an embodiment of the present application, the first storage instruction added to the storage sequence is screened using the two conditions of serial number and access address, thereby ensuring that the RAW conflict maintained in the storage sequence is a truly valid conflict, reducing the performance loss caused by invalid conflict relationships; at the same time, it also avoids redundant data occupying the CPU's memory, allowing the CPU to have more speculative data samples within a limited storage space, thereby improving the accuracy of speculative loading.
[0059] After determining to add the first storage instruction to the storage sequence, the CPU can write data corresponding to the first storage instruction and the first load instruction in the preset SSIT and LFST, so that the first storage instruction can be found for speculation before the first load instruction is executed next time.
[0060] In one possible implementation, the SSIT includes multiple banks, and the identifier of each bank corresponds to the low-bit information of the PC; the CPU can write the SSID of the first load instruction into the bank corresponding to the first low-bit information based on the first low-bit information of the PC of the first load instruction.
[0061] The SSIT blocks are subtables of the SSIT. The table entry index of each block is the first index, and the table entry content is the SSID. For example, the SSIT includes eight blocks, each of which has an identifier corresponding to a different state of the lowest three bits of the PC. When the lowest three bits of the PC of the first storage instruction are all 0, the CPU may add the first storage instruction to the block identified as 0. It will be understood that this is merely an example and not a limitation. The SSIT may also include more blocks, such as 64, and each block's identifier may correspond to a different state of the lowest six bits of the PC.
[0062] The first low-order information refers to the low-order information of the PC of the first storage instruction.
[0063] The first index refers to an index calculated from the PC of the first load instruction or the PC of the first store instruction, and the first index is used to index the table entry corresponding to the first load instruction or the first store instruction in the SSIT.
[0064] Correspondingly, the second index is the SSID in the SSIT, which is used to index the table entry corresponding to the instruction pair of the first load instruction-the first store instruction having a RAW conflict in the LFST.
[0065] Optionally, the second index includes the SSID and the block identifier, for example, (block identifier) 0 + (SSID) 99.
[0066] In the embodiment of the present application, by separately storing the information of the first load instruction using the PC low-bit information, invalid conflict relationships caused by the same first index calculated by PCs of different load instructions can be avoided, thereby reducing performance losses caused by invalid conflict relationships.
[0067] Among them, since the hardware limits the table size of SSIT and LFST, the calculated first index and second index will fall within a certain range, which may cause incorrect RAW conflict relationships to be recorded between unrelated load instructions and storage instructions; therefore, in some solutions, the CPU may introduce the rule that "the same storage instruction can only exist in one storage sequence at the same time" to avoid incorrect RAW conflict relationships.
[0068] However, when a store instruction has a RAW conflict with multiple load instructions, the store instruction may oscillate in the store sequence of the multiple load instructions, causing both load instructions to fail to execute normally.
[0069] For example, the storage sequence for load instruction 1 includes store instruction A, while the storage sequence for load instruction 2 is empty. When a RAW conflict occurs between load instruction 2 and store instruction A, store instruction A is added to the storage sequence for load instruction 2, and store instruction A is removed from the storage sequence for load instruction 1, leaving it empty. Obviously, in this example, when the CPU executes load instructions 1 and 2, store instruction A may oscillate between the two storage sequences, and the CPU will alternately retrieve and remove store instruction A from the two storage sequences. This will cause the execution of load instructions 1 and 2 to always be in a state of unspecified or incorrect speculation, significantly impacting CPU performance.
[0070] Therefore, in a possible implementation proposed in the present application, when the first storage instruction is added to the storage sequence, a counter corresponding to the first storage instruction is added, and the initial count of the counter is zero; whenever a RAW conflict is detected between the first storage instruction and the first load instruction, the counter is incremented by one; if the counter is greater than or equal to a preset threshold, the first storage instruction is retained in the storage sequence until the storage sequence is cleared.
[0071] Exemplarily, the preset threshold may be 3.
[0072] Among them, when the CPU is powered on again or after a certain number of clock cycles during operation, the CPU will clear the storage sequence of some or all load instructions to avoid speculation errors caused by outdated RAW conflict relationships.
[0073] In an embodiment of the present application, by counting the RAW conflicts between the first load instruction and the first storage instruction that occur multiple times, and when the count value is greater than or equal to a preset threshold, the first storage instruction is allowed to reside in the storage sequence of the first load instruction, thereby stopping the storage instruction oscillation as quickly as possible and reducing the performance loss caused by the storage instruction oscillation.
[0074] In one possible implementation, the first load instruction comprises the oldest load instruction in the global history register.
[0075] The CPU executes instructions out of order during the execute phase, but commits them in order during the final commit phase of the pipeline. During the commit phase, instructions that were executed out of order are returned to the order specified in the program, primarily through the reorder buffer (ROB). Specifically, after an instruction completes, it enters the commit phase, at which point the instruction's status in the ROB is updated to completed. The ROB is essentially a first-in, first-out queue; completed instructions can only be retired in order after all previous instructions have completed.
[0076] Before the instruction is retired, its status is speculative. The CPU will retire the instruction only when it determines that there are no exceptional situations to handle for this instruction.
[0077] Among them, the global history register records the execution results of the previous N branch instructions, where N is greater than 1; when the instruction is retired, the instruction will be deleted from the global history register.
[0078] Therefore, in this possible implementation, the oldest load instruction in the global history register that has a RAW conflict relationship is used as the first load instruction, and this oldest load instruction is preferentially dispatched, executed, and retired, which can prevent too many instructions from piling up in the commit stage of the pipeline and affecting the CPU's judgment of the program execution situation.
[0079] Optionally, when the CPU adds the first store instruction to the store sequence, it writes the SSID of the first load instruction to the SSIT; it writes the serial number and SSID of the first store instruction to the store instruction list.
[0080] Among them, in the case of preferentially processing the oldest load instruction, the CPU can only allocate the table entry content in the SSIT for this oldest load instruction (which is also the first load instruction); in addition, the CPU also presets another table, which can be called the store instruction list, specifically used to store the store instructions that have a RAW conflict with the oldest load instruction, that is, after obtaining the first store instruction corresponding to the oldest load instruction, the CPU can write the relevant data of the first store instruction to the store instruction list.
[0081] Specifically, reference can be made to Figure 3 , Figure 3 which is a logical schematic diagram of an instruction processing provided by an embodiment of this application.
[0082] In Figure 3 In a specific example, the CPU first calculates the first index 99 and SSID 102 based on the PC of the first load instruction; then it writes the SSID 102 to the table entry corresponding to the first index 99 in BANK0 of the SSIT, and the SSID item corresponding to the first store instruction in the store instruction list; then it uses this SSID 102 as the second index, writes the SN B of the first store instruction to the LFST, and writes the SN B of the first store instruction to the store instruction list; finally, it sets the identifiers of the corresponding table entries in the SSIT and LFST to 1 to indicate that these table entries are enabled.
[0083] Through the specially preset store instruction list, the CPU can directly perform speculative dispatch of the oldest load instruction based on this store instruction list and dispatch the oldest load instruction as quickly as possible.
[0084] Step 106: The CPU distributes the first load instruction based on the storage sequence.
[0085] Among them, after the CPU determines that all the storage instructions in the storage sequence of the first load instruction have been executed, it can distribute the first load instruction to the corresponding memory for data reading.
[0086] In the embodiments of the present application, by using the serial number of the first storage instruction as the basis for updating the storage sequence, the number of storage instructions in the storage sequence can be streamlined; by using the access address of the first storage instruction as the basis for updating the storage sequence, it can better ensure that there is a RAW conflict between the storage instructions and the first load instruction in the storage sequence. Therefore, through the method provided in the embodiments of the present application, it is possible to improve the quality of the storage sequence, improve the accuracy of CPU speculative loading, and reduce the performance loss caused by CPU speculative errors under the condition that the quantity and form of indexes are limited.
[0087] In another embodiment, the present application also provides a processor, which is used to execute the method provided in the above embodiments.
[0088] In some other possible embodiments, the present application also provides an electronic device, specifically, reference can be made to Figure 4 , Figure 4 which is a schematic structural diagram of the electronic device provided in the embodiments of the present application. As Figure 4 shown, the electronic device 400 provided in the embodiments of the present application includes a processor 410 and a memory 420.
[0089] Among them, the processor 410 may include one or more processing cores. The processor 410 uses various interfaces and lines to connect various parts within the computing device 400, and by running or executing instructions, programs, code sets or instruction sets stored in the memory 420, and calling data stored in the memory 420, it executes the method provided in any one or more of the above embodiments. Optionally, the processor 410 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 410 may integrate one or a combination of a CPU, a graphics processing unit (GPU), and a modem. It can be understood that the above modem may not be integrated into the processor 410 and may be implemented separately through a communication chip.
[0090] The memory 420 may include a random access memory (RAM), or may also include a read-only memory (ROM). Optionally, the memory 420 includes a non-transitory computer-readable storage medium. The memory 420 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 420 may include a storage program area. Among them, the storage program area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), and instructions for implementing the method of the embodiment of the present application.
[0091] Among them, the processor 410 and the memory 420 are communicatively connected through a bus within the computing device 400. This bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0092] In another embodiment of the present application, there is also provided a computer-readable storage medium. Computer-executable instructions are stored in the computer-readable storage medium. When at least one processor of the device executes the computer-executable instructions, the device executes the above Figure 1 or Figure 2 the method flow described in any one of the embodiments.
[0093] In another embodiment of the present application, there is also provided a computer program product. The computer program product includes computer-executable instructions. The computer-executable instructions are stored in a computer-readable storage medium. At least one processor of the device can read the computer-executable instructions from the computer-readable storage medium. Executing the computer-executable instructions by at least one processor causes the device to execute the above Figure 1 or Figure 2 the method flow described in any one of the embodiments.
[0094] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of this application.
[0095] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0096] In several embodiments provided by the embodiments of this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0097] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0098] In addition, in each embodiment of the embodiments of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0099] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
Claims
1. A CPU instruction processing method, characterized in that, Applied to a central processing unit (CPU); the method includes: Obtaining a first load instruction and a first store instruction having a RAW (Read-After-Write) conflict; Updating the store sequence of the first load instruction based on the sequence number and / or access address of the first store instruction, where the store sequence is a set of store instructions having a RAW conflict with the first load instruction; Distributing the first load instruction based on the store sequence; The updating the store sequence of the first load instruction based on the sequence number and / or access address of the first store instruction includes: If the sequence number of the first store instruction does not exist in a preset memory access dependency mapping table, adding the first store instruction to the store sequence; If the sequence number of the first store instruction exists in the memory access dependency mapping table, keeping the store sequence unchanged.
2. The method according to claim 1, characterized in that, The updating the store sequence of the first load instruction based on the sequence number and / or access address of the first store instruction includes: If the sequence number of the first store instruction does not exist in a preset memory access dependency mapping table and the access address of the first store instruction is the same as the access address of the first load instruction, adding the first store instruction to the store sequence; If the sequence number of the first store instruction exists in the memory access dependency mapping table or the access address of the first store instruction is different from the access address of the first load instruction, keeping the store sequence unchanged.
3. The method according to any one of claims 1-2, characterized in that, After updating the store sequence of the first load instruction, the method further includes: In the case of adding the first store instruction to the store sequence, adding a counter corresponding to the first store instruction, and the initial count of the counter is zero; Whenever a RAW conflict occurs between the first store instruction and the first load instruction, incrementing the counter by one; If the counter is greater than or equal to a preset threshold, keeping the first store instruction in the store sequence until the store sequence is emptied.
4. The method according to any one of claims 1-2, characterized in that, The first load instruction includes the oldest load instruction in the global history register.
5. The method according to claim 4, wherein The processor includes a Store Sequence Identification Table (SSIT) and a store instruction list; the method further includes: In the case of adding the first store instruction to the store sequence, writing the store sequence identification number (SSID) of the first load instruction into the SSIT; Writing the sequence number and SSID of the first store instruction into the store instruction list.
6. The method according to any one of claims 1-2, characterized in that, The processor includes an SSIT, and the SSIT includes multiple blocks, and the identifier of each block corresponds to the low-order information of the program counter (PC); after obtaining the first load instruction and the first store instruction having a RAW conflict, the method further includes: Based on the first low-order information of the PC of the first load instruction, writing the store sequence identification number (SSID) of the first load instruction into the block corresponding to the first low-order information.
7. A processor, characterized in that, The processor is configured to execute the CPU instruction processing method according to any one of claims 1-6.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor; The memory is coupled to the processor; The memory is used to store computer program instructions; The processor is used to call the computer program instructions to execute the CPU instruction processing method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium; wherein, when the computer-executable instructions are executed by the processor, the CPU instruction processing method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Handling effective address synonyms in a load-store unit that operates without address translation
CN111133421A
Method for establishing loading and storing instruction dependency, processor and medium
CN117785285A