An instruction processing method and device, an electronic device, and a readable storage medium
The efficiency of the target instruction is determined by the efficiency of the target instruction.
Patent Information
- Application Number
- CN202511292684.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-10
AI Technical Summary
In the existing technology, the processor has difficulty efficiently solving the technical problem of determining instruction boundaries.
By comparing the target instruction boundary information with that obtained from the instruction cache, the boundary of the target instruction stream is determined, thereby improving the efficiency of instruction determination.
The efficiency of the target instruction has been achieved.
Smart Images

Figure CN120780361B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to an instruction processing method and device, electronic equipment and readable storage medium. BACKGROUND
[0002] With the rapid development of science and technology, various applications have higher requirements for processor performance. In the case of increasing data processing amount and faster processing speed, the existing processor gradually exposes some deficiencies in the process of reading and processing instructions.
[0003] In a modern processor, instructions need to go through the processes of instruction fetching, decoding and execution. Usually, the processor needs to obtain multiple instructions per cycle, usually continuously reads bytes, and then divides the read byte string into instructions. For fixed-length instructions, the division process is relatively simple; for variable-length instructions, the boundaries of each instruction need to be determined through decoding to find valid instructions, which greatly increases the decoding difficulty. The complexity of determining the instruction boundary in the existing processor increases the decoding delay, and further affects the overall performance of the processor. SUMMARY
[0004] Embodiments of the present application provide an instruction processing method and device, electronic equipment and readable storage medium, which can improve the efficiency of reading and executing instructions.
[0005] In a first aspect, an instruction processing method is disclosed, and the method comprises:
[0006] obtaining target instruction stream and target instruction boundary information from the cache according to the memory access request; the target instruction boundary information is used to indicate the boundary information of each instruction in the target instruction stream;
[0007] determining the boundary of each instruction in the target instruction stream according to the instruction start information and the target instruction boundary information; the instruction start information is used to represent whether the target field in the target instruction stream is a complete instruction;
[0008] cutting the target instruction stream according to the boundary of each instruction, and processing each instruction obtained by cutting.
[0009] In a second aspect, an instruction processing device is disclosed, and the device comprises:
[0010] an obtaining module, configured to obtain target instruction stream and target instruction boundary information from the cache according to the memory access request; the target instruction boundary information is used to indicate the boundary information of each instruction in the target instruction stream;
[0011] The boundary determining module is configured to determine the boundary of each instruction in the target instruction stream according to instruction start information and the target instruction boundary information; and the instruction start information is used to represent whether a target field in the target instruction stream is a complete instruction.
[0012] The processing module is configured to split the target instruction stream according to the boundary of each instruction and process each instruction obtained by the splitting.
[0013] In a third aspect, an electronic device is provided, which includes a processor, and a memory configured to store instructions executable by the processor; and the processor is configured to execute the instructions to implement the method in the first aspect.
[0014] In a fourth aspect, a computer readable storage medium is provided, and when instructions in the computer readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to implement the method in the first aspect.
[0015] The embodiments of the present application have the following advantages:
[0016] The target instruction stream and the target instruction boundary information are obtained from the cache according to the memory access request, and the boundary of each instruction in the target instruction stream is determined according to the boundary information of each instruction indicated by the target instruction boundary information. The storage space in the cache is used to pre-store the target instruction boundary information, so as to reduce the delay in reading instructions and improve the efficiency of determining the instruction boundary. Since the target instruction boundary information is stored in the cache instead of being calculated on site after the data is read by the processor front end, the complexity of the instruction reading circuit is reduced, so as to save the occupied area of the processor front end. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative labor.
[0018] Figure 1 is a step flow chart of an instruction processing method embodiment of the present application;
[0019] Figure 2 is a step flow chart of another instruction processing method embodiment of the present application;
[0020] Figure 3 is a pipeline architecture schematic diagram of an instruction processing method of the present application;
[0021] Figure 4is a pipeline architecture schematic diagram of another instruction processing method of the present application;
[0022] Figure 5 is a pipeline architecture schematic diagram of another instruction processing method of the present application;
[0023] Figure 6 is a structural block diagram of an instruction processing device of the present application;
[0024] Figure 7 is a structural block diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0026] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a class, and are not limited to the number of objects, for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents an "or" relationship between the associated objects before and after it. The term "multiple" in the embodiments of the present application means two or more, and other quantifiers are similar.
[0027] Method embodiments
[0028] Reference Figure 1 , a step flowchart of an instruction processing method embodiment of the present application is shown, which can specifically include the following steps:
[0029] Step 101, obtaining a target instruction stream and target instruction boundary information from the cache according to a memory access request; the target instruction boundary information is used to indicate the boundary information of each instruction in the target instruction stream;
[0030] Step 102, determining the boundary of each instruction in the target instruction stream according to the instruction start information and the target instruction boundary information; the instruction start information is used to represent whether a target field in the target instruction stream is a complete instruction;
[0031] Step 103, splitting the target instruction stream according to the boundary of each instruction, and processing each instruction obtained by splitting.
[0032] The instruction processing method provided in the application can be applied to a high-performance Reduced Instruction Set Computing-V (RISCV) processor. In the high-performance RISCV processor, the ability of the processor front end to read instructions is crucial. For an 8-way processor, the front end needs to ensure that 8 instructions can be retrieved per cycle on average to meet the needs of the back end. Therefore, the peak instruction retrieval capability of the front end generally needs to read 16 instructions per cycle. For the RISCV instruction set, in the case of enabling C extension, each instruction can be 2 bytes or 4 bytes, and the length of 16 instructions can be 32-64 bytes. Therefore, the front end of the 8-way RISCV processor needs to retrieve 64 bytes of instructions per cycle and perform decoding to ensure that the needs of the back end can be met in the worst case.
[0033] It should be noted that 8 issue refers to the core capability of the processor back end, indicating that the processor can send at most 8 instructions from the decoding or scheduling stage to the execution unit for execution at the same time in one clock cycle. Issue refers to the dispatch of instructions to the execution unit, and does not mean that the instructions are necessarily executed and completed in the clock cycle. The processor front end is responsible for fetching instructions from memory and performing preliminary processing to prepare them for execution by the back end. The processor front end is located at the beginning of the processor pipeline and performs tasks such as instruction fetching, decoding, branch prediction, etc. The instruction fetching task refers to reading the instruction stream from the instruction cache according to the program counter or the results of the branch predictor. The decoding task refers to parsing the fetched binary instruction code into micro-operations or control signals that the processor can understand, such as determining the instruction type, operands, target registers, etc. The branch prediction task refers to predicting the direction and target address of program branches to keep the instruction pipeline full and avoid stalling due to waiting for branch results. Program branches refer to conditional statements or loop statements, and direction refers to jumping or not jumping. The processor back end is located in the middle and later stages of the processor pipeline and receives instructions processed by the front end. The processor back end is mainly responsible for actually executing instructions, managing data flow and resources, and ensuring that instructions are completed correctly according to program logic. The main tasks include scheduling, execution, memory access, write back, etc. Scheduling refers to dispatching decoded instructions from the front end to the appropriate idle execution unit, involving complex logic such as resolving data dependencies and resource conflicts. The execution task refers to the execution unit executing the operations specified by the instruction. The execution unit specifically includes integer operation units, floating point operation units, load or store units, branch units, etc. The memory access task refers to handling load and store requests for data cache or memory, i.e., read data requests and write data requests. The write back task refers to writing the execution results back to the register file.
[0034] C extension is an optional extension defined in RISCV instruction set standard, which introduces a set of 16-bit length compressed instruction set for RISCV instruction set. Therefore, after the C extension is enabled, the instruction length acquired and processed by the processor is no longer fixed at 32 bits, but can be 16 bits or 32 bits. This makes the instruction stream "irregular", and the number of instructions contained in the number of bytes read by the processor front end when reading instructions is no longer fixed. Since the number of instructions contained in the number of bytes read by the processor front end each time and the instruction length are no longer fixed, the data stream read each time needs to be segmented to obtain the correct instruction before the instruction is processed, otherwise decoding errors may occur. For example, for 64-byte instructions retrieved, the boundaries of the instructions need to be determined first, so that the 64-byte binary data can be cut into independent instructions and decoded. This process needs to be executed in series, because the segmentation result of the previous instruction will affect the next instruction. Serial computing of the boundaries of 32 instructions will have a great impact on the timing, and may cause the timing to be unable to converge.
[0035] Taking the high-performance open-source processor core Kunming Lake V2 with 6-way issue as an example, the front end reads instructions in a sequential segmentation instruction manner, retrieves 32 bytes of instructions per cycle, calculates the boundaries of 16 instructions in a cycle, and sends the segmented instructions to the processor back end. However, the existing technology has a problem that the time consumed by the instruction fetch unit to calculate the boundaries of the instructions in series on site after retrieving the instructions is long, and the area occupied by the hardware is large, which affects the efficiency of decoding and processing the instructions, and further affects the overall performance of the processor.
[0036] In the embodiments of the present application, in order to solve the above technical problems, additional information describing the boundaries of the instructions is stored in the instruction cache together with the instruction data, so as to reduce the difficulty and complexity of determining the boundaries of the instructions after reading the instruction data from the cache, improve the efficiency of determining the boundaries of the instructions, and further improve the efficiency of processing the instructions and the overall performance of the processor.
[0037] The memory access request is a data read request sent by the instruction fetch unit to the instruction cache, the target instruction stream is a byte sequence returned from the instruction cache, and the target instruction boundary information is instruction boundary metadata returned from the instruction cache. The target instruction boundary information indicates which bytes in the target instruction stream belong to an instruction and which byte boundaries are instruction boundaries.
[0038] It should be noted that the instruction fetch unit generates the memory access request according to the program counter and the branch prediction result, and the cache receives the memory access request, and a circuit with a control logic part in the cache processes the memory access request, determines the storage address of the target instruction stream and the target instruction boundary information according to the memory access request, and returns the target instruction stream and the target instruction boundary information from the storage address to the instruction fetch unit. The instruction fetch unit determines the instruction boundary in the target instruction stream according to the target instruction boundary information, can mark the instruction boundary in the target instruction stream, cuts the target instruction stream into independent instructions according to the boundary mark, and can distribute the independent instructions to the multiple decoding channels in parallel; and the decoder converts the instructions in the machine code form into micro-operations, and sends the instructions converted into the micro-operations to the back-end execution unit through the emission queue.
[0039] The target instruction boundary information and the target instruction stream are stored separately, the instruction fetch unit can read the target instruction boundary information and the target instruction stream at the same time, or can read the target instruction boundary information first and then read the target instruction stream; for example, the metadata of the target instruction stream is returned in the first cycle, and the instruction data is returned in the second cycle. Since the target instruction boundary information is read in advance, the instruction boundary in the target instruction stream can be determined according to the target instruction boundary information when the target instruction stream is read, without the need to calculate the instruction boundary in the target instruction stream after the target instruction stream is completely read, so that the efficiency of determining the instruction boundary in the target instruction stream is improved, and the efficiency of cutting the target instruction stream according to the instruction boundary and sending the cut instructions to the back end of the processor for processing is improved.
[0040] The instruction start information is recorded after the instruction is read in the last cycle, and is used to indicate whether the starting position of the target instruction stream read in the current cycle contains an incomplete instruction, that is, an instruction across the cache line. For example, if the instruction at the end of the last read instruction is 4 bytes but only 2 bytes are stored, the first two bytes of the target instruction stream obtained by reading in the current cycle are incomplete instructions, and the starting position of the target instruction stream is not an instruction boundary. The boundaries of the instructions in the target instruction stream are associated with each other, and the boundary of a subsequent instruction is determined on the basis of the boundary of a connected previous instruction and the length of the previous instruction. Therefore, to determine the boundaries of the instructions in the target instruction stream, it is necessary to first determine the instruction boundary of the starting position of the target instruction stream according to the instruction start information, and then determine the boundaries of the instructions in the target instruction stream according to the target instruction boundary information. The starting address of the read instruction is carried in the memory access request, and the starting position of the target instruction stream is the position pointed to by the starting address.
[0041] Optionally, step 101 can specifically include the following steps:
[0042] The substep 1011 comprises: obtaining the target instruction stream and third instruction boundary information from the cache according to the memory access address; the third instruction boundary information is all the instruction boundary information already existing in the cache;
[0043] The substep 1012 comprises: verifying the third instruction boundary information according to the memory access address.
[0044] The substep 1013 comprises: in the case where the instruction boundary information corresponding to the memory access address exists in the third instruction boundary information, determining the instruction boundary information corresponding to the memory access address as the target instruction boundary information.
[0045] The substep 1014 comprises: in the case where the instruction boundary information corresponding to the memory access address does not exist in the third instruction boundary information, taking the memory access address as an instruction boundary, and determining the target instruction boundary information of the target instruction stream according to the length indicated by the last two bits of every two bytes after the memory access address; the target instruction boundary information is used to indicate whether every two byte boundaries in the target instruction stream are instruction boundaries.
[0046] The memory access request comprises a memory access address, the memory access address is taken as an instruction boundary, and the target instruction boundary information corresponding to the target instruction stream is determined according to the length indicated by the last two bits of every two bytes after the memory access address.
[0047] The target instruction boundary information is used to indicate whether every two byte boundaries in the target instruction stream are instruction boundaries.
[0048] In the embodiment of the application, the memory access address comprises a start address when reading an instruction, and the start of a complete instruction pointed to by the start address, that is, in the case where the start address points to a complete instruction, it is determined whether the instruction boundary information corresponding to the start address exists in the existing instruction boundary information. The instruction boundary information corresponding to the start address is not only instruction boundary information corresponding to an address, but also includes instruction boundary information after the start address in the cache line where the start address is located. It should be noted that when the processor performs sequential instruction fetching, it is determined that the next instruction pointed to by the start address is a complete instruction, or when the processor executes a jump, branch or call instruction, the program counter directly executes a new, non-sequential target address, and the target address itself must be the start boundary of an instruction.
[0049] The existing instruction boundary result is checked according to the start address. If the current start address has been recorded as an instruction boundary in the existing instruction boundary result, the check is passed, and the existing instruction boundary result is directly used to determine the instruction boundary of the target instruction stream. If the current start address is not recorded as an instruction boundary in the existing instruction boundary result, the current start address is taken as an instruction boundary, and the instruction boundary after the current address is calculated. The existing instruction boundary result can be stored in a cache or in a fetch unit. The check on the existing instruction boundary result and the calculation of the instruction boundary after the current address can be performed by the fetch unit or by a control logic circuit part in the cache.
[0050] When the instruction boundary information corresponding to the start address exists in the existing instruction boundary information, the instruction boundary information corresponding to the start address is directly determined as the target instruction boundary. When the instruction boundary information corresponding to the start address does not exist in the existing instruction boundary information, the target instruction boundary information corresponding to the target instruction stream to be read out is calculated based on the start address. After the program counter sends the target address to the fetch unit, the fetch unit generates a memory access request based on the target address. Since the target address is necessarily the start boundary of an instruction, the target address, that is, the start address, can be directly taken as the start instruction boundary of the target instruction stream, and the boundary of each instruction in the target instruction stream is determined in combination with the instruction length indicated by the lower two bits of each 2-byte unit.
[0051] For example, the start address is marked as a marker, and the current 2-byte unit is read from the start address. The lower two bits are checked. If the lower two bits are not 00, 10, or 01, it indicates that the current 2-byte unit is a 16-bit compressed instruction, and the start of the next instruction is aligned at the next 2-byte unit. The position aligned at the next 2-byte unit is marked as a boundary. If the lower two bits are 11, it indicates that the current 2-byte unit is the first half of a 32-bit instruction, and this instruction occupies 4 bytes. The start of the next instruction skips the next 2-byte unit and is located 4 bytes away from the current boundary. The position 4 bytes away is marked as a boundary. The above operation is repeated until the required instruction length is reached or the end address of the read instruction is reached.
[0052] The processor determines the instruction fetching mode of the current memory access request according to context information instead of only according to the address value carried by the memory access request. For example, when the branch predictor predicts that a branch instruction will jump, the target address is calculated, and an explicit control signal is carried when sending the instruction fetching request to the instruction fetching unit, indicating to the instruction fetching unit that the instruction fetching request is a jump prediction request, and the start address points to a complete instruction. In the case of sequential instruction fetching, it can also be determined according to the last period of reading the instruction whether the instruction at the end of the instruction stream read in the last period is a complete instruction. In the case where the instruction boundary corresponding to the start address does not exist in the existing instruction boundary information, the target instruction boundary information corresponding to the target instruction stream is calculated, and the newly calculated target instruction boundary information can also be written back to the cache.
[0053] In the embodiment of the application, the target instruction stream and the target instruction boundary information are obtained from the cache according to the memory access request, and the boundary of each instruction in the target instruction stream is determined according to the boundary information of each instruction indicated by the target instruction boundary information. The storage space in the cache is used to pre-store the target instruction boundary information, which reduces the delay of the critical path when reading instructions and improves the efficiency of determining the instruction boundary. Since the target instruction boundary information is stored in the cache instead of being calculated on site by the processor front end after reading data, the complexity of the instruction reading circuit is also reduced, thereby saving the occupied area of the processor front end.
[0054] Referring to Figure 2 , a step flowchart of another instruction processing method embodiment of the application is shown, which can specifically include the following steps:
[0055] Step 201, querying whether the memory access request hits the cache;
[0056] Step 202, in the case where the memory access request hits the cache, obtaining the target instruction stream and the target instruction boundary information from the cache;
[0057] Step 203, in the case where the memory access request does not hit the cache, sending the memory access request to the next level cache, receiving the target instruction stream returned by the next level cache, and calculating the target instruction boundary information corresponding to the target instruction stream.
[0058] For steps 201 to 203, the memory access request includes a fetch address and an instruction length, and the fetch address includes a tag field (Tag), an index field (Index), and an offset field (Offset). It should be noted that the reply of the instruction cache to the memory access request includes metadata, a tag, and instruction data. The metadata refers to data describing the instruction data, and is used to describe attribute information of the instruction data. The tag refers to a tag of a cache line where the instruction is located. The instruction data refers to a target instruction stream containing an actually executable binary instruction. The tag is used to identify the cache line where the instruction data is located. The metadata is auxiliary information used to describe the instruction data, such as a check code. In the embodiment of the present application, the instruction data refers to a target instruction stream, and the metadata includes not only the usual auxiliary information such as a check code, but also calculated target instruction boundary information.
[0059] The cache group is determined according to the index field in the memory access request, the tags of all cache lines in the cache group are read, and the tag of each cache line is matched with the tag carried in the physical address. In the case of successful matching, it is indicated that the memory access request hits the cache. In the case of failed matching, it is indicated that the memory access request misses the cache.
[0060] In the case of that the memory access request hits the cache, the data in the cache is accessed according to the offset in the hit cache line, and the target instruction stream and the target instruction boundary information are obtained in parallel. In the case of that the memory access request misses the cache, the memory access request is sent to the next level cache according to the hierarchical structure among the caches. The access time difference of different levels of caches is usually particularly large. The L1 cache generally needs only one cycle to obtain the result, while the L2 cache may need dozens to hundreds of cycles, and the memory request may need hundreds to thousands of cycles. For example, the current cache is L1, in the case of that the memory access request misses L1, the memory access request is sent to L2 cache. In the case of that L2 cache hits, L2 cache returns the target instruction stream. In the case of that L2 cache misses, the memory access request is sent to cache L3. In the case of that cache L3 hits, L3 cache returns the target instruction stream. In the case of that all caches miss, the main memory is accessed.
[0061] It should be noted that only the target instruction stream is saved in other levels, and the target instruction boundary information is calculated when the target instruction stream is written into the L1 cache for the first time. When the target instruction boundary information and the target instruction stream are accessed for the first time, the target instruction boundary information and the target instruction stream are written into the L1 cache, and thereafter, the target instruction stream and the target instruction boundary information need to be accessed, and the target instruction stream and the target instruction boundary information can be obtained from the L1 cache. When the target instruction stream and the target instruction boundary information are not stored in the L1 cache, the target instruction stream is obtained from the cache or the memory of other levels, and the corresponding target instruction boundary information is calculated based on the target instruction stream.
[0062] Optionally, step 203 can specifically include the following steps:
[0063] Sub-step 2031, performing AND operation on the lowest two bits of each two bytes in the target instruction stream to obtain the target instruction boundary information corresponding to the target instruction stream; the target instruction boundary information is used to indicate the length of each instruction.
[0064] It should be noted that in the RISCV instruction set, all instructions are aligned with 2 bytes as the basic unit, and the instruction length is represented by the lowest two bits of each two bytes. The length of each instruction can be 2 bytes or 4 bytes. Therefore, to determine the length of an instruction, the lowest two bits of each two bytes can be checked. Specifically, when the lowest two bits of each two bytes are 11, it indicates that the current instruction is a 4-byte instruction or a 32-bit instruction, that is, the 2-byte unit represents the first half of a 32-bit standard instruction, and the next 2-byte unit must be combined to form a complete instruction. When the lowest two bits of each two bytes are 00, 01 or 10, it indicates that the current instruction is a 16-bit instruction. After performing AND operation on the lowest two bits, the AND operation result of 11 is 1, indicating that the current 2-byte unit is the first half of a 32-bit instruction, and the next 2-byte unit is the second half of a 32-bit instruction, which needs to be combined with the next 2-byte unit to form a complete 32-bit instruction. The AND operation result of 00, 01 or 10 is 0, indicating that the current byte unit is the end of a 32-bit instruction or a complete 16-bit instruction. The AND operation result of the lowest two bits of each 2-byte unit is spliced to obtain the target instruction boundary information, which indirectly reflects whether each 2-byte unit is the start of a 32-bit instruction.
[0065] For example, assuming that the cache line size is 64 bytes, and the length of the instruction stream to be read each time is also 64 bytes, a 64-byte instruction data block, i.e., a target instruction stream, is received from the memory or a lower-level cache. The 64-byte instruction stream is divided into 32 consecutive 2-byte units, each of which is represented by . Therefore, there are 32 consecutive 2-byte units in the 64-byte instruction stream, i.e., to . For each unit, the lowest two bits of the 16-bit value are extracted. For example, the lowest two bits of are recorded as , the lowest two bits of are recorded as , and so on. The lowest two bits of each 2-byte unit are subjected to bitwise AND operation. That is, for each 2-byte unit , the value of is calculated.The result of the AND operation. If the last two bits are 11, the result of the AND operation is 1. If the last two bits are 01, 00 or 10, the result of the AND operation is 0. This operation result indicates whether the cell is the beginning of a 32-bit instruction, the result of 1 indicates that it is the beginning of a 32-bit instruction, and the result of 0 indicates that it is not the beginning of a 32-bit instruction, which can be the beginning of a complete 16-bit instruction or the middle part of a 32-bit instruction. The AND operation results of 32 cells are combined into a 32-bit bit vector, which is called target instruction boundary information. Each bit of the 32-bit target instruction boundary information corresponds to a cell, and the i-th bit corresponds to the i-th cell. If the i-th bit is 1, it means that the i-th cell is the beginning of a 32-bit instruction; if the i-th bit is 0, it means that the i-th cell is not the beginning of a 32-bit instruction, which can be the beginning of a complete 16-bit instruction or the middle part of a 32-bit instruction. The result of the AND operation.
[0066] In the embodiment of the present application, it is not necessary to check the instruction header byte by byte after reading the instruction, but the boundary information of the instruction is calculated in advance and stored when the instruction is loaded into the cache. When the processor front end needs to access the instruction, the instruction data and the boundary information can be read out at the same time, and the front end does not need to perform time-consuming serial decoding, but can directly determine how to split the instruction according to the boundary information.
[0067] Optionally, step 203 can specifically include the following steps:
[0068] Step S11, taking the zeroth byte boundary in the target instruction stream as an instruction boundary, determining first instruction boundary information of the target instruction stream according to the length indicated by the last two bits of each two bytes;
[0069] Step S12, taking the second byte boundary in the target instruction stream as an instruction boundary, determining second instruction boundary information of the target instruction stream according to the length indicated by the last two bits of each two bytes;
[0070] Step S13, splicing the first instruction boundary information and the second instruction boundary information to obtain target instruction boundary information corresponding to the target instruction stream;
[0071] The first instruction boundary information and the second instruction boundary information are used to indicate whether each two byte boundary in the target instruction stream is an instruction boundary.
[0072] It should be noted that there are two cases for the boundary of the instruction at the starting position of the target instruction stream, one of which is across the cache line, and the other of which is not across the cache line. Crossing the cache line means that this is a 32-bit instruction, and the first half of the 2-byte cell is in the previous cache line and the second half of the 2-byte cell is in the current cache line. Whether the instruction at the starting position of the target instruction stream crosses the cache line will cause the subsequent calculation of the boundaries of the instructions in the target instruction stream to also have corresponding two cases.
[0073] In the embodiments of the present application, the corresponding instruction boundaries of the target instruction stream in the two cases are calculated by respectively assuming that the instruction at the starting position of the target instruction stream crosses the cache line and does not cross the cache line. Specifically, assuming that the instruction at the starting position of the target instruction stream does not cross the cache line, the first instruction boundary of the target instruction stream is from byte 0; assuming that the instruction at the starting position of the target instruction stream crosses the cache line, this is a 32-bit instruction, the 2-byte unit in the first half is at the end of the previous cache line, and the 2-byte unit in the second half is at the beginning of the current cache line, therefore, the first instruction boundary in the target instruction stream is from byte 2.
[0074] In the case of determining the instruction boundary at the starting position of the target instruction stream, the boundary position of each instruction can be directly determined according to the instruction length indicated by the last two bits of each 2-byte unit, therefore, the first instruction boundary information and the second instruction boundary information directly reflect the boundary position of each instruction in the target instruction stream, rather than the length of the instruction.
[0075] Exemplarily, assuming that the read target instruction stream is 64 bytes, containing 32 continuous 2-byte units, each 2-byte unit is , the 32 continuous 2-byte units are to ; assuming that is a legal instruction starting boundary, from , the instruction length is judged according to the last two bits of each 2-byte unit, and the subsequent boundary is deduced. If the last two bits of are 11, it is the first half of a 32-bit instruction. It must constitute a complete instruction with . Therefore, the boundary of the next instruction is at , and is not a boundary. If the last two bits are not 11, it is a 16-bit instruction, and the boundary of the next instruction is at . The first instruction boundary information is a 32-bit bit vector, each bit corresponds to a 2-byte unit. In the case that one bit of the first instruction boundary information is 1, it indicates that the unit is a starting boundary of an instruction; in the case that one bit of the first instruction boundary information is 0, it indicates that the unit is not a starting boundary, but the second half of a 32-bit instruction. Similarly, assuming that is an instruction starting boundary. is marked as a non-boundary, from At the beginning, the same logical rule is applied to parse, and the boundary of the subsequent byte unit is pushed to obtain the second instruction boundary information of another 32-bit bit vector. The first instruction boundary information and the second instruction boundary information have the same format. The first instruction boundary information and the second instruction boundary information are spliced to obtain 64-bit target instruction boundary information, which represents two instruction boundary conditions.
[0076] Optionally, step 203 can specifically include the following steps:
[0077] Step S21, receiving the first instruction stream returned by the next level cache; the length of the first instruction stream is the same as the byte length that can be stored by the cache line, and the first instruction stream includes the target instruction stream;
[0078] Step S22, calculating the fourth instruction boundary information corresponding to the first instruction stream;
[0079] Step S23, determining the target instruction stream from the first instruction stream according to the start address and the end address, and determining the target instruction boundary information from the fourth instruction boundary information.
[0080] For steps S21 to S23, it should be noted that when the target instruction stream is obtained from the next level cache in the case that the memory request misses the cache, it is obtained in the unit of cache line, and when the instruction boundary information is calculated, it is also calculated in the unit of cache line.
[0081] Specifically, the memory request misses the L1 instruction cache, and a request is sent to the L2 cache. The L2 cache determines the target cache line according to the address indicated by the request, and returns the data of the entire target cache line to the L1 cache. Even if the processor core only wants part of the instruction data in the cache line, the L2 cache returns the entire cache line containing the target instruction stream. The boundaries of each instruction in the entire cache line are calculated according to the lowest two bits of each 2-byte unit, that is, the fourth instruction boundary information. It should be noted that when calculating the boundaries of each instruction in the entire cache line, it is necessary to determine whether the instruction at the start position of the cache line is complete or incomplete. The lowest two bits of each 2-byte unit in the entire cache line can be ANDed to obtain the first instruction boundary information. Alternatively, it can be assumed that the instruction at the start position of the entire cache line is complete or incomplete, respectively, and the start position is regarded as the instruction boundary or the second byte boundary is regarded as the instruction boundary, respectively. The boundaries of each instruction in the entire cache line are calculated according to the lowest two bits of each 2-byte unit. When the processor performs a jump instruction, the entire cache line is also read, and the instruction boundary information is calculated from the start address indicated by the memory request in the cache line until the end of the entire cache line.
[0082] After obtaining the instruction data of the whole cache line and the fourth instruction boundary information, the position of the target instruction stream in the whole cache line is determined according to the start address and the end address indicated by the memory access request, and the data between the start address and the end address is cut from the data in the whole cache line to obtain the target instruction stream. The instruction boundary information corresponding to the instruction data is cut from the fourth instruction boundary information to obtain the target instruction boundary information.
[0083] In step 204, the boundary of each instruction in the target instruction stream is determined according to the instruction start information and the target instruction boundary information.
[0084] This step can refer to the above-mentioned step 102, and details are not described here.
[0085] Optionally, step 204 can specifically include:
[0086] In step S31, in the case that the instruction start information indicates that the first two bytes or the first 4 bytes in the target instruction stream are complete instructions, the zeroth byte boundary in the target instruction stream is determined as the instruction boundary, and the boundary of each instruction in the target instruction stream is determined according to the length of the instruction indicated by the target instruction boundary information.
[0087] In step S32, in the case that the instruction start information indicates that the first two bytes in the target instruction stream are not complete instructions, the second byte boundary of the target instruction stream is determined as the instruction boundary, and the boundary of each instruction is determined according to the length of the instruction indicated by the target instruction boundary information.
[0088] It should be noted that, for steps S31 and S32, the target instruction boundary information is the result of the AND operation of the lower two bits of each 2-byte unit in the target instruction stream. The instruction start information indicates whether the last instruction of the instruction data obtained by the last reading is complete, and thus indirectly reflects whether the start instruction of the current reading instruction is complete.
[0089] In the case that the instruction start information indicates that the first two bytes in the target instruction stream are not a complete instruction, it is indicated that the first two bytes in the target instruction stream are the second half of a 32-bit instruction, therefore, the 0th byte boundary is not the boundary of a complete instruction, but the middle of an instruction, the first two bytes are skipped, and the second byte boundary is the boundary of an instruction. Since the target instruction boundary information is the result of an AND operation on the lower two bits of each 2-byte unit in the target instruction stream, the target instruction boundary information can indicate whether the current 2-byte unit is the first half of a 32-bit instruction or a complete 16-bit instruction. In the case that the second byte is the boundary of an instruction, each subsequent target instruction boundary information is traversed. In the case that the target instruction boundary information is 1, it is indicated that the current 2-byte unit is the first half of a 32-bit instruction, and a complete 32-bit instruction needs to be formed in combination with the next 2-byte unit, and the next instruction boundary is marked at a 4-byte unit position apart from the current 2-byte unit as the instruction boundary. In the case that the target instruction boundary information is 0, it is indicated that the current 2-byte unit is a complete 16-bit instruction, and the next instruction boundary is marked at the boundary of the next 2-byte unit as the instruction boundary.
[0090] In the case that the instruction start information indicates that the first two bytes in the target instruction stream are a complete instruction or the first four bytes are a complete instruction, it is indicated that the end of the instruction data read in the last instruction reading cycle is a complete instruction. Therefore, the start position of the current target instruction stream is the start boundary of an instruction, and the length of the instruction can be 32 bits or 16 bits, which needs to be determined according to the target instruction boundary information corresponding to the start position. In the case that the target instruction boundary information corresponding to the start position is 0, it is indicated that the first instruction in the target instruction stream is a complete 16-bit instruction. In the case that the target instruction boundary information corresponding to the start position is 1, it is indicated that the first instruction in the target instruction stream is a 32-bit instruction. Similarly, based on subsequent target instruction boundary information, the boundary of each instruction is determined.
[0091] In the embodiments of the present application, the target instruction boundary information indirectly indicates whether the instruction is a 32-bit instruction, and after the fetch unit reads the target instruction stream, the boundary of each instruction in the target instruction stream is determined based on the target instruction boundary information, rather than the lower two bits of each 2-byte unit, and the boundary of each instruction in the target instruction stream is calculated in real time and in series, which reduces the complexity of determining the instruction boundary by the fetch unit, shortens the length of the critical path of the processor front end, and improves the efficiency of determining the instruction boundary.
[0092] Optionally, step 204 can specifically include:
[0093] Step S41, in the case that the instruction start information indicates that the first two bytes or the first four bytes in the target instruction stream are a complete instruction, determining the instruction boundary indicated by the first instruction boundary information in the target instruction boundary information as the boundary of each instruction.
[0094] Step S42, in the case that the instruction start information indicates that the first two bytes in the target instruction stream are not a complete instruction, determining the instruction boundary indicated by the second instruction boundary information in the target instruction boundary information as the boundary of each instruction.
[0095] The first instruction boundary is the boundary of each instruction in the target instruction stream indicated assuming that the zeroth byte boundary in the target instruction stream is an instruction boundary; and the second instruction boundary is the boundary of each instruction in the target instruction stream indicated assuming that the second byte boundary in the target instruction stream is an instruction boundary.
[0096] In the case that the instruction start information indicates that the instruction at the first 2-byte unit in the target instruction stream is not the second half of a 32-bit instruction, the data of the first 2-byte unit in the target instruction stream is a 16-bit instruction or the first half of a 32-bit instruction, that is, the first two bytes or the first four bytes of the target instruction stream are a complete instruction. That is, there is no incomplete instruction at the start position of the target instruction stream, and the zeroth byte boundary is an instruction boundary, so the instruction boundary indicated by the first instruction boundary information is directly selected as the boundary of each instruction in the target instruction stream.
[0097] In the case that the instruction start information indicates that the instruction at the first 2-byte unit in the target instruction stream is the second half of a 32-bit instruction, the zeroth byte is not an instruction boundary, but the middle of an instruction, and the second byte boundary is an instruction boundary, and the second instruction boundary information is the instruction boundary calculated by taking the second byte boundary as an instruction boundary. Therefore, the instruction boundary indicated by the second instruction boundary information is taken as the boundary of each instruction in the target instruction stream.
[0098] In the embodiment of the present application, the target instruction boundary information is calculated when the target instruction stream is written into the cache, and the fetch unit can directly obtain the target instruction boundary information from the cache, only needs to select one boundary combination from the target instruction boundary information according to the instruction start information, and does not need to calculate the instruction boundary in real time when the fetch unit reads the target instruction stream. The boundary calculation step is placed in the miss cache which is not sensitive to delay, which can further shorten the length of the processor front-end critical path, improve the reading instruction bandwidth, improve the instruction bandwidth, and simplify the hardware control logic.
[0099] Optionally, step 204 can specifically include:
[0100] The instruction boundary indicated by the target instruction boundary information is determined as the boundary of each instruction in the target instruction stream.
[0101] In the embodiment of the present application, the target instruction boundary information is the instruction boundary calculated by taking the start address as the instruction boundary. Therefore, the instruction boundary indicated by the target instruction boundary information can be directly determined as the boundary of each instruction in the target instruction stream.
[0102] In step 205, a continuous bit sequence in the boundary range is extracted as a complete instruction according to the boundary of each instruction in the target instruction stream.
[0103] In step 206, the complete instruction is parsed to obtain the operation code and the operand of the complete instruction.
[0104] In step 207, the operation corresponding to the complete instruction is performed according to the operation code and the operand.
[0105] For steps 205 to 207, the target instruction stream is the original instruction byte sequence read from the instruction cache, and the complete instruction is the independent 16-bit instruction or 32-bit instruction cut out. Specifically, the target instruction stream can be cut out according to the determined boundary by a many-to-many selector. The many-to-many selector is a highly parallel hardware switch network, the input end of which is connected to each byte in the target instruction stream, and the output end of which is connected to the input of a plurality of decoders. The many-to-many selector takes the boundary result as a control signal and routes each byte of a plurality of instructions to the correct decoding channel. For example, the 0th, 2nd and 6th bytes in the target instruction stream are the instruction boundaries indicated by the target instruction boundary information, and the 0th, 2nd and 6th bytes in the target instruction stream are cut out to obtain the 16-bit instruction and the 32-bit instruction.
[0106] The operation code is a binary field in the instruction that defines the basic operation, such as addition, subtraction or jump. The operand is the operation object in the instruction, which can be a register operand or an immediate number. The cut-out instruction is sent to the decoder, which is a hardware module that can translate binary instructions into internal control signals of the processor. After the decoder translates the instruction, it is sent to the execution unit, which is responsible for actual calculation according to the instruction.
[0107] Referring to Figure 3 , a pipeline architecture schematic diagram of an instruction processing method of the present application is shown, the additional information refers to the target instruction boundary information, the start information refers to the instruction start information, the instruction data refers to the target instruction stream, and the instruction request refers to the memory request.
[0108] In the first stage of the pipeline, the memory access request is sent to the cache, and the target instruction stream and the target instruction boundary information are obtained from the cache, and the target instruction boundary information indirectly indicates whether the current 2-byte unit is the first half of the 32-bit instruction.
[0109] In the second stage of the pipeline, it is determined according to the instruction start information whether the target instruction stream start position is the second half of the 32-bit instruction, and in the case of determining the case of the instruction at the target instruction stream start position, the length of the instruction indicated by the target instruction boundary information is used to calculate the boundary of each instruction in the target instruction stream.
[0110] In the third stage of the pipeline, the target instruction stream is divided based on the calculated instruction boundary result, each instruction after the division is sent to the decoder for decoding, and the decoded instruction is executed through the processor back end.
[0111] Referring to Figure 4 , another pipeline architecture diagram of the instruction processing method of the application is shown, the additional information refers to the target instruction boundary information, the start information refers to the instruction start information, the instruction data refers to the target instruction stream, and the instruction fetch request refers to the memory access request.
[0112] In the first stage of the pipeline, the memory access request is sent to the cache, and the target instruction stream and the target instruction boundary information are obtained from the cache.
[0113] In the second stage of the pipeline, the target instruction boundary information includes first instruction boundary information and second instruction boundary information, and one of the first instruction boundary information and the second instruction boundary information is selected as the boundary of each instruction in the target instruction stream according to the start information. And based on the calculated instruction boundary result, the target instruction stream is divided, each instruction after the division is sent to the decoder for decoding, and the decoded instruction is executed through the processor back end.
[0114] Referring to Figure 5 , another pipeline architecture diagram of the instruction processing method of the application is shown, in which the start information refers to the instruction start information, the instruction data refers to the target instruction stream, and the instruction fetch request refers to the memory access request.
[0115] In the first stage of the pipeline, the memory access request is sent to the cache, and the target instruction stream and the target instruction boundary information are obtained from the cache, and the target instruction boundary information is all the boundary information already in the cache.
[0116] In the second stage of the pipeline, the existing boundary information is checked according to the start address, if the start address corresponding boundary information exists in the existing boundary information, the existing boundary information is determined as the boundary of each instruction in the target instruction stream. And based on the existing instruction boundary result, the target instruction stream is segmented, each instruction after segmentation is sent to the decoder for decoding, and the decoded instruction is executed through the processor backend.
[0117] If the start address corresponding boundary information does not exist in the existing boundary information, the start address is taken as the instruction boundary, and the boundary of each instruction in the target instruction stream is calculated. In the third stage of the pipeline, based on the calculated instruction boundary result, the target instruction stream is segmented, each instruction after segmentation is sent to the decoder for decoding, and the decoded instruction is executed through the processor backend.
[0118] To sum up, the instruction processing method provided by the application sends a memory access request to a cache, obtains a target instruction stream and target instruction boundary information from the cache, wherein the target instruction boundary information can indirectly indicate the length of an instruction, so that the boundary of the target instruction stream can be calculated based on the target instruction boundary information when reading the instruction, the length of the critical path of the processor front end is shortened, the process of calculating the instruction boundary when reading the instruction is simplified, and the efficiency of calculating the instruction boundary is improved. The target instruction boundary information can directly indicate the boundary of each instruction in the target instruction stream, so that the instruction boundary indicated by the target instruction boundary information can be directly determined as the boundary of each instruction in the target instruction stream when reading the instruction, the target instruction stream is segmented, the length of the critical path of the processor front end is further shortened, and the corresponding hardware area overhead is reduced. Due to the foregoing improved operation, the efficiency of determining the instruction boundary is improved, and the efficiency of subsequent instruction segmentation and instruction execution according to the boundary is also improved, thereby improving the overall performance of the processor.
[0119] Reference Figure 6 An instruction processing apparatus of the application is shown, which can specifically include the following modules:
[0120] The obtaining module 310 is configured to obtain a target instruction stream and target instruction boundary information from a cache according to a memory access request; the target instruction boundary information is used to indicate the boundary information of each instruction in the target instruction stream;
[0121] The boundary determination module 320 is configured to determine the boundary of each instruction in the target instruction stream according to instruction start information and the target instruction boundary information; the instruction start information is used to represent whether the target field in the target instruction stream is a complete instruction;
[0122] The processing module 330 is configured to segment the target instruction stream according to the boundary of each instruction, and process each instruction obtained by segmentation.
[0123] Optionally, the obtaining module comprises:
[0124] a querying module, configured to query whether the memory access request hits a cache;
[0125] a first obtaining submodule, configured to obtain, in a case where the memory access request hits the cache, a target instruction stream and target instruction boundary information from the cache;
[0126] a calculating module, configured to, in a case where the memory access request does not hit the cache, send the memory access request to a next-level cache, receive a target instruction stream returned by the next-level cache, and calculate target instruction boundary information corresponding to the target instruction stream.
[0127] Optionally, the calculating module comprises:
[0128] a first calculating submodule, configured to perform an AND operation on the last two bits of each two bytes in the target instruction stream to obtain the target instruction boundary information corresponding to the target instruction stream; the target instruction boundary information is used to indicate the length of each instruction.
[0129] Optionally, the calculating module comprises:
[0130] a first boundary information determining module, configured to take the zeroth byte boundary in the target instruction stream as an instruction boundary, and determine first instruction boundary information of the target instruction stream according to the length indicated by the last two bits of each two bytes;
[0131] a second boundary information determining module, configured to take the second byte boundary in the target instruction stream as an instruction boundary, and determine second instruction boundary information of the target instruction stream according to the length indicated by the last two bits of each two bytes;
[0132] a splicing module, configured to splice the first instruction boundary information and the second instruction boundary information to obtain the target instruction boundary information corresponding to the target instruction stream;
[0133] The first instruction boundary information and the second instruction boundary information are used to indicate whether each two byte boundary in the target instruction stream is an instruction boundary.
[0134] Optionally, the obtaining module comprises:
[0135] a second obtaining submodule, configured to obtain, according to the memory access address, a target instruction stream and third instruction boundary information from the cache; the third instruction boundary information is all instruction boundary information already in the cache;
[0136] a verifying module, configured to verify the third instruction boundary information according to the memory access address.
[0137] The third boundary information determining module is configured to determine the instruction boundary information corresponding to the memory access address as target instruction boundary information if the instruction boundary information corresponding to the memory access address exists in the third instruction boundary information.
[0138] The fourth boundary information determining module is configured to, if the instruction boundary information corresponding to the memory access address does not exist in the third instruction boundary information, determine the memory access address as an instruction boundary, and determine target instruction boundary information of the target instruction stream according to the length indicated by the last two bits of every two bytes after the memory access address; the target instruction boundary information is used to indicate whether every two byte boundary in the target instruction stream is an instruction boundary.
[0139] Optionally, the computing module comprises:
[0140] The receiving module is configured to receive the first instruction stream returned by the next-level cache; the length of the first instruction stream is the same as the byte length that can be stored by the cache line, and the first instruction stream comprises the target instruction stream.
[0141] The second computing submodule is configured to compute fourth instruction boundary information corresponding to the first instruction stream.
[0142] The fifth boundary information determining module is configured to determine the target instruction stream from the first instruction stream according to the start address and the end address, and determine the target instruction boundary information from the fourth instruction boundary information.
[0143] Optionally, the boundary determining module comprises:
[0144] The first boundary determining submodule is configured to, if the instruction start information indicates that the first two bytes or the first four bytes in the target instruction stream are complete instructions, determine the zeroth byte boundary in the target instruction stream as an instruction boundary, and determine the boundary of each instruction in the target instruction stream according to the length of the instruction indicated by the target instruction boundary information.
[0145] The second boundary determining submodule is configured to, if the instruction start information indicates that the first two bytes in the target instruction stream are not complete instructions, determine the second byte boundary of the target instruction stream as an instruction boundary, and determine the boundary of each instruction according to the length of the instruction indicated by the target instruction boundary information.
[0146] Optionally, the boundary determining module comprises:
[0147] a third boundary determination submodule, configured to determine, as the boundary of each instruction, the instruction boundary indicated by first instruction boundary information in the target instruction boundary information, in a case where the instruction start information indicates that the first two bytes or the first four bytes in the target instruction stream are a complete instruction;
[0148] a fourth boundary determination submodule, configured to determine, as the boundary of each instruction, the instruction boundary indicated by second instruction boundary information in the target instruction boundary information, in a case where the instruction start information indicates that the first two bytes in the target instruction stream are not a complete instruction.
[0149] Optionally, the boundary determination module comprises:
[0150] a fifth boundary determination submodule, configured to determine, as the boundary of each instruction in the target instruction stream, the instruction boundary indicated by the target instruction boundary information.
[0151] Optionally, the processing module comprises:
[0152] an extraction module, configured to extract, as a complete instruction, a continuous bit sequence within the boundary range according to the boundary of each instruction from the target instruction stream;
[0153] an analysis module, configured to analyze the instruction to obtain an operation code and an operand of the instruction for each complete instruction;
[0154] an execution module, configured to execute an operation corresponding to the instruction according to the operation code and the operand.
[0155] In summary, the instruction processing apparatus provided in the present application sends a memory access request to a cache, and obtains a target instruction stream and target instruction boundary information from the cache. The target instruction boundary information can indirectly indicate the length of an instruction, so that the boundary of the target instruction stream can be calculated based on the target instruction boundary information when reading the instruction, the length of the critical path of the processor front end is shortened, the process of calculating the instruction boundary when reading the instruction is simplified, and the efficiency of calculating the instruction boundary is improved. The target instruction boundary information can directly indicate the boundary of each instruction in the target instruction stream, so that the instruction boundary indicated by the target instruction boundary information can be determined as the boundary of each instruction in the target instruction stream when reading the instruction, the target instruction stream is segmented, the length of the critical path of the processor front end is further shortened, and the corresponding hardware area overhead is reduced. Due to the foregoing improvement operation, the efficiency of determining the instruction boundary is improved, and the efficiency of subsequent instruction segmentation and instruction execution according to the boundary is also improved, and the overall performance of the processor is improved.
[0156] For the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts are described in the part of the method embodiment.
[0157] Each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other.
[0158] As to the processor in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be described in detail here.
[0159] Referring to Figure 7 is a structural block diagram of an electronic device for instruction processing provided by an embodiment of the present application. As shown in Figure 7 , the electronic device includes a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface complete communication with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions make the processor execute the instruction processing method of the foregoing embodiments.
[0160] The processor can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, digital signal processor), an ASIC (Application Specific Integrated Circuit, application specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable device, transistor logic device, hardware component or any combination thereof. The processor can also be a combination that realizes computing functions, such as one or more microprocessor combinations, combinations of DSP and microprocessor, etc.
[0161] The communication bus can include a channel for transmitting information between the memory and the communication interface. The communication bus can be a PCI (Peripheral Component Interconnect, peripheral component interconnect) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only one line is used in the specification, but it does not mean that there is only one bus or one type of bus.
[0162] The memory can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0163] The embodiment of the present application also provides a non-transitory computer readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device (a server or a terminal), the processor can execute the instruction processing method shown in the embodiment of the present application. Figure 1 The instruction processing method shown in the embodiment of the present application.
[0164] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same and similar parts between each embodiment can be referred to each other.
[0165] Those skilled in the art should understand that the embodiments of the embodiment of the present application can be provided as a method, an apparatus or a computer program product. Therefore, the embodiment of the present application can adopt a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, the embodiment of the present application can adopt a computer program product in the form of one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.
[0166] The embodiment of the present application is described with reference to the flowchart and / or block diagram of the method, terminal device (system) and computer program product according to the embodiment of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be realized by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a machine that implements the function specified in the flow Figure 1 The apparatus that realizes the function specified in one flow or multiple flows and / or one block or multiple blocks. Figure 1 The apparatus that realizes the function specified in one flow or multiple flows and / or one block or multiple blocks.
[0167] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a predetennined manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart or flowsheets and / or block or blocks. Figure 1 of the flowchart or flowsheets and / or block or blocks. Figure 1 of the flowchart or flowsheets and / or block or blocks.
[0168] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the function specified in the flowchart or flowsheets and / or block or blocks. Figure 1 of the flowchart or flowsheets and / or block or blocks. Figure 1 of the flowchart or flowsheets and / or block or blocks.
[0169] While preferred embodiments of the application have been described, modifications and alterations thereto will occur to those skilled in the art upon reading the preceding description. In particular, it will be apparent to one of ordinary skill in the art that the present application can be implemented in other ways than those specifically set forth herein. Accordingly, the appended claims are intended to include within their scope all such modifications and alterations as fall within the scope of the present application. It will be apparent to those skilled in the art that aspects of the present application can be implemented in part or in whole on a number of different types of computing devices. For example, the present application can be implemented on a number of different types of computing devices, such as a personal computer, a workstation, a server, a handheld device, a mobile device, a network appliance, a distributed computing environment, and the like. The present application can be implemented on any other type of computing device.
[0170] Finally, it should be noted that, in this document, the terms "computer", "server" and "database" are used generically to refer to any computing device or devices suitable to implement the present application, and can include multiple such devices or systems linked together by one or more communications networks.
[0171] The above provides a kind of instruction processing method, device, electronic equipment and readable storage medium provided by the present application, detailed introduction is carried out, the principle and implementation mode of the present application are described in this document with specific examples;The above example is only for helping to understand the method of the present application and its core idea;For those skilled in the art, according to the idea of the present application, there will be changes in specific implementation mode and application range;Summarized above, the content of the specification should not be understood as the limitation of the present application.
Claims
1. An instruction processing method, characterized by, The method comprises: acquiring a target instruction stream and target instruction boundary information from a cache according to a memory access request; the target instruction boundary information is used to indicate boundary information of each instruction in the target instruction stream; determining the boundary of each instruction in the target instruction stream according to instruction start information and the target instruction boundary information; the instruction start information is used to represent whether a target field in the target instruction stream is a complete instruction; segmenting the target instruction stream according to the boundary of each instruction and processing each instruction segmented; the acquiring of the target instruction stream and the target instruction boundary information from the cache according to the memory access request comprises: inquiring whether the memory access request hits the cache; in the case that the memory access request does not hit the cache, sending the memory access request to a next-level cache, receiving a target instruction stream returned by the next-level cache, and calculating target instruction boundary information corresponding to the target instruction stream; the calculating of the target instruction boundary information corresponding to the target instruction stream comprises: performing AND operation on the last two bits of each two bytes in the target instruction stream to obtain the target instruction boundary information corresponding to the target instruction stream; the target instruction boundary information is used to indicate the length of each instruction.
2. The method of claim 1, wherein, the calculating of the target instruction boundary information corresponding to the target instruction stream comprises: taking the zeroth byte boundary in the target instruction stream as an instruction boundary, and determining first instruction boundary information of the target instruction stream according to the length indicated by the last two bits of each two bytes; taking the second byte boundary in the target instruction stream as an instruction boundary, and determining second instruction boundary information of the target instruction stream according to the length indicated by the last two bits of each two bytes; splicing the first instruction boundary information and the second instruction boundary information to obtain the target instruction boundary information corresponding to the target instruction stream; wherein, the first instruction boundary information and the second instruction boundary information are used to indicate whether each two byte boundary in the target instruction stream is an instruction boundary.
3. The method of claim 1, wherein, the memory access request comprises a memory access address; the acquiring of the target instruction stream and the target instruction boundary information from the cache according to the memory access request comprises: acquiring the target instruction stream and third instruction boundary information from the cache according to the memory access address; the third instruction boundary information is all the instruction boundary information already in the cache; verifying the third instruction boundary information according to the memory access address; in the case that there is instruction boundary information corresponding to the memory access address in the third instruction boundary information, determining the instruction boundary information corresponding to the memory access address as the target instruction boundary information; in the case that there is no instruction boundary information corresponding to the memory access address in the third instruction boundary information, taking the memory access address as an instruction boundary, and determining the target instruction boundary information corresponding to the target instruction stream according to the length indicated by the last two bits of each two bytes after the memory access address; the target instruction boundary information is used to indicate whether each two byte boundary in the target instruction stream is an instruction boundary.
4. The method of claim 1, wherein, The memory access request comprises a start address and an end address of a read instruction; the target instruction stream returned by the next level cache is received, and target instruction boundary information corresponding to the target instruction stream is calculated, comprising: The first instruction stream returned by the next level cache is received; the length of the first instruction stream is the same as the byte length that can be stored by a cache line, and the first instruction stream comprises the target instruction stream; Fourth instruction boundary information corresponding to the first instruction stream is calculated; The target instruction stream is determined from the first instruction stream according to the start address and the end address, and the target instruction boundary information is determined from the fourth instruction boundary information.
5. The method of claim 1, wherein, The determination of the boundary of each instruction in the target instruction stream according to the instruction start information and the target instruction boundary information comprises: In a case where the instruction start information indicates that the first two bytes or the first 4 bytes in the target instruction stream are a complete instruction, the zeroth byte boundary in the target instruction stream is determined as an instruction boundary, and the boundary of each instruction in the target instruction stream is determined according to the length of the instruction indicated by the target instruction boundary information; In a case where the instruction start information indicates that the first two bytes in the target instruction stream are not a complete instruction, the second byte boundary of the target instruction stream is determined as an instruction boundary, and the boundary of each instruction is determined according to the length of the instruction indicated by the target instruction boundary information.
6. The method of claim 2, wherein, The determination of the boundary of each instruction in the target instruction stream according to the instruction start information and the target instruction boundary information comprises: In a case where the instruction start information indicates that the first two bytes or the first 4 bytes in the target instruction stream are a complete instruction, the instruction boundary indicated by the first instruction boundary information in the target instruction boundary information is determined as the boundary of each instruction; In a case where the instruction start information indicates that the first two bytes in the target instruction stream are not a complete instruction, the instruction boundary indicated by the second instruction boundary information in the target instruction boundary information is determined as the boundary of each instruction.
7. The method of claim 3, wherein, The determination of the boundary of each instruction in the target instruction stream according to the instruction start information and the target instruction boundary information comprises: The instruction boundary indicated by the target instruction boundary information is determined as the boundary of each instruction in the target instruction stream.
8. The method of claim 1, wherein, The splitting of the target instruction stream according to the boundary of each instruction and the processing of each instruction obtained by the splitting comprise: A continuous bit sequence within the boundary of each instruction is extracted from the target instruction stream as a complete instruction according to the boundary of each instruction; Each complete instruction is parsed to obtain an operation code and an operand of the complete instruction; An operation corresponding to the complete instruction is performed according to the operation code and the operand.
9. An instruction processing apparatus, characterized by, The apparatus comprises: An acquisition module configured to acquire a target instruction stream and target instruction boundary information from a cache according to a memory access request; the target instruction boundary information is used to indicate boundary information of each instruction in the target instruction stream; The boundary determination module is configured to determine a boundary of each instruction in the target instruction stream according to instruction start information and the target instruction boundary information; the instruction start information is used to represent whether a target field in the target instruction stream is a complete instruction; The processing module is configured to split the target instruction stream according to the boundary of each instruction, and process each instruction obtained by splitting. The obtaining module is specifically configured to: query whether the memory access request hits the cache; in a case where the memory access request does not hit the cache, send the memory access request to a next-level cache, receive a target instruction stream returned by the next-level cache, and calculate target instruction boundary information corresponding to the target instruction stream; the calculation of the target instruction boundary information corresponding to the target instruction stream includes: performing an AND operation on the last two bits of each two bytes in the target instruction stream to obtain the target instruction boundary information corresponding to the target instruction stream; the target instruction boundary information is used to indicate a length of each instruction.
10. An electronic device, comprising: comprise: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method of any one of claims 1 to 8.
11. A computer readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can perform the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Instruction distribution method, processor, chip and electronic equipment
CN115658150A