Data processing method and device, electronic equipment and storage medium

By using a head buffer in the processor's execution module to cache and distribute instructions, and requesting the transmit queue to stop sending instructions when necessary, the problems of processor instruction processing timing delay and performance bottlenecks in the prior art are solved, and more efficient instruction processing and performance improvement are achieved.

CN120066586APending Publication Date: 2025-05-30ARM TECH CHINA CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510222889.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The transmit queues and execution modules in existing processors have timing delays and performance bottlenecks when processing instructions, especially when the processing logic of the core core and VPU is inconsistent.

Method used

By introducing a head buffer into the execution module of the processor, matching instructions in the transmit queue are cached to the head buffer, and if necessary, request the transmit queue to stop sending instructions to avoid buffer overload.

Benefits of technology

Improves the efficiency of instruction distribution, reduces interference between core core and VPU, improves the overall performance of the processor, and simplifies the logic driving of the transmit queue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066586A_ABST
    Figure CN120066586A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, electronic equipment and a computer readable storage medium, and relates to the field of processors. The method comprises the following steps: receiving an instruction matched with an execution module in an IQ, and caching the instruction to a head buffer; reading the cached instruction from the head buffer, decoding the instruction to obtain a microcode, caching the microcode to an inter-stage register, reading the cached microcode from the inter-stage register, and processing the cached microcode to obtain a processing result of the corresponding instruction; and when the microcode cached in the inter-stage register reaches a first cache threshold value, stopping reading the cached instruction from the head buffer, and when the instruction cached in the head buffer reaches a second cache threshold value, requesting the transmitting queue to stop transmitting the instruction, so that the transmitting queue stops transmitting the instruction to the execution module. According to the embodiment of the invention, the time sequence of the transmitting queue is better, and the performance of the execution module is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of processors. Specifically, this application relates to a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] The issue queue (IQ) in a processor is a key component in the processor pipeline and is a cache in the processor for storing instructions that have been register-renamed but not yet sent to the execution module (core or VPU) for execution. Its main function is to select instructions whose source operands are all ready according to certain rules and send them to the execution module for execution. Summary of the Invention

[0003] Embodiments of this application provide a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can solve the above problems of the prior art. The technical solutions are as follows: According to one aspect of the embodiments of this application, a data processing method is provided, which is applied to an execution module in a processor. The execution module is a core or a vector processing unit (VPU). The method includes: Receiving an instruction in the issue queue (IQ) that matches the execution module, and caching the instruction in the head buffer; Reading the cached instruction from the head buffer for decoding to obtain microcode, caching the microcode in the inter-stage register, and reading the cached microcode from the inter-stage register for processing to obtain a processing result of the corresponding instruction; When the microcode cached in the inter-stage register reaches a first cache threshold, stop reading the cached instruction from the head buffer. When the instruction cached in the head buffer reaches a second cache threshold, request the issue queue to stop sending instructions, so that the sending queue stops sending instructions to the execution module.

[0004] According to one aspect of the embodiments of this application, a data processing method is provided, which is applied to the issue queue in a processor. The method includes: Aggregating instructions; Sending the instruction to an execution module that matches the instruction. The execution module is a core or a VPU, so that the execution module caches the instruction in the head buffer, reads the cached instruction from the head buffer for decoding to obtain microcode, caches the microcode in the inter-stage register, and reads the cached microcode from the inter-stage register for processing to obtain a processing result of the corresponding instruction; In response to a request from the first execution module to stop sending instructions, stop sending instructions to the first execution module; wherein, the first execution module is an execution module in which the cached instructions in the head buffer reach a second cache threshold.

[0005] According to another aspect of the embodiments of the present application, a data processing device is provided, which is applied to an execution module in a processor. The execution module is a core or a vector processing unit (VPU). The device includes: An instruction receiving unit, configured to receive instructions matching the execution module in an issue queue (IQ), and cache the instructions in a head buffer; A processing unit, configured to read the cached instructions from the head buffer for decoding to obtain microcodes, cache the microcodes in an inter-stage register, read the cached microcodes from the inter-stage register for processing, and obtain a processing result of the corresponding instruction; A request unit, configured to stop reading the cached instructions from the head buffer when the microcodes cached in the inter-stage register reach a first cache threshold, and request the issue queue to stop sending instructions when the cached instructions in the head buffer reach a second cache threshold, so that the issue queue stops sending instructions to the execution module.

[0006] According to another aspect of the embodiments of the present application, a data processing device is provided, which is applied to an issue queue in a processor. The device includes: A summarizing unit, configured to summarize instructions; An instruction sending unit, configured to send instructions to an execution module matching the instructions. The execution module is a core or a VPU, so that the execution module caches the instructions in a head buffer, reads the cached instructions from the head buffer for decoding to obtain microcodes, caches the microcodes in an inter-stage register, reads the cached microcodes from the inter-stage register for processing, and obtains a processing result of the corresponding instruction; A stop sending unit, configured to stop sending instructions to the first execution module in response to a request from the first execution module to stop sending instructions; wherein, the first execution module is an execution module in which the cached instructions in the head buffer reach a second cache threshold.

[0007] According to another aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method provided in the first aspect or the second aspect above.

[0008] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in the first aspect or the second aspect are implemented.

[0009] According to one aspect of an embodiment of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the method provided in the first aspect or the second aspect are implemented.

[0010] The technical solution provided by the embodiment of the present application has the following beneficial effects: the execution module receives the instructions matching the execution module in the transmission queue and caches the instructions in the head buffer. Compared with the solution in which all the instructions in the transmission queue are received by the core core, the embodiment of the present application can distribute the instructions to different execution modules for processing earlier, the core core and the VPU work independently of each other, and the performance is improved. The execution module reads the cached instructions from the head buffer for decoding to obtain the microcode, caches the microcode to the inter-stage register, reads the cached microcode from the inter-stage register for processing, and obtains the processing result of the corresponding instruction. Since the core core only needs to decode the instructions matching itself, the timing of the core core in the DE stage is better and will not be interfered by the VPU. The pause signal of the transmission queue comes from the full state of the head buffer, and no complex logic drive is required. When the inter-stage register cache is large, only the buffer is stopped from reading instructions. Only when the head buffer cache is large, the transmission queue is requested to stop sending instructions, and the decoding process and the process of receiving instructions are decoupled, so that the timing of the transmission queue is also improved, and the performance of the execution module is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.

[0012] Figure 1 A hardware architecture diagram of a pipeline of a processor provided in an embodiment of the present application; Figure 2 A flowchart of a data processing method provided in an embodiment of the present application; Figure 3 A flowchart of a data processing method provided in an embodiment of the present application; Figure 4 A hardware architecture diagram of a pipeline of a processor provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application; Figure 6Schematic structural diagram of a data processing device provided by an embodiment of the present application; Figure 7 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0013] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0014] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an" and "the" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the art of the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0015] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0016] First, several terms related to the present application will be introduced and explained: In processor design, the issue queue stage (IQ stage), decode stage (DE stage), and execute stage (EX stage) are key parts in pipeline processing, where IQ Stage (Issue Queue Stage): It is mainly responsible for selecting eligible instructions from the issue queue and sending them to the Functional Unit (FU) for execution. The issue queue stores the instructions that have been renamed but not yet executed. The Issue queue is the core component of this stage, and its design affects the entire issue stage. The issue queue can be designed as centralized or distributed, corresponding to the relationship between the functional unit and the issue queue. In addition, the issue queue can also be divided into data capture type and non-data capture type according to at which stage in the pipeline the register values are read.

[0017] DE Stage (Decode Stage): It is mainly responsible for fetching instructions from the instruction memory and decoding them into operation codes and operands. The decoded instructions will be sent to the issue queue to wait for execution. In the decode stage, the processor will identify the type of the instruction (such as addition, subtraction, multiplication, etc.) and determine the positions of the operands. In addition, the decode stage may also involve operations such as register reading and instruction prefetching.

[0018] The EX Stage (Execution Stage) is a key part of the processor pipeline, responsible for actually executing instructions and completing data operations and processing. In the execution stage, the processor usually contains multiple computing units (such as the Arithmetic Logic Unit ALU, multiplier, etc.) for executing different types of instructions. For example, the execution stage (EX Stage) of the CV32E40P processor contains two modules, cv32e40p_alu and cv32e40p_mult, which are responsible for executing operations such as addition, subtraction, AND, OR, shift, etc. and multiplication operations respectively. In the execution stage, the processor will call the corresponding functional unit for calculation according to the type of the instruction and the operands. The calculation result will be written back to the register or used for the execution of subsequent instructions. In addition, the execution stage may also involve operations such as the processing of conditional branch instructions and the assignment of control signals.

[0019] The Vector Processing Unit (VPU) is a part of the Central Processing Unit (CPU) that implements a direct operation instruction set for one-dimensional arrays (vectors). Different from scalar processors that can only process one data at a time, the vector processing unit can process multiple data elements simultaneously, thus significantly improving the ability of data parallel processing. This ability has significant advantages in fields such as numerical simulation, scientific computing, and graphics processing.

[0020] Inter-stage register: It mainly involves data transfer and control between different levels or stages inside the processor. It is used to transfer data between different stages inside the processor, ensure the correctness and integrity of data, help the processor coordinate the work of different processing stages, ensure the smooth progress of the entire processing flow, and temporarily store intermediate results during the processing for subsequent processing stages to use.

[0021] Head Buffer is usually a data buffer used to temporarily store the data to be processed or transferred before data transfer or processing. It is located at the beginning of the data stream, so it is named "head buffer".

[0022] UOP generator: UOP (Micro-Operation) generator is short for micro-operation generator. In computer architecture, the UOP generator is responsible for decomposing high-level instructions (such as machine instructions) into a series of low-level micro-operations. These micro-operations are the basic operations that the processor can directly execute. The UOP generator can ensure that the processor executes instructions in an efficient and accurate manner.

[0023] Figure 1 This is a hardware architecture diagram of a pipeline of a processor provided by an embodiment of the present application. As shown in the figure, there is an issue queue in the IQ stage for storing instructions waiting to be executed; In the decode stage, a decoder and a uop generator shared by the core and the vpu are deployed in the core. A vpu decoder, which is a dedicated decoder for the vpu, is deployed in the vpu.

[0024] Inter-stage registers are deployed between the IQ stage and the DE stage, and between the DE stage and the EX stage. Its function is to store the data of the previous stage every cycle and drive the combinational logic of the next stage. For example, in cycle T0, the instruction is in the IQ stage, and in cycle T1, the instruction is stored from the IQ to the inter-stage register of the DE. In cycle T1, it can drive the decoder logic of the DE to perform decoding.

[0025] There is a stall source in each stage. When the pipeline finds that it cannot process an instruction and needs to pause, it will send a stall signal to a younger stage. For example, the EX stage will send a stall signal to the DE stage. When the DE stage receives the stall signal from the EX stage, it will send it to the IQ stage simultaneously or together. Figure 1 The data processing method shown includes the following steps: S101. Summarize the instructions into the IQ queue; S102. Sequentially send each instruction to the core according to the arrangement order of the instructions in the IQ queue; S103. The core decodes the received instruction to obtain microcode; S104. The core determines whether the microcode is required by itself or by the VPU. If it is required by itself, the core processes the microcode through its own computing unit. If it is required by the VPU, the core sends the microcode to the VPU; S105. If the VPU finds that it cannot handle the microcode, it notifies the core to stop. Stopping means stopping decoding and stopping sending microcode. At the same time, the core collects the stall signals from the EX stage of the core and the EX stage of the VPU in the DE stage and sends them to the IQ queue uniformly.

[0026] The following problems exist in this process: 1. The timing delay required in the DE stage of the core (that is, the stage of decoding, generating, and transmitting microcode in steps S103 - S104) is relatively long; 2. Within one clock cycle, all stall signals of the core and the VPU are returned to the IQ queue uniformly. The stall signals include the stall signals of the core in the EX stage, the stall signals of the VPU in the EX stage, and the stall signals of the core in the DE stage. The stall of the IQ is affected by the core and the VPU, resulting in poor performance of the IQ queue, and the timing also deteriorates accordingly after being affected.

[0027] 3. Because the core needs to collect the stall signals of the VPU to decide whether to stop sending instructions to the VPU, the core is affected by the stall signals of the VPU, resulting in poor timing in the DE stage of the core.

[0028] The data processing method, device, electronic device, computer-readable storage medium, and computer program product provided by this application aim to solve the above technical problems in the prior art.

[0029] The technical solutions of the embodiments of this application and the technical effects produced by the technical solutions of this application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0030] A data processing method is provided in the embodiments of this application, as Figure 2 shown. This method is applied to an execution module in a processor. The execution module refers to the core or the VPU. This method includes: S201. Receive the instructions in the transmit queue IQ that match the execution module, and cache the instructions in the head buffer.

[0031] It should be noted that, different from Figure 1 the problem that the processing logics of the core and the VPU are inconsistent in the shown embodiment, the embodiment of the present application has made an adjustment, and the processing logics of the core and the VPU are consistent. Moreover, different from Figure 1 the situation in the shown embodiment where the transmit queue sends all instructions to the core, the transmit queue of the present application will determine which execution module to send the instruction to according to whether the instruction matches the execution module. Thus, the core and the VPU will each only receive the instructions that match themselves, enabling the instructions in the transmit queue to be processed more promptly.

[0032] Furthermore, different from the embodiment where the core stores instructions by setting inter-stage registers in the IQ stage and the DE stage, the embodiment of the present application adopts a solution of setting a head buffer between the IQ stage and the DE stage of the execution module to replace the inter-stage register. On the one hand, the head buffer can dynamically adjust its size and content as needed to adapt to different data transmission requirements, while the inter-stage register is usually of a fixed size and has a more specialized function. Therefore, the head buffer has higher flexibility and adaptability compared to the inter-stage register. On the other hand, in terms of data transmission, the head buffer may help reduce the waiting time and improve the continuity and efficiency of data transmission. It can pre-allocate space before the data arrives and cache and forward the data during the data transmission process. The inter-stage register is mainly used to temporarily store data for quick access within the processor or between the processor and external devices.

[0033] S202. Read the cached instructions from the head buffer for decoding to obtain microcodes, cache the microcodes in the inter-stage register, and read the cached microcodes from the inter-stage register for processing to obtain the processing results of the corresponding instructions.

[0034] Regarding Figure 1 the problem that the core has poor timing due to the core completely decoding the instructions in Figure 1 and splitting the decoder of the core in

[0035] The execution module in the embodiment of the present application reads the cached instructions from the head buffer for decoding. The stage of obtaining the microcode is the DE stage. After obtaining the microcode, the microcode is stored in the inter-stage register between the DE stage and the EX stage. The execution module reads the cached microcode from the inter-stage register for processing to obtain the processing result of the corresponding instruction.

[0036] S203. When the microcode cached in the inter-stage register reaches the first cache threshold, stop reading the cached instructions from the head buffer. When the instructions cached in the head buffer reach the second cache threshold, request the issue queue to stop sending instructions so that the issue queue stops sending instructions to the execution module.

[0037] It should be noted that when the speed of the processor for processing the microcode is lower than the speed of caching the microcode, the microcode cached in the inter-stage register will gradually increase. When the inter-stage register reaches the first cache threshold, the embodiment of the present application stops reading the cached instructions from the head buffer to achieve the purpose of stopping decoding. When stopping reading the cached instructions from the head buffer, the processor is still obtaining instructions from the issue queue and storing them in the head buffer. Therefore, the number of instructions in the head buffer is still increasing. When the instructions cached in the head buffer reach the second cache threshold, the processor requests the issue queue to stop sending instructions. Thus, the stall signal of the IQ queue comes from the full state of the head buffer, and no complex logic drive is required, and the timing of the IQ queue is also better.

[0038] The data processing method provided by the embodiment of the present application is applied to the execution module in the processor. The execution module is the core core or the vector processing unit VPU. The execution module caches the instructions received from the issue queue that match the execution module into the head buffer. Compared with the solution where all the instructions in the issue queue are received by the core core, the embodiment of the present application can distribute the instructions to different execution modules for processing earlier. The core core and the VPU work independently without affecting each other, and the performance is improved. The execution module reads the cached instructions from the head buffer for decoding to obtain the microcode, caches the microcode in the inter-stage register, and reads the cached microcode from the inter-stage register for processing to obtain the processing result of the corresponding instruction. Since the core core only needs to decode the instructions that match itself, the timing of the core core in the DE stage is better and it will not be interfered by the VPU. The stall signal of the issue queue comes from the full state of the head buffer, and no complex logic drive is required. When there are more microcodes cached in the inter-stage register, only stop reading instructions from the buffer. Only when there are more instructions cached in the head buffer, request the issue queue to stop sending instructions, decoupling the decoding process and the instruction receiving process, making the timing of the issue queue better and improving the performance of the execution module.

[0039] Based on the above embodiments, as an alternative embodiment, the header buffer of the embodiments of the present application is a multi-entry structure. It can be understood that the header buffer may specifically refer to the buffer located at the beginning of the data structure, which is used to store metadata about the data structure itself or the data therein, or to temporarily store the data to be processed. The multi-entry structure further indicates that the data structure contains multiple entries, and these entries may have the same type or format and are stored in the data structure according to a certain order or rule to efficiently store and access a large amount of data.

[0040] The header buffer of the embodiments of the present application includes at least two entries, and at least two of the entries include two entries with fixed head pointers. The two entries with fixed head pointers respectively point to the first and second instructions in sequence. By operating in this way, on the one hand, by fixing the head pointers, the execution module can quickly locate the first and second instructions in sequence without traversing the entire header buffer. This can significantly improve the performance for the execution module that needs to frequently access the latest or oldest instructions. On the other hand, with fixed pointers, the insertion and deletion operations of instructions become simpler and more efficient. The execution module can easily add a new instruction after the latest instruction or delete an old instruction before the oldest instruction. On the other hand, since only two pointers need to be maintained instead of the complete index or linked list structure of the entire header buffer, memory space can be saved, which is particularly important for the execution module with limited resources.

[0041] Based on the above embodiments, as an alternative embodiment, the number of entries in the header buffer is 4. The determined 4 entries in the embodiments of the present application are the result of performance evaluation and area evaluation. If the core and VPU are too busy to process, they can be temporarily placed in the header buffer. In theory, the larger the header buffer, the better, but considering the limited area, it is set to 4, which can also meet the performance requirements.

[0042] Based on the above embodiments, as an alternative embodiment, the instruction that matches the execution module refers to an instruction for which the execution module has the ability to decode the encoding format of the instruction.

[0043] The issue queue of the present application determines the execution module that matches the instruction according to the encoding format of the instruction. For example, formats such as MMX, SSE, SSE2, SSE3, Sup-SSE3, SSE4.1, SSE4.2, etc. are suitable for core processing, while vector instruction sets, graphics processing instructions, etc. are suitable for VPU processing.

[0044] Based on the above embodiments, as an alternative embodiment, the execution module includes a decoder and a micro-operation generator, and the decoder and the micro-operation generator are configured to decode the cached instructions read from the head buffer to obtain microcodes.

[0045] The decoder in the embodiment of the present application is configured to parse an instruction to obtain a task to be executed, and the micro-operation generator is configured to generate specific operations of the task in each cycle.

[0046] Based on the above embodiments, as an alternative embodiment, when the cached instructions in the head buffer reach a second cache threshold, requesting the issue queue to stop sending instructions includes: In response to the cached instructions in the head buffer reaching the second cache threshold in the current cycle, requesting the issue queue to stop sending instructions in the next cycle.

[0047] It should be noted that when the cached instructions in the head buffer of the present application reach the second preset threshold in the current cycle, the issue queue is not requested to stop sending instructions in the current cycle, but in the next cycle, so that the performance is better than that of a single-entry inter-stage register in any case.

[0048] Based on the above embodiments, as an alternative embodiment, requesting the issue queue to stop sending instructions includes: Sending a first signal to the issue queue, where the first signal is used to indicate that the cached instructions in the head buffer reach the second cache threshold.

[0049] In the embodiment of the present application, by informing the issue queue that the cache of the head buffer has reached the threshold and prompting the issue queue to stop sending instructions in the way of informing the status information, the modification of the code of the processing logic of the issue queue can be reduced, and at the same time, the decoding logic and the logic driving of the computing unit are isolated.

[0050] Please refer to Figure 3 , which exemplarily shows the data processing method in the embodiment of the present application. The method is applied to the issue queue in a processor, and the method includes: S301. Aggregate instructions.

[0051] It should be noted that before the IQ stage of the processor, there are usually younger stages, such as: the IA stage and the lW stage. The instruction fetch unit IFU in the processor is responsible for loading (or accessing) the next instruction to be executed from the instruction cache or the main memory in the IA stage. IFU waits for the instruction loading to complete in the IW stage, and the loaded instruction will be sent by IFU to the issue queue.

[0052] S302. Send the instruction to the execution module that matches the instruction. The execution module is the core or the VPU, so that the execution module caches the instruction in the head buffer, reads the cached instruction from the head buffer for decoding to obtain microcode, caches the microcode in the inter-stage register, and reads the cached microcode from the inter-stage register for processing to obtain the processing result of the corresponding instruction.

[0053] It should be noted that, different from Figure 1 the problem that the processing logics of the core and the VPU are inconsistent in the shown embodiment, the embodiment of the present application has made adjustments, and the processing logics of the core and the VPU are consistent. Moreover, different from Figure 1 the shown embodiment where the issue queue sends all instructions to the core, the issue queue of the present application will determine which execution module to send the instruction to according to whether the instruction matches the execution module. Thus, the core and the VPU will each only receive instructions that match themselves, enabling the instructions in the issue queue to be processed more promptly.

[0054] Furthermore, different from the embodiment where the core stores instructions by setting inter-stage registers in the IQ stage and the DE stage, the embodiment of the present application adopts a solution of setting a head buffer between the IQ stage and the DE stage of the execution module to replace the inter-stage register. On the one hand, the head buffer can dynamically adjust its size and content as needed to adapt to different data transmission requirements, while the inter-stage register is usually of a fixed size and has a more specialized function. Therefore, the head buffer has higher flexibility and adaptability compared to the inter-stage register. On the other hand, in terms of data transmission, the head buffer may help reduce the waiting time and improve the continuity and efficiency of data transmission. It can pre-allocate space before the data arrives and cache and forward the data during the data transmission process. The inter-stage register is mainly used to temporarily store data for quick access within the processor or between the processor and external devices.

[0055] Regarding Figure 1 the problem that the core has poor timing due to the core completely decoding the instructions in Figure 1 the decoder of the core in

[0056] The execution module in the embodiment of the present application reads the cached instructions from the head buffer for decoding. The stage of obtaining the microcode is the DE stage. After obtaining the microcode, the microcode is stored in the inter-stage register between the DE stage and the EX stage. The execution module reads the cached microcode from the inter-stage register for processing to obtain the processing result of the corresponding instruction.

[0057] S303. In response to the request of the first execution module to stop sending instructions, stop sending instructions to the first execution module.

[0058] The first execution module in the embodiment of the present application is the execution module when the cached instructions in the head buffer reach the second cache threshold.

[0059] It should be noted that when the speed of the processor for processing microcode is lower than the speed of caching microcode, the microcode cached in the inter-stage register will gradually increase. When the inter-stage register reaches the first cache threshold, the present embodiment stops reading the cached instructions from the head buffer to achieve the purpose of stopping decoding. When stopping reading the cached instructions from the head buffer, the processor is still obtaining instructions from the issue queue and storing them in the head buffer. Therefore, the number of instructions in the head buffer is still increasing. When the cached instructions in the head buffer reach the second cache threshold, the processor requests the issue queue to stop sending instructions. Thus, the stall signal of the IQ queue comes from the full state of the head buffer, and no complex logic drive is required, and the timing of the IQ queue is also better.

[0060] The data processing method provided by the embodiment of the present application is applied to the issue queue in the processor. The issue queue aggregates instructions and sends the instructions to the execution module that matches the instructions. Compared with the scheme where all instructions in the issue queue are received by the core, the embodiment of the present application can distribute the instructions to different execution modules for processing earlier. The core and the VPU work independently without affecting each other, and the performance is improved. The execution module reads the cached instructions from the head buffer for decoding to obtain the microcode, caches the microcode in the inter-stage register, and reads the cached microcode from the inter-stage register for processing to obtain the processing result of the corresponding instruction. Since the core only needs to decode the instructions that match itself, the timing of the core in the DE stage is better and it will not be interfered by the VPU. The stall signal of the issue queue comes from the full state of the head buffer, and no complex logic drive is required. When there are more cached microcodes in the inter-stage register, only stop reading instructions from the buffer. Only when there are more cached instructions in the head buffer, request the issue queue to stop sending instructions, decouple the decoding process and the process of receiving instructions, make the timing of the issue queue better, and improve the performance of the execution module.

[0061] Based on the above embodiments, as an alternative embodiment, the instruction matching the execution module refers to an instruction for which the execution module has the ability to decode the encoding format of the instruction. The issue queue of the present application determines the execution module matching the instruction according to the encoding format of the instruction. For example, formats such as MMX, SSE, SSE2, SSE3, Sup-SSE3, SSE4.1, SSE4.2, etc. are applicable to core processing, while vector instruction sets, graphics processing instructions, etc. are suitable for VPU processing.

[0062] Based on the above embodiments, as an alternative embodiment, in response to a first execution module requesting to stop sending instructions, it includes: Receiving a first signal sent by the first execution module, where the first signal is used to indicate that the instructions cached in the head buffer reach a second cache threshold; Determining that the first execution module requests to stop sending instructions according to the first signal.

[0063] In the embodiment of the present application, by informing the issue queue that the cache of the head buffer has reached the threshold and prompting the issue queue to stop sending instructions by informing the status information, it is possible to reduce the modification of the code of the processing logic of the issue queue, and at the same time isolate the decoding logic and the logic driving of the computing unit.

[0064] Please refer to Figure 4 , which exemplarily shows a hardware architecture diagram of a pipeline of a processor provided by another embodiment of the present application. As shown in the figure, the IQ queue aggregates instructions. The IQ queue sends the instructions to the matching execution module according to the encoding format of the instructions. The execution module is a core or a VPU. The execution module caches the instructions in the head buffer at the DE stage. The decoder and the micro-operation generator read the cached instructions from the head buffer to perform decoding to obtain microcodes, cache the microcodes in the inter-stage register at the EX stage, read the cached microcodes from the inter-stage register for processing to obtain the processing results of the corresponding instructions. When the microcodes cached in the inter-stage register reach the first cache threshold, it triggers the pause source of the execution module at the EX stage. Further, the pause source at the EX stage triggers the pause source at the DE stage, and the decoder and the micro-operation generator stop reading the cached instructions from the head buffer. When the instructions cached in the head buffer reach the second cache threshold, the execution module sends a full status signal to the issue queue, prompting that the instructions cached in the head buffer reach the second cache threshold, so that the issue queue stops sending instructions to the execution module.

[0065] The embodiment of the present application provides a data processing device, which is applied to an execution module in a processor. The execution module is a core or a vector processing unit VPU, as Figure 5As shown in the figure, the data processing device may include: an instruction receiving unit 501, a processing unit 502, and a request unit 503. Among them, The instruction receiving unit 501 is configured to receive the instructions in the issue queue IQ that match the execution module, and cache the instructions into the head buffer; The processing unit 502 is configured to read the cached instructions from the head buffer for decoding to obtain microcodes, cache the microcodes into the inter-stage register, and read the cached microcodes from the inter-stage register for processing to obtain the processing results of the corresponding instructions; The request unit 503 is configured to stop reading the cached instructions from the head buffer when the microcodes cached in the inter-stage register reach the first cache threshold, and request the issue queue to stop sending instructions when the instructions cached in the head buffer reach the second cache threshold, so that the issue queue stops sending instructions to the execution module.

[0066] The device according to the embodiment of the present application can execute the data processing method on the execution module side provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the data processing method on the execution module side according to the embodiments of the present application. For the detailed function descriptions of the modules of the device, reference may specifically be made to the descriptions in the corresponding methods shown above, and details are not described herein again.

[0067] The embodiment of the present application provides a data processing device, which is applied to the issue queue in a processor, as Figure 6 As shown in the figure, the data processing device may include: a summarizing unit 601, an instruction sending unit 602, and a stop sending unit 603. Among them, The summarizing unit 601 is configured to summarize instructions; The instruction sending unit 602 is configured to send instructions to the execution module that matches the instructions. The execution module is a core or a VPU, so that the execution module caches the instructions into the head buffer, reads the cached instructions from the head buffer for decoding to obtain microcodes, caches the microcodes into the inter-stage register, and reads the cached microcodes from the inter-stage register for processing to obtain the processing results of the corresponding instructions; The stop sending unit 603 is configured to stop sending instructions to the first execution module in response to a request from the first execution module to stop sending instructions; Among them, the first execution module is the execution module whose cached instructions in the head buffer reach the second cache threshold.

[0068] The device of the embodiment of the present application can execute the data processing method on the transmission queue side provided by the embodiment of the present application, and the implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the data processing method on the transmission queue side of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, which will not be repeated here.

[0069] In an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory. The processor executes the above-mentioned computer program to implement the steps of the data processing method. Compared with the related art, it can be achieved that the execution module caches the instructions to the head buffer by receiving the instructions matching the execution module in the transmission queue. Compared with the solution that all the instructions in the transmission queue are received by the core, the embodiment of the present application can distribute the instructions to different execution modules for processing earlier, and the core and VPU work independently without affecting each other, so the performance is improved. The execution module reads the cached instructions from the head buffer to decode and obtain microcode, and caches the microcode. To the inter-stage register, read the cached microcode from the inter-stage register for processing, and obtain the processing result of the corresponding instruction. Since the core core only needs to decode the instructions that match itself, the timing of the core core in the DE stage is better, and it will not be interfered by the VPU. The pause signal of the transmit queue comes from the full state of the head buffer, and no longer requires complex logic drive. When the inter-stage register cache is large, only the buffer is stopped from reading instructions. Only when the head buffer cache is large, the transmit queue is requested to stop sending instructions, decoupling the decoding process and the process of receiving instructions, so that the timing of the transmit queue is also improved, and the performance of the execution module is improved.

[0070] In an alternative embodiment, an electronic device is provided, such as Figure 7 As shown, Figure 7 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0071] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0072] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 4002 is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.

[0073] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited here.

[0074] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0075] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0076] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0077] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown in the drawings or described in words.

[0078] It should be understood that although the flowcharts of the embodiments of the present application indicate each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0079] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also belong to the protection scope of the embodiments of the present application.

Claims

1. A data processing method, characterized in that: An execution module applied to a processor, wherein the execution module is a core or a vector processing unit VPU, and the method comprises: Receive instructions matching the execution module in a transmission queue IQ, and cache the instructions in a head buffer headbuffer; Reading the cached instructions from the head buffer to decode and obtain microcode, caching the microcode to the inter-stage register, reading the cached microcode from the inter-stage register to process, and obtaining the processing result of the corresponding instruction; When the microcode cached in the inter-level register reaches a first cache threshold, the reading of cached instructions from the head buffer is stopped; when the cached instructions in the head buffer reach a second cache threshold, the transmit queue is requested to stop sending instructions, so that the transmit queue stops sending instructions to the execution module.

2. The method according to claim 1, characterized in that The header buffer is a multi-entry structure, and the header buffer includes two entries with fixed head pointers, and the two entries point to the first and second instructions in sequence respectively.

3. The method according to claim 1, characterized in that The instructions matched with the execution module refer to instructions for which the execution module has the capability of decoding the encoding format of the instructions.

4. According to the method shown in claim 1, the execution module includes a decoder and a micro-operation generator, and the decoder and the micro-operation generator are used to decode the cached instructions read from the header buffer to obtain microcode.

5. The method according to claim 1, characterized in that When the instructions cached in the head buffer reach a second cache threshold, requesting the transmit queue to stop sending instructions comprises: In response to the cached instructions in the head buffer reaching a second cache threshold in a current cycle, the issue queue is requested to stop sending instructions in a next cycle.

6. The method according to claim 1, characterized in that The requesting the transmitting queue to stop sending instructions comprises: A first signal is sent to the transmit queue, the first signal being used to indicate that the instructions cached in the head buffer have reached a second cache threshold.

7. A data processing method, characterized in that: Applied to a transmission queue in a processor, the method comprises: Aggregate instructions; Sending the instruction to an execution module that matches the instruction, the execution module being a core or a VPU, so that the execution module caches the instruction in a head buffer, reads the cached instruction from the head buffer for decoding to obtain a microcode, caches the microcode in an inter-stage register, reads the cached microcode from the inter-stage register for processing, and obtains a processing result of the corresponding instruction; In response to a request from the first execution module to stop sending instructions, stop sending instructions to the first execution module; The first execution module is an execution module when the instructions cached in the head buffer reach a second cache threshold.

8. The method according to claim 7, characterized in that The instructions matched with the execution module refer to instructions for which the execution module has the capability of decoding the encoding format of the instructions.

9. The method according to claim 7, characterized in that: The step of stopping sending instructions in response to the first execution module requesting: receiving a first signal sent by the first execution module, where the first signal is used to indicate that the instructions cached in the header buffer have reached a second cache threshold; Determine, according to the first signal, that the first execution module requests to stop sending instructions.

10. A data processing device, characterized in that: An execution module applied to a processor, the execution module being a core or a vector processing unit VPU, the device comprising: An instruction receiving unit, used for receiving instructions matching the execution module in the transmission queue IQ, and caching the instructions into a head buffer; A processing unit, configured to read the cached instructions from the header buffer, decode and obtain microcodes, cache the microcodes to inter-stage registers, read the cached microcodes from the inter-stage registers for processing, and obtain processing results of corresponding instructions; A request unit is used to stop reading cached instructions from the head buffer when the microcode cached in the inter-stage register reaches a first cache threshold, and to request the transmit queue to stop sending instructions when the cached instructions in the head buffer reach a second cache threshold, so that the transmit queue stops sending instructions to the execution module.

11. A data processing device, characterized in that: The device is applied to a transmission queue in a processor, and comprises: A summary unit, used to summarize instructions; An instruction sending unit, used to send the instruction to an execution module matching the instruction, the execution module being a core or a VPU, so that the execution module caches the instruction to a header buffer, reads the cached instruction from the header buffer for decoding to obtain a microcode, caches the microcode to an inter-stage register, reads the cached microcode from the inter-stage register for processing, and obtains a processing result of the corresponding instruction; a stop sending unit, configured to stop sending instructions to the first execution module in response to a request by the first execution module to stop sending instructions; The first execution module is an execution module when the instructions cached in the head buffer reach a second cache threshold.

12. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the data processing method according to any one of claims 1 to 9.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 9 is implemented.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Finger fetching module and related equipment

    CN121300857A