Coprocessor, data processing method, processor system and electronic equipment
By designing a coprocessor with pipeline computing unit and cache module, the problem of low computational processing efficiency of coordinate rotation coprocessors is solved, and more efficient processing and mitigation of pipeline blocking is achieved.
Patent Information
- Application Number
- CN202510319911.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
The current coordinate rotation coprocessor is not efficient when processing multiple instructions at the same time.
A coprocessor is designed, including a control unit and an N-level pipeline calculation unit, and each stage pipeline calculation unit includes a pipeline module and a cache module. Computation instructions are processed through the pipeline iterative calculation, and in the case where predictions are blocked, the results are stored in the cache module to avoid pipeline blocking.
Improve the processing efficiency of the coprocessor, alleviate the situation of pipeline blockage, and maintain the order of calculation instructions.
Smart Images

Figure CN120144185A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to a coprocessor, a data processing method, a processor system, and an electronic device. Background Art
[0002] A coprocessor is a processor developed and applied to assist a Central Processing Unit (CPU) in completing processing tasks that it cannot execute or that have low execution efficiency and effectiveness, and is used to relieve the specific processing tasks of the system microprocessor.
[0003] A coordinate rotation coprocessor refers to a coprocessor with coordinate rotation calculation capabilities. Generally, a coordinate rotation coprocessor is implemented based on the Coordinate Rotation Digital Computer (CORDIC) algorithm.
[0004] However, the current coordinate rotation coprocessors have low arithmetic processing efficiency when processing multiple instructions simultaneously. Summary of the Invention
[0005] The present disclosure provides a coprocessor, a data processing method, a processor system, and an electronic device to solve the problem of low arithmetic processing efficiency of the coprocessor.
[0006] In a first aspect, the present disclosure provides a coprocessor,
[0007] comprising: a control unit and an N-stage pipelined computing unit connected in sequence, where N is an integer greater than 1; each stage of the pipelined computing unit includes a pipelined module and a cache module connected in sequence; and adjacent cache modules are connected to each other;
[0008] The control unit is configured to receive a calculation instruction sent by a main processor, parse the calculation instruction to obtain a mathematical operation and operands corresponding to the calculation instruction, and input the mathematical operation and operands corresponding to the calculation instruction into the N-stage pipelined computing unit in the order in which the calculation instruction is received;
[0009] The N-stage pipelined computing unit is configured to perform pipelined iterative calculation processing on the calculation instruction through N-stage pipelined modules to obtain a calculation result corresponding to the calculation instruction;
[0010] The control unit is further configured to, during the execution of the calculation instruction by the N-stage pipeline calculation unit, in the case where the result prediction output by the pipeline calculation unit is blocked: store the result output by the pipeline calculation unit with the prediction blocked in the cache module in the pipeline calculation unit with the prediction blocked; or store the result output by the pipeline calculation unit with the prediction blocked in the cache module in the next-stage pipeline calculation unit of the pipeline calculation unit with the prediction blocked; and send the calculation result corresponding to the calculation instruction to the main processor.
[0011] In a possible design, the control unit is further configured to:
[0012] Determine, in the reverse order from the last to the first of the pipeline calculation units, the predicted data path conditions of each pipeline calculation unit in the next clock cycle; the predicted data path conditions of the pipeline module are predicted to be blocked or not blocked; the prediction being blocked includes: prediction bypass or prediction backpressure;
[0013] In the next clock cycle, control the behavior of each pipeline calculation unit based on the predicted data path conditions of each pipeline calculation unit; the behavior of the pipeline calculation unit refers to the processing method of the result output by the pipeline module. Among them, in the case where the predicted data path condition of the pipeline module is prediction bypass, the behavior of the pipeline calculation unit includes storing the result output by the pipeline calculation unit in the cache module in the next-stage pipeline calculation unit of the pipeline calculation unit; in the case where the pipeline calculation unit predicts backpressure, the behavior of the pipeline calculation unit includes storing the result output by the pipeline calculation unit in the cache module in the pipeline calculation unit.
[0014] In a possible design, the control unit specifically is configured to:
[0015] Determine the predicted data path condition of the last-stage pipeline calculation unit in the next clock cycle according to the tag value of the last-stage pipeline calculation unit in the N-stage pipeline calculation unit and the current situation of the main processor receiving the calculation result; the tag value of the pipeline calculation unit refers to the number of clock cycles required to complete the calculation instruction in the pipeline calculation unit.
[0016] In the reverse order, starting from the (N - 1)-th pipeline computing unit, each pipeline computing unit in the N-stage pipeline computing unit is processed as follows: Determine the predicted data path situation of the pipeline computing unit in the next clock cycle according to the condition information of the pipeline computing unit; the condition information of the pipeline computing unit includes at least one of the following: the tag value of the pipeline computing unit, the data path situation of the next-stage pipeline computing unit of the pipeline computing unit in the next clock cycle, the arbitration result of the pipeline computing unit, and the information of the main processor currently receiving the calculation result; the arbitration result of the pipeline computing unit is used to indicate whether the result output by the pipeline computing unit in the next clock cycle is sent to the main processor.
[0017] In a possible design, the control unit is specifically configured to:
[0018] For the first (N - 1) pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit is successful, and the main processor currently indicates not to receive the calculation result, then determine that the predicted data path situation of the pipeline computing unit is predicted bypass; in the next clock cycle, store the result output by the pipeline computing unit in the cache module in the next-stage pipeline computing unit of the pipeline computing unit.
[0019] For the first (N - 1) pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit fails, and the predicted data path situation of the next-stage pipeline computing unit of the pipeline computing unit is not predicted backpressure, then determine that the predicted data path situation of the pipeline computing unit is predicted bypass; in the next clock cycle, store the result output by the pipeline computing unit in the cache module in the next-stage pipeline computing unit of the pipeline module.
[0020] For the first (N - 1) pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit fails, and the predicted data path situation of the next-stage pipeline computing unit of the pipeline computing unit is predicted backpressure, then determine that the predicted data path situation of the pipeline computing unit is predicted backpressure; in the next clock cycle, store the result output by the pipeline computing unit in the cache module in the pipeline computing unit.
[0021] For the first N - 1 levels of pipeline computing units, if the tag value of the pipeline computing unit is greater than zero and the predicted data path condition of the next - level pipeline computing unit of the pipeline computing unit is predicted backpressure, determine that the predicted data path condition of the pipeline computing unit is predicted backpressure; in the next clock cycle, store the result output by the pipeline computing unit in the cache module in the pipeline computing unit.
[0022] In a possible design, the control unit is specifically configured to:
[0023] For the first N - 1 levels of pipeline computing units, if the tag value of the pipeline computing unit is greater than zero and the next - level pipeline computing unit of the pipeline computing unit is predicted to be unblocked, determine that the predicted data path condition of the pipeline computing unit is predicted to be unblocked; in the next clock cycle, input the result output by the pipeline computing unit to the next - level pipeline computing unit of the pipeline computing unit.
[0024] For the first N - 1 levels of pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the pipeline module arbitration is successful, and the main processor currently indicates receiving the calculation result, determine that the predicted data path condition of the pipeline computing unit is predicted to be unblocked; in the next clock cycle, use the result output by the pipeline computing unit as the calculation result corresponding to the calculation instruction.
[0025] In a possible design, the control unit is specifically configured to:
[0026] For the last - level pipeline computing unit in the N - level pipeline computing unit, if the tag value of the last - level pipeline computing unit is zero and the main processor currently indicates that it can receive the calculation result, determine that the predicted data path condition in the next clock cycle of the last - level pipeline computing unit is predicted to be unblocked;
[0027] For the last - level pipeline computing unit in the N - level pipeline module, if the tag value of the last - level pipeline computing unit is zero and the main processor currently indicates that it cannot receive the calculation result, determine that the predicted data path condition in the next clock cycle of the last - level pipeline computing unit is predicted backpressure.
[0028] In a possible design, the control module is further configured to:
[0029] If there is a pipeline computing unit with a tag value of zero in the N-stage pipeline computing unit, obtain the last pipeline computing unit among the pipeline computing units with a tag value of zero as the first pipeline computing unit; and determine that the arbitration result of the other pipeline computing units in the N-stage pipeline computing unit except the first pipeline computing unit is arbitration failure;
[0030] If there is no valid data after the first pipeline computing unit in the pipeline computing unit, determine that the arbitration result of the first pipeline computing unit is arbitration success;
[0031] If there is valid data after the first pipeline computing unit in the pipeline computing unit, determine that the arbitration result of the first pipeline computing unit is arbitration failure.
[0032] In a possible design, the computing instruction is established based on the R-type instruction in the reduced instruction set RISC-V.
[0033] In a second aspect, the present disclosure provides a data processing method applied to a coprocessor. The coprocessor includes: a control unit and an N-stage pipeline computing unit connected in sequence, where N is an integer greater than 1; each stage of the pipeline computing unit includes a pipeline module and a cache module connected in sequence; and adjacent cache modules are connected to each other;
[0034] The control unit receives a computing instruction sent by the main processor;
[0035] The control unit parses the computing instruction to obtain the mathematical operation and operands corresponding to the computing instruction; and inputs the mathematical operation and operands corresponding to the computing instruction into the N-stage pipeline computing unit in the order of receiving the computing instruction;
[0036] The N-stage pipeline computing unit performs pipeline iterative calculation processing on the computing instruction through N-stage pipeline modules to obtain the calculation result corresponding to the computing instruction;
[0037] During the process of the N-stage pipeline computing unit executing the computing instruction, when the result prediction output by the pipeline computing unit is blocked, the control unit: stores the result output by the pipeline computing unit with the prediction blocked in the cache module corresponding to the pipeline computing unit with the prediction blocked; or stores the result output by the pipeline computing unit with the prediction blocked in the cache module corresponding to the next-stage pipeline computing unit of the pipeline computing unit with the prediction blocked;
[0038] The control unit sends the calculation result corresponding to the computing instruction to the main processor.
[0039] In some embodiments, the method further includes:
[0040] The control unit determines, in the order from the last to the first of the pipeline computing units, the predicted data path conditions of each pipeline computing unit in the next clock cycle; the predicted data path conditions are that the pipeline module is predicted to be blocked or not blocked; the predicted blocked includes: predicted bypass or predicted backpressure;
[0041] The control unit controls the behaviors of each pipeline computing unit based on the predicted data path conditions of each pipeline computing unit in the next clock cycle; the behavior of the pipeline computing unit refers to the processing method of the result output by the pipeline module. Among them, when the predicted data path condition of the pipeline module is predicted bypass, the behavior of the pipeline computing unit includes storing the result output by the pipeline computing unit in the cache module of the next-level pipeline computing unit of the pipeline computing unit; when the pipeline computing unit predicts backpressure, the behavior of the pipeline computing unit includes storing the result output by the pipeline computing unit in the cache module of the pipeline computing unit.
[0042] In some embodiments, the control unit determines, in the order from the last to the first of the pipeline computing units, the predicted data path conditions of each pipeline computing unit in the next clock cycle, including:
[0043] The control unit determines the predicted data path condition of the last pipeline computing unit in the next clock cycle according to the tag value of the last pipeline computing unit in the N-level pipeline computing unit and the current situation of the main processor receiving the calculation result; wherein, the tag value of the pipeline computing unit refers to the number of clock cycles required to complete the calculation instruction in the pipeline computing unit.
[0044] The control unit processes each pipeline computing unit in the N-level pipeline computing unit in order from the last to the first as follows: determines the predicted data path condition of the pipeline computing unit in the next clock cycle according to the condition information of the pipeline computing unit; wherein, the condition information of the pipeline computing unit includes at least one of the following: the tag value of the pipeline computing unit, the data path condition of the next-level pipeline computing unit of the pipeline computing unit in the next clock cycle, the arbitration result of the pipeline computing unit, and the information of the main processor currently receiving the calculation result; the arbitration result of the pipeline computing unit is used to indicate whether the result output by the pipeline computing unit is sent to the main processor in the next clock cycle.
[0045] In some embodiments, determining the predicted data path condition of the pipeline computing unit in the next clock cycle according to the condition information of the pipeline computing unit includes:
[0046] For the first N - 1 levels of pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit is successful, and the main processor currently indicates not to receive the calculation result, then determine that the predicted data path condition of the pipeline computing unit is predicted bypass;
[0047] In the next clock cycle, the control unit controls the behaviors of the respective pipeline computing units based on the predicted data path conditions of the respective pipeline computing units, including:
[0048] In the next clock cycle, the control unit stores the result output by the pipeline computing unit in the cache module of the next - level pipeline computing unit of the pipeline computing unit.
[0049] In some embodiments, determining the predicted data path condition of the pipeline computing unit in the next clock cycle according to the condition information of the pipeline computing unit includes:
[0050] For the first N - 1 levels of pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit fails, and the predicted data path condition of the next - level pipeline computing unit of the pipeline computing unit is not predicted backpressure, then determine that the predicted data path condition of the pipeline computing unit is predicted bypass;
[0051] In the next clock cycle, the control unit controls the behaviors of the respective pipeline computing units based on the predicted data path conditions of the respective pipeline computing units, including:
[0052] In the next clock cycle, the control unit stores the result output by the pipeline computing unit in the cache module of the next - level pipeline computing unit of the pipeline module.
[0053] In some embodiments, determining the predicted data path condition of the pipeline computing unit in the next clock cycle according to the condition information of the pipeline computing unit includes:
[0054] For the first N - 1 levels of pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit fails, and the predicted data path condition of the next - level pipeline computing unit of the pipeline computing unit is predicted backpressure, then determine that the predicted data path condition of the pipeline computing unit is predicted backpressure;
[0055] In the next clock cycle, the control unit controls the behaviors of each pipeline computing unit respectively based on the predicted data path conditions of each pipeline computing unit, including:
[0056] In the next clock cycle, the control unit stores the result output by the pipeline computing unit in the cache module in the pipeline computing unit.
[0057] In some embodiments, determining the predicted data path condition of the pipeline computing unit in the next clock cycle according to the condition information of the pipeline computing unit includes:
[0058] For the first N - 1 levels of pipeline computing units, if the tag value of the pipeline computing unit is greater than zero, and the predicted data path condition of the next - level pipeline computing unit of the pipeline computing unit is predicted backpressure, then determine that the predicted data path condition of the pipeline computing unit is predicted backpressure;
[0059] In the next clock cycle, the control unit controls the behaviors of each pipeline computing unit respectively based on the predicted data path conditions of each pipeline computing unit, including:
[0060] In the next clock cycle, the control unit stores the result output by the pipeline computing unit in the cache module in the pipeline computing unit.
[0061] In some embodiments, determining the predicted data path condition of the pipeline computing unit in the next clock cycle according to the condition information of the pipeline computing unit includes:
[0062] For the first N - 1 levels of pipeline computing units, if the tag value of the pipeline computing unit is greater than zero, and the next - level pipeline computing unit of the pipeline computing unit is predicted to be unblocked, then determine that the predicted data path condition of the pipeline computing unit is predicted to be unblocked;
[0063] In the next clock cycle, the control unit controls the behaviors of each pipeline computing unit respectively based on the predicted data path conditions of each pipeline computing unit, including:
[0064] In the next clock cycle, the control unit inputs the result output by the pipeline computing unit to the next - level pipeline computing unit of the pipeline computing unit.
[0065] In some embodiments, determining the predicted data path condition of the pipeline computing unit in the next clock cycle according to the condition information of the pipeline computing unit includes:
[0066] For the first N-1 levels of pipeline computing units, if the tag value of a pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline module is successful, and the main processor currently indicates receiving the calculation result, then it is determined that the predicted data path condition of the pipeline computing unit is predicted unblocked;
[0067] In the next clock cycle, the control unit controls the behaviors of the respective pipeline computing units based on the predicted data path conditions of the respective pipeline computing units, including:
[0068] In the next clock cycle, the control unit uses the result output by the pipeline computing unit as the calculation result corresponding to the calculation instruction.
[0069] In some embodiments, the control unit determines the predicted data path condition of the last pipeline computing unit in the N-level pipeline computing units in the next clock cycle according to the tag value of the last pipeline computing unit in the N-level pipeline computing units and the current situation of the main processor receiving the calculation result, including:
[0070] For the last pipeline computing unit in the N-level pipeline computing units, if the tag value of the last pipeline computing unit is zero and the main processor currently indicates that it can receive the calculation result, the control unit determines that the predicted data path condition of the last pipeline computing unit in the next clock cycle is predicted unblocked;
[0071] In the next clock cycle, the control unit controls the behaviors of the respective pipeline computing units based on the predicted data path conditions of the respective pipeline computing units, including:
[0072] In the next clock cycle, the control unit uses the result output by the last pipeline computing unit as the calculation result corresponding to the calculation instruction.
[0073] In some embodiments, the control unit determines the predicted data path condition of the last pipeline computing unit in the N-level pipeline computing units in the next clock cycle according to the tag value of the last pipeline computing unit in the N-level pipeline computing units and the current situation of the main processor receiving the calculation result, including:
[0074] For the last pipeline computing unit in the N-level pipeline module, if the tag value of the last pipeline computing unit is zero and the main processor currently indicates that it cannot receive the calculation result, the control unit determines that the predicted data path condition of the last pipeline computing unit in the next clock cycle is predicted backpressure;
[0075] In the next clock cycle, the control unit controls the behaviors of the respective pipeline computing units based on the predicted data path conditions of the respective pipeline computing units, including:
[0076] In the next clock cycle, the control unit stores the result output by the last-stage pipeline computing unit in the cache module in the last-stage pipeline computing unit.
[0077] In some embodiments, the method further includes:
[0078] If there is a pipeline computing unit with a tag value of zero in the N-stage pipeline computing unit, the control unit obtains the last pipeline computing unit among the pipeline computing units with a tag value of zero as the first pipeline computing unit.
[0079] The control unit determines that the arbitration results of the other pipeline computing units in the N-stage pipeline computing unit except the first pipeline computing unit are arbitration failures.
[0080] If there is no valid data after the first pipeline computing unit in the pipeline computing unit, the control unit determines that the arbitration result of the first pipeline computing unit is arbitration success.
[0081] If there is valid data after the first pipeline computing unit in the pipeline computing unit, the control unit determines that the arbitration result of the first pipeline computing unit is arbitration failure.
[0082] In some embodiments, the computing instruction is established based on the R-type instruction in the reduced instruction set RISC-V.
[0083] In a third aspect, the present disclosure provides a processor system, including a coprocessor and a main processor as described in any one of the above first aspects, and the main processor is communicatively connected to the coprocessor.
[0084] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the above second aspect is implemented.
[0085] In a fifth aspect, the present disclosure provides an electronic device, including the processor system as described in the above third aspect.
[0086] The coprocessor, data processing method, processor system, and electronic device provided by the embodiments of the present disclosure. The coprocessor establishes N-level pipeline computing units connected in sequence, and a corresponding cache module is connected after each pipeline computing unit. When any pipeline computing unit is about to be blocked, the result output by the pipeline computing unit is stored in its own cache module, or when the next-level cache module of the pipeline computing unit does not store data, the result output by the pipeline computing unit is stored in its next-level cache module. The coprocessor can perform processing with different numbers of iterations and can process instructions with different precision requirements. When the pipeline is about to be blocked, by temporarily storing the result output by the pipeline computing unit in the next-level cache module, when the subsequent pipeline computing units are not blocked, the result output by the blocked pipeline computing unit is stored backward in the idle cache module, freeing up storage space for the previous pipeline and alleviating the pipeline blockage situation. In addition, storing the result output by the pipeline computing unit backward maintains the order of the computing instructions in the original pipeline computing unit, realizes the order preservation of the computing instructions, and improves the processing efficiency of the coprocessor. Description of the Drawings
[0087] Figure 1 It is a schematic diagram of the architecture of a processor system provided by the embodiments of the present disclosure;
[0088] Figure 2 It is a schematic diagram of the structure of a coprocessor provided by the embodiments of the present disclosure;
[0089] Figure 3 It is a format of an R-type instruction provided by the embodiments of the present disclosure;
[0090] Figure 4 It is an arbitration schematic diagram of a pipeline computing unit provided by the embodiments of the present disclosure;
[0091] Figure 5 It is a flowchart of a data processing method provided by the embodiments of the present disclosure;
[0092] Description of the Reference Numerals:
[0093] 200: Coprocessor; 240: Control unit; 210: First-level pipeline computing unit; 220: Second-level pipeline computing unit; 230: Nth-level pipeline computing unit; 211: First-level pipeline module; 212: First-level cache module; 221: Second-level pipeline module; 222: Second-level cache module; 231: Nth-level pipeline module; 232: Nth-level cache module. Detailed Embodiments
[0094] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0095] In the present disclosure, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, or B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item(s) or plural item(s). For example, at least one (item) of a alone, b alone, or c alone may represent: a alone, b alone, c alone, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b, and c, where a, b, and c may be single or multiple. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0096] The orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "upper", "lower", "left", "right", "front", "rear", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present disclosure.
[0097] The terms "connected" and "coupled" should be understood in a broad sense. For example, the "connection" or "coupling" of a circuit structure may refer not only to a physical connection but also to an electrical connection or a signal connection. For example, it may be directly connected, that is, a physical connection, or indirectly connected through at least one intermediate element, as long as the circuit is connected. It may also be the internal connection of two elements; a signal connection may refer not only to a signal connection through a circuit but also to a signal connection through a media medium, such as radio waves. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific circumstances.
[0098] First, the names involved in the present disclosure will be explained.
[0099] Cordic is a numerical solution method similar to binary search. Cordic can gradually approximate the exact value through successive binary iterations. Cordic can implement operations such as trigonometric functions, hyperbolic functions, multiplication, division, exponentiation, and logarithm. In theory, each iteration doubles the precision, that is, one more significant digit is added. Exemplarily, for 32-bit integers / fixed-point numbers, 32 iterations are required to obtain an exact result within 32 bits. Taking the calculation of trigonometric functions as an example, in an embedded system with limited hardware or resources, Cordic recursively approximates the values of trigonometric functions through a series of continuous rotation operations with almost no multiplication and division operations.
[0100] A coprocessor is a chip used to relieve the system microprocessor of specific processing tasks. The coprocessor mentioned in this disclosure refers to a coprocessor with coordinate rotation calculation capabilities.
[0101] The coprocessor provided by this disclosure can be applied in a processor system. The following combines Figure 1 to introduce a processor system provided by this disclosure.
[0102] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the architecture of a processor system provided by an embodiment of this disclosure. As Figure 1 shown, the processor system includes a main processor and a coprocessor. The main processor can send calculation tasks to the coprocessor. The coprocessor processes the calculation tasks and returns the processing results to the main processor. Among them, the main processor can also be called the CPU, Figure 1 which is shown as the CPU in Figure 1 . Exemplarily, the coprocessor can be based on Cordic, and the coprocessor can also be called a Cordic coprocessor.
[0103] Since Cordic is based on multiple iterative calculations to obtain the calculation result, the coprocessor provided by the embodiments of the present disclosure can perform iterative calculations on calculation instructions in the form of pipeline calculation. Among them, the form of pipeline calculation means that the iterative calculation in Cordic is divided into multiple pipeline modules, and multiple pipeline modules can perform calculation processing in parallel at the same time. Thus, the coprocessor can process multiple calculation instructions simultaneously.
[0104] The present disclosure provides a coprocessor. An N-level pipeline calculation unit connected in sequence is established in the coprocessor, and a corresponding cache module is connected after each pipeline calculation unit. When any pipeline calculation unit is about to be blocked, the result output by the pipeline calculation unit is stored in its own cache module, or when the next-level cache module of the pipeline calculation unit does not store data, the result output by the pipeline calculation unit is stored in its next-level cache module. The coprocessor can perform different numbers of iterative processing and can process instructions with different precision requirements. When the pipeline is about to be blocked, by temporarily storing the result output by the pipeline calculation unit in the next-level cache module, when the subsequent pipeline calculation unit is not blocked, the result output by the blocked pipeline calculation unit is stored backward in the idle cache module, freeing up storage space for the previous pipeline and alleviating the situation of pipeline blockage. In addition, storing the result output by the pipeline calculation unit backward maintains the order of the calculation instructions in the original pipeline calculation unit, realizes the order preservation of the calculation instructions, and improves the processing efficiency of the coprocessor.
[0105] The following uses specific embodiments to illustrate the coprocessor provided by the present disclosure in detail.
[0106] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a coprocessor provided by an embodiment of the present disclosure. As Figure 2 shown, the coprocessor 200 provided in this embodiment may include a control unit 240 and an N-level pipeline calculation unit connected in sequence. The control unit 240 is respectively connected to each level of the pipeline calculation unit in the N-level pipeline calculation unit, and N is an integer greater than 1. Among them, Figure 2 exemplarily shows that the N-level pipeline calculation unit includes: a first-level pipeline calculation unit 210, a second-level pipeline calculation unit 220,..., and an N-level pipeline calculation unit 230.
[0107] Among them, each level of the pipeline calculation unit in the N-level pipeline calculation unit includes a pipeline module and a cache module connected in sequence. It can be understood that each level of cache module is connected to the next-level pipeline module, Figure 2Exemplarily shown in [the figure] is that the first-level pipelined computing unit 210 includes a first-level pipeline module 211 and a first-level cache module 212 connected in sequence, the second-level pipelined computing unit 220 includes a second-level pipeline module 221 and a second-level cache module 222 connected in sequence, ……, the Nth-level pipelined computing unit 230 includes an Nth-level pipeline module 231 and an Nth-level cache module 232 connected in sequence. The first-level pipeline module 211 is connected to the second-level pipeline module 221 through the first-level cache module 212, the second-level pipeline module 221 is connected to the third-level pipeline module through the second-level cache module 222, ……, the (N - 1)th-level pipeline module is connected to the Nth-level pipeline module 231 through the (N - 1)th-level cache module. Usually, the result obtained by the pipeline module can be stored in its corresponding cache module. The result output by the pipelined computing unit includes the result obtained by the pipeline module of itself through calculation, or the data stored in its own cache module.
[0108] Further, each pipeline module is based on Cordic.
[0109] Among them, every two adjacent cache modules are connected. In the present disclosure, the path between two connected cache modules can be referred to as a bypass. Figure 2 Exemplarily shown in [the figure] is that the first-level cache module 212 is connected to the second-level cache module 222. It can be understood that Figure 2 the following connection relationships not shown in [the figure] also exist: the second-level cache module 222 is connected to the third-level cache module, ……, the (N - 1)th-level cache module is connected to the Nth-level cache module 232.
[0110] Further, the coprocessor 200 may include 4-level pipelined computing units, and the number of iterations of each level of pipelined computing unit is the same.
[0111] That is, the N-level pipeline module includes: a first pipeline module, a second pipeline module, a third pipeline module, and a fourth pipeline module; the number of iterative calculations in the first pipeline module, the second pipeline module, the third pipeline module, and the fourth pipeline module is the same.
[0112] The cache module includes: a first cache module, a second cache module, a third cache module, and a fourth cache module.
[0113] The first pipeline module, the first cache module, the second pipeline module, the second cache module, the third pipeline module, the third cache module, the fourth pipeline module, and the fourth cache module are connected in sequence.
[0114] Further, 8 iterations can be performed in each level of pipeline module.
[0115] In this embodiment, a four-stage pipelined computing unit may be provided. Generally, a four-stage pipelined computing unit can meet various precision requirements in application scenarios.
[0116] The control unit 240 is configured to receive the computing instructions sent by the main processor.
[0117] Among them, the main processor is communicatively connected to the coprocessor 200. The main processor can send computing instructions to the coprocessor, and the control unit 240 receives the computing instructions sent by the main processor. The main processor may be the CPU in the processor system in the foregoing Figure 1 illustrated embodiment.
[0118] Furthermore, the computing instructions received from the main processor may be established based on the R-type instructions in RISC-V.
[0119] The control unit 240 may further be configured to: parse the computing instructions to obtain the mathematical operations and operands corresponding to the computing instructions, and input the mathematical operations and operands corresponding to the computing instructions into the N-stage pipelined computing unit in the order in which the computing instructions are received.
[0120] Exemplarily, please refer to Figure 3 , Figure 3 which is a format of an R-type instruction provided by an embodiment of the present disclosure. Among them, the func7, xd, xs1, and xs2 fields may be used for decoding, and rs1, rs2, and rd are used to transfer source operands and calculation results. In digital power applications, these mathematical functions are required to implement some matrix transformations such as rotation transformations. In the present disclosure, the R-type instructions are customized to determine the computing instructions.
[0121] Furthermore, the funct7 field may be defined to transfer mathematical operations (such as trigonometric functions, logarithms, and complex number operations). The computing instructions may be used to calculate mathematical functions of 32-bit fixed-point numbers, such as sine and cosine, exponential / logarithm, square root, etc.
[0122] RISC-V is an open-source instruction set architecture (ISA) that supports freely defining extended instruction sets. In combination with the coprocessor, custom extended instructions can be designed for dedicated mathematical operations (such as trigonometric functions, logarithms, and complex number operations).
[0123] The coprocessor interface of RISC-V can directly support the acceleration unit of Cordic and efficiently cooperate with the main processor through a new instruction or register access mechanism.
[0124] With the help of the RISC-V toolchain (such as GCC, LLVM, debugger, and simulator), the coprocessor can be called through standardized instructions and compiler toolchains without additional complex interface design.
[0125] The above N-stage pipeline computing unit is used to perform pipeline iterative computing processing on computing instructions through an N-stage pipeline module to obtain the computing result corresponding to the computing instruction.
[0126] The N-stage pipeline computing unit can perform pipeline iterative computing processing on computing instructions in the order of receiving the computing instructions, and obtain the computing result corresponding to the computing instruction. Among them, the control unit can parse the computing instruction to obtain the mathematical operation and operands and input them into the N-stage pipeline computing unit. The N-stage pipeline computing unit performs pipeline iterative computing processing on the computing instruction according to the mathematical operation and operands corresponding to the computing instruction. In the present disclosure, inputting the data operation and operands corresponding to the computing instruction into the N-stage pipeline computing unit is simply referred to as inputting the computing instruction into the N-stage pipeline computing unit.
[0127] After receiving each computing instruction, the computing instruction is input into the N-stage pipeline computing unit for pipeline iterative processing, and different computing instructions can be processed in parallel between the pipeline computing units. Among them, the iterative computing processing of the pipeline computing unit refers to sequentially processing the computing instruction according to the arrangement and connection order of the units in the pipeline computing unit. Among them, the computing process of each pipeline module itself is independent, and each pipeline module processes one computing instruction at a time, that is, each pipeline module does not process multiple computing instructions at the same time. Therefore, the computing processes of the computing instructions in the pipeline module do not affect each other.
[0128] For example, in the first clock cycle, the received computing instruction A is input into the pipeline computing unit, and the computing instruction A is first processed by the first-stage pipeline module 211. In the second clock cycle, the received computing instruction B is input into the pipeline computing unit, and the computing instruction B is first processed by the first-stage pipeline module 211, and the computing instruction A is processed by the second-stage pipeline module 221 at this time. It can be seen that when the computing instruction A has not been completed, the computing instruction B can also be processed at the same time, and the computing processes between the computing instruction A and the computing instruction B do not affect each other.
[0129] The above control unit 240 is used to, during the process of the pipeline computing unit performing computing, in the case where the result prediction output by the pipeline computing unit is blocked: store the result output by the pipeline computing unit with the prediction blocked in the cache module in the pipeline computing unit with the prediction blocked; or, store the result output by the pipeline computing unit with the prediction blocked in the cache module in the next-stage pipeline computing unit of the pipeline computing unit with the prediction blocked, so that the pipeline computing unit with the prediction blocked is not blocked.
[0130] Among them, the pipeline module and the cache module included in the pipeline computing unit are at the same level as the pipeline computing unit, and there is a one-to-one correspondence between the pipeline module and the cache module. The cache module corresponding to the pipeline module refers to the cache module that is connected to and after the pipeline module. Exemplarily, as Figure 2 the cache module corresponding to the first-level pipeline module 211 in Figure 2 is the first-level cache module 212. It can also be said that the first-level pipeline module 211 corresponds to the first-level cache module 212...., as
[0131] The control unit 240 can predict the next situation of the pipeline computing unit, so as to control the destination of the result output by the pipeline computing unit. Among them, due to reasons such as the main processor being unable to receive the calculation result currently, pipeline blocking may occur, that is, the result output by the pipeline computing unit is blocked. If the result output by the pipeline computing unit is predicted to be blocked, then next, the result output by the pipeline computing unit can be stored in the cache module in the pipeline computing unit. If it is predicted that there is no stored data in the next-level cache module, then next, the result output by the pipeline computing unit can be stored in the cache module in the next-level pipeline computing unit of the pipeline computing unit. Since adjacent two-level cache modules are connected, therefore, the result output by the pipeline computing unit can be transmitted to the next-level cache module through the cache module in the pipeline computing unit and stored in the next-level cache module.
[0132] Exemplarily, in the first clock cycle, the received calculation instruction A is input to the first-level pipeline computing unit 221, and the calculation instruction A is first processed by the first-level pipeline module 211.
[0133] In the second clock cycle, the received calculation instruction B is input into the pipeline calculation unit. The calculation instruction B first undergoes calculation processing through the first-stage pipeline module 211. At this time, the calculation instruction A undergoes calculation processing through the second-stage pipeline module 221. The control unit 240 predicts the results output by each pipeline in the third clock cycle: After passing through the second-stage pipeline module 221, the calculation instruction A is calculated and completed, and the calculation result corresponding to the calculation instruction A can be sent to the main processor; after passing through the first-stage pipeline module 211, the calculation instruction B is calculated and completed. Since the calculation result corresponding to the calculation instruction A needs to be sent to the main processor, at this time, the calculation result corresponding to the calculation instruction B, that is, the result output by the first-stage pipeline module 211, cannot be sent to the main processor, and the result output by the first-stage pipeline module 211 is blocked. At the same time, the second-level cache module 222 does not need to store the result output by the second-stage pipeline module 221. Therefore, it is predicted that the result output by the first-stage pipeline module 211 is stored in the second-level cache module 222.
[0134] In the third clock cycle, control the result output by the second-stage pipeline module 221, that is, the calculation result corresponding to the calculation instruction A, to be sent to the main processor. Control the result output by the first-stage pipeline module 211, that is, the calculation result corresponding to the calculation instruction B, to be stored in the second-level cache module 222.
[0135] The control unit 240 is also used to send the calculation result corresponding to the calculation instruction to the main processor.
[0136] Through the control of the results output by each pipeline module in the pipeline calculation unit by the control unit 240, it is possible to send the calculation results corresponding to the calculation instructions to the main processor in the order of receiving the calculation instructions. That is, for the calculation instructions received relatively earlier, their calculation results are also sent to the main processor relatively earlier.
[0137] In practical applications, when the main processor needs the coprocessor 200 to process a computing task, it generates a corresponding computing instruction and sends the computing instruction to the coprocessor 200. After receiving the computing instruction sent by the main processor, the control unit 240 inputs the computing instruction into the first-stage pipeline computing unit and continues to perform computing processing in the subsequent pipeline computing units according to the number of cycles required for the computing instruction. Since there are N-level pipeline computing units and each pipeline computing unit contains a cache module, during the pipeline processing, the control unit 240 predicts whether the results output by each pipeline computing unit will be blocked. If it is predicted that the result output by a certain pipeline computing unit will be blocked, then when the result is output by this pipeline computing unit, the control unit 240 temporarily stores the output result in the cache module of this pipeline computing unit, or when the cache module of the next-level pipeline computing unit of this pipeline computing unit is idle, the control unit 240 temporarily stores the output result in the cache module of the next-level pipeline computing unit of this pipeline computing unit. After the pipeline computing unit completes the calculation of the computing instruction, it outputs the corresponding calculation result, which is sent to the main processor by the control unit 240.
[0138] For the coprocessor provided in this embodiment, the computing instruction sent by the main processor is established based on the R-type instruction in the Reduced Instruction Set Computing (RISC-V), and can perform the transfer of computing instructions for dedicated mathematical operations. The coprocessor can efficiently cooperate with the main processor. The main processor can directly call the coprocessor through a custom RISC-V instruction without additional complex interface design. The coprocessor is based on Cordic to establish N-level pipeline computing units connected in sequence, and a corresponding cache module is connected after each pipeline computing unit. When any pipeline computing unit is about to be blocked, the result output by this pipeline computing unit is stored in its own cache module, or when the next-level cache module of this pipeline computing unit does not store data, the result output by this pipeline computing unit is stored in its next-level cache module. The coprocessor can perform processing with different numbers of iterations and can process instructions with different precision requirements. When the pipeline is about to be blocked, by temporarily storing the result output by the pipeline computing unit in the next-level cache module, when the subsequent pipeline computing units are not blocked, the result output by the blocked pipeline computing unit is stored backward in the idle cache module, freeing up storage space for the previous pipeline and alleviating the situation of pipeline blockage. In addition, storing the result output by the pipeline computing unit backward maintains the order of the computing instructions in the original pipeline computing unit, realizes the order preservation of the computing instructions, and improves the processing efficiency of the coprocessor.
[0139] In a possible design, the control unit 240 predicts the next situation of the pipeline computing unit, so as to control the destination of the result output by the pipeline computing unit in the pipeline computing unit. Specifically, it can be implemented in the following ways:
[0140] The control unit 240 determines the predicted data path situation of each pipeline computing unit in the next clock cycle in the reverse order from the back of the pipeline computing unit.
[0141] Among them, the predicted data path situation includes that the pipeline computing unit is predicted to be blocked or the pipeline computing unit is predicted not to be blocked. Being predicted to be blocked includes: predicted bypass or predicted backpressure. Predicted bypass means that in the next clock cycle, the result output by the pipeline computing unit is stored in the cache module corresponding to the next-level pipeline computing unit of the pipeline computing unit. Since the adjacent cache modules are connected, when it is predicted that no data needs to be stored in the next-level cache module, the result output by the pipeline computing unit at this level can be stored in the next-level cache module through bypass. Predicted backpressure means that in the next clock cycle, the result output by the pipeline computing unit is stored in the cache module corresponding to the pipeline computing unit. Since it is predicted that data will be stored in the next-level cache module, the result output by this level of pipeline computing unit cannot be stored in the next-level cache module, so the result output by the pipeline computing unit is stored in this level of cache module.
[0142] In the next clock cycle, the control unit 240 controls the behavior of each pipeline computing unit based on the predicted data path situation of each pipeline computing unit.
[0143] Among them, the behavior of the pipeline computing unit refers to the processing method of the result output by the pipeline computing unit. When the predicted data path situation of the pipeline computing unit is predicted bypass, the behavior of the pipeline computing unit includes storing the result output by the pipeline computing unit in the cache module corresponding to the next-level pipeline computing unit of the pipeline computing unit. When the pipeline computing unit predicts backpressure, the behavior of the pipeline computing unit includes storing the result output by the pipeline computing unit in the cache module corresponding to the pipeline computing unit.
[0144] In practical applications, at each clock cycle, the control unit 240 predicts the results output by each pipeline computing unit in the next clock cycle. Since the destination of the results output by the pipeline computing unit affects the results output by the previous-stage pipeline computing unit. For example, when the third-stage pipeline computing unit is blocked, the results output by this pipeline computing unit need to be stored in its corresponding cache module or the next-level cache module. For the second-stage pipeline computing unit, if the results it outputs are also blocked, the storage location of the results output by the third-stage pipeline computing unit affects whether the results output by the second-stage pipeline computing unit are stored in the second-stage cache module or the third-stage cache module. Therefore, the control unit 240 predicts the results output by each pipeline computing unit in the reverse order of the pipeline computing units from the back to the front.
[0145] When the control unit 240 predicts that the results output by a certain-stage pipeline computing unit will be blocked, if there is no need to store data in the next-level cache module of this pipeline computing unit, the predicted data path situation of this pipeline computing unit is predicted bypass. In the next clock cycle, the behavior that the control unit 240 can control this pipeline computing unit is: storing the results output by this pipeline computing unit in its next-level cache module through bypass.
[0146] When the control unit 240 predicts that the results output by a certain-stage pipeline computing unit will be blocked, if there is data stored in the next-level cache module of this pipeline computing unit, the predicted data path situation of this pipeline computing unit is predicted backpressure. In the next clock cycle, the behavior that the control unit 240 can control this pipeline computing unit is: storing the results output by this pipeline computing unit in its corresponding cache module.
[0147] In this embodiment, the control unit 240 determines the predicted data path situations of each pipeline computing unit in the next clock cycle in the reverse order of the pipeline computing units from the back to the front. In the next clock cycle, the control unit 240 controls the behaviors of each pipeline computing unit based on the predicted data path situations of each pipeline computing unit. Thus, the process of the pipeline computing unit executing the computing task is reasonably managed and controlled, enabling the pipeline computing to proceed orderly. When the pipeline is about to be blocked, by temporarily storing the results output by the pipeline computing unit in the next-level cache module, when the subsequent pipeline computing units are not blocked, the results output by the blocked pipeline computing unit are stored backward in the idle cache module, thereby freeing up storage space for the previous pipeline and alleviating the situation of pipeline blockage. In addition, storing the results output by the pipeline computing unit backward maintains the order of the computing instructions in the original pipeline computing unit, realizes the order preservation of the computing instructions, and improves the processing efficiency of the coprocessor.
[0148] In a possible design, when the control unit 240 predicts the data path conditions of each pipeline computing unit in the next clock cycle, it can determine the predicted data path conditions of each pipeline computing unit based on the number of calculation cycles that the calculation instructions in each pipeline computing unit still need to perform, and the situation of the output results already calculated by the calculation instructions in each pipeline computing unit being sent to the main processor. The following will be described in detail with specific embodiments.
[0149] The control unit 240 is specifically configured to determine the predicted data path conditions of the last pipeline computing unit in the next clock cycle according to the tag value of the last pipeline computing unit in the N-stage pipeline computing unit and the current situation of the main processor receiving the calculation results.
[0150] The tag value of the pipeline computing unit refers to the number of clock cycles required to complete the calculation instruction currently executed by the pipeline computing unit. Each stage of the pipeline computing unit can complete the calculation of the pipeline computing unit in one calculation cycle, and this calculation cycle is equal to the clock cycle. For example, if the calculation instruction C needs to be calculated by 3 pipeline computing units to obtain the calculation result, then it takes 3 calculation cycles to complete the calculation instruction C, that is, it takes 3 clock cycles to complete the calculation instruction C. When the calculation instruction C is input to the first-stage pipeline computing unit, the tag value of the first-stage pipeline computing unit is 3. After being processed by the first-stage pipeline computing unit, the output result tag value is decreased by 1, that is, 3 - 1 = 2. The result output by the first-stage pipeline computing unit is input to the second-stage pipeline computing unit, and at the same time, the tag value of the second-stage pipeline computing unit is assigned as 2. After being processed by the second-stage pipeline computing unit, the output result tag value is decreased by 1, that is, 2 - 1 = 1. The result output by the second-stage pipeline computing unit is input to the third-stage pipeline computing unit, and at the same time, the tag value of the third-stage pipeline computing unit is assigned as 1. The tag value of the third-stage pipeline computing unit is 1, so in the next clock cycle, the output result tag value of the third-stage pipeline computing unit will be assigned as 0, indicating that in the next clock cycle, the output result of the third-stage pipeline computing unit is the calculation result of the calculation instruction C.
[0151] The control unit 240 is specifically configured to start from the (N - 1)th pipeline computing unit in reverse order and perform the following processing on each stage of the pipeline computing unit in the N-stage pipeline computing unit respectively: determine the predicted data path conditions of the pipeline computing unit in the next clock cycle according to the condition information of the pipeline computing unit.
[0152] Among them, the conditional information of the pipeline computing unit may include, but is not limited to, at least one of the following: the tag value of the pipeline computing unit, the data path condition of the next-level pipeline computing unit of the pipeline computing unit in the next clock cycle, the arbitration result of the pipeline computing unit, and the information of the main processor currently receiving the calculation result.
[0153] Among them, the tag value of the pipeline computing unit can indicate whether the calculation instruction executed in the current pipeline computing unit is calculated completed. Therefore, the tag value of the pipeline computing unit will affect the behavior of the pipeline computing unit.
[0154] Based on the above discussion, the data path condition of the next-level pipeline computing unit of the pipeline computing unit in the next clock cycle may affect the behavior of the pipeline computing unit.
[0155] Among them, the arbitration result of the pipeline computing unit is used to indicate whether the result output by the pipeline computing unit is sent to the main processor in the next clock cycle. The arbitration result may include arbitration success or arbitration failure. The arbitration result of the pipeline computing unit indicating arbitration success means that the result output by the pipeline computing unit is sent to the main processor. The arbitration result of the pipeline computing unit indicating arbitration failure means that the result output by the pipeline computing unit is not sent to the main processor. In the next clock cycle, when there is a result output by the pipeline computing unit that needs to be sent to the main processor, that is, when the tag value of the pipeline computing unit is 0, it is necessary to determine whether the result output by the pipeline computing unit can be sent to the main processor. For example, in the next clock cycle, if the tag values of multiple pipeline computing units are 0, then arbitration needs to be performed on these multiple pipeline computing units, that is, to determine whether each pipeline computing unit can send its own output result to the main processor. Among them, the arbitration result of the pipeline computing unit determined to be sent to the main processor is arbitration success, and the arbitration result of the pipeline computing unit determined not to be sent to the main processor is arbitration failure. Another example, in the next clock cycle, there is only one pipeline computing unit with a tag value of 0. However, the pipeline computing unit behind this pipeline computing unit still has valid data, that is, the pipeline computing unit behind this pipeline computing unit is still performing calculation processing. Then, in order to ensure the order, the result output by this pipeline computing unit cannot be sent to the main processor either, so the arbitration result of this pipeline computing unit is arbitration failure.
[0156] Among them, the information of the main processor currently receiving the calculation result is used to indicate whether the main processor can currently receive the calculation result. For example, the main processor can Figure 1 through the handshake channel as in, output a high level when it can receive the calculation result, and output a low level when it cannot receive the calculation result.
[0157] In practical applications, in the current clock cycle, the control unit 240 starts from the last-stage pipeline computing unit and predicts the data path conditions of each pipeline computing unit in the reverse order. For the last-stage pipeline computing unit, since it is already the last stage and there is no subsequent pipeline computing unit, the result output by the last-stage pipeline computing unit is the calculation result of the calculation instruction it executes. Moreover, if the last-stage pipeline computing unit is blocked, bypass processing cannot be performed either. Therefore, in the next clock cycle, the behavior of the last-stage pipeline computing unit can be to send the calculation result to the main processor or to store the output result in the last-level cache module. For other pipeline computing units except the last-stage pipeline computing unit, the prediction is performed step by step in the reverse order of the pipeline computing units: according to the condition information of each pipeline computing unit, the predicted data path condition of the pipeline computing unit is determined.
[0158] In this embodiment, starting from the last-stage pipeline computing unit forward, according to the factors affecting the data path of the pipeline computing unit, the behavior prediction of the pipeline computing unit is performed step by step. The factors affecting the data path of the pipeline computing unit covered are relatively comprehensive, and the behavior prediction of the pipeline computing unit in the next clock cycle can be carried out orderly and effectively, reasonably controlling the process of the pipeline computing unit executing the computing task, making the pipeline computing proceed orderly. When the pipeline is about to be blocked, by temporarily storing the result output by the pipeline computing unit in the next-level cache module, when the subsequent pipeline computing unit is not blocked, the result output by the blocked pipeline computing unit is stored backward in the idle cache module, thereby freeing up storage space for the previous pipeline and alleviating the pipeline blockage situation. In addition, storing the result output by the pipeline computing unit backward maintains the order of the calculation instructions in the original pipeline computing unit, realizes the order preservation of the calculation instructions, and improves the processing efficiency of the coprocessor.
[0159] In a possible design, during the process of the control module predicting the N-level pipeline computing unit, the arbitration result of each pipeline module can be determined first. The following is an exemplary introduction to a method for the control module to perform arbitration.
[0160] The control module is further configured to:
[0161] If there is a pipeline computing unit with a tag value of zero in the N-level pipeline computing unit, obtain the last pipeline computing unit among the pipeline computing units with a tag value of zero as the first pipeline computing unit; and determine that the arbitration result of other pipeline computing units in the N-level pipeline computing unit except the first pipeline computing unit is arbitration failure.
[0162] If there is no valid data after the first pipeline computing unit in the pipeline computing unit, determine that the arbitration result of the first pipeline computing unit is arbitration success.
[0163] If there is valid data after the first pipeline computing unit in the pipeline computing unit, determine that the arbitration result of the first pipeline computing unit is arbitration failure.
[0164] In practical applications, one or more pipeline computing units participate in arbitration simultaneously. If there is no valid data after the data of the pipeline computing unit with the last position, it means that the computing instruction executed in this pipeline computing unit is the first computing instruction received in the pipeline computing unit currently, and its corresponding computing result should be sent to the main processor before the computing results corresponding to other computing instructions. Then, the arbitration result of this level of pipeline data is arbitration success. Correspondingly, the arbitration results of other pipeline computing units are arbitration failure. If there is still valid data after the pipeline computing unit with the last position, it means that there are still computing instructions in the current pipeline computing unit that have not been sent to the main processor before receiving the computing instruction executed in this pipeline computing unit. Therefore, the arbitration result of this pipeline computing unit is arbitration failure, and the arbitration results of other pipeline computing units are also arbitration failure. It should be noted that in this case, the arbitration results of all pipeline computing units participating in arbitration are arbitration failure.
[0165] Further, an effective bit can be set for each pipeline computing unit to indicate whether the data in this pipeline computing unit is valid data. When the control unit 240 performs arbitration, it can quickly determine whether there is valid data after the data of the pipeline computing unit with the last position according to the effective bit of the current pipeline computing unit. Exemplarily, if the effective bit of the pipeline computing unit is 0, it can indicate that there is no valid data in this pipeline computing unit; if the effective bit of the pipeline computing unit is 1, it can indicate that there is valid data in this pipeline computing unit.
[0166] In this embodiment, by determining whether there is valid data after the last pipeline computing unit participating in arbitration, a reasonable and simple arbitration method is given, which can ensure the order of pipeline computing results.
[0167] In some scenarios, when the control unit 240 predicts the data path situation of the N-level pipeline computing unit, it is specifically determined according to different conditions. The following uses specific embodiments to elaborate on the data path situations determined under different conditions.
[0168] First, the prediction situation of the last-level pipeline computing unit is described.
[0169] In a possible design, the control unit 240 is specifically configured to, for the last pipeline computing unit in the N-stage pipeline computing unit, if the tag value of the last pipeline computing unit is zero and the main processor currently indicates that it can receive the calculation result, determine that the predicted data path situation in the next clock cycle of the last pipeline computing unit is predicted to be unblocked.
[0170] In practical applications, when the control unit 240 makes a prediction for the last pipeline computing unit, if the tag value of the last pipeline computing unit is zero, it indicates that the calculation instruction executed in the last pipeline computing unit has been calculated and needs to be sent to the main processor. If the main processor currently indicates that it can receive the calculation result, it is determined that the predicted data path situation in the next clock cycle of the last pipeline computing unit is predicted to be unblocked. In the present disclosure, this situation where the prediction is unblocked can also be referred to as predicted output. In the next clock cycle, the result output by the last pipeline computing unit is used as the calculation result corresponding to the calculation instruction, and the sending module sends the calculation result corresponding to the calculation instruction to the main processor.
[0171] In a possible design, the control unit 240 is specifically configured to, for the last pipeline computing unit in the N-stage pipeline computing unit, if the tag value of the last pipeline computing unit is zero and the main processor currently indicates that it cannot receive the calculation result, determine that the predicted data path situation in the next clock cycle of the last pipeline computing unit is predicted backpressure.
[0172] In practical applications, when the control unit 240 makes a prediction for the last pipeline computing unit, if the tag value of the last pipeline computing unit is zero, it indicates that the calculation instruction executed in the last pipeline computing unit has been calculated and needs to be sent to the main processor. If the main processor currently indicates that it cannot receive the calculation result, it is determined that the predicted data path situation in the next clock cycle of the last pipeline computing unit is predicted backpressure. In the next clock cycle, the result output by the last pipeline computing unit is stored in the last-level cache module.
[0173] The prediction situation of the first N-1 pipeline computing units is introduced below.
[0174] In a possible design, for the first N-1 pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit is successful, and the main processor currently indicates that it does not receive the calculation result, then it is determined that the predicted data path situation of the pipeline computing unit is predicted bypass; in the next clock cycle, the result output by the pipeline computing unit is stored in the cache module in the next-level pipeline computing unit of the pipeline computing unit.
[0175] In a possible design, for the first N - 1 levels of pipeline computing units, if the tag value of a pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit fails, and the predicted data path condition of the next - level pipeline computing unit of the pipeline computing unit is not predicted backpressure, then determine that the predicted data path condition of the pipeline computing unit is predicted bypass; in the next clock cycle, store the result output by the pipeline computing unit in the cache module in the next - level pipeline computing unit of the pipeline module.
[0176] In a possible design, for the first N - 1 levels of pipeline computing units, if the tag value of a pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit fails, and the predicted data path condition of the next - level pipeline computing unit of the pipeline computing unit is predicted backpressure, then determine that the predicted data path condition of the pipeline computing unit is predicted backpressure; in the next clock cycle, store the result output by the pipeline computing unit in the cache module in the pipeline computing unit.
[0177] In a possible design, for the first N - 1 levels of pipeline computing units, if the tag value of a pipeline computing unit is greater than zero, and the predicted data path condition of the next - level pipeline computing unit of the pipeline computing unit is predicted backpressure, then determine that the predicted data path condition of the pipeline computing unit is predicted backpressure; in the next clock cycle, store the result output by the pipeline computing unit in the cache module in the pipeline computing unit.
[0178] In a possible design, for the first N - 1 levels of pipeline computing units, if the tag value of a pipeline computing unit is greater than zero, and the next - level pipeline computing unit of the pipeline computing unit is predicted not to be blocked, then determine that the predicted data path condition of the pipeline computing unit is predicted not to be blocked; in the next clock cycle, input the result output by the pipeline computing unit to the next - level pipeline computing unit of the pipeline computing unit.
[0179] In a possible design, for the first N - 1 levels of pipeline computing units, if the tag value of a pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline module is successful, and the main processor currently indicates receiving the calculation result, then determine that the predicted data path condition of the pipeline computing unit is predicted not to be blocked; in the next clock cycle, use the result output by the pipeline computing unit as the calculation result corresponding to the calculation instruction.
[0180] In practical applications, for each computing unit in the first N - 1 levels of pipeline computing units, the control unit can make predictions through the following steps 1 - step 11 and control the behavior of the pipeline computing unit in the next clock cycle.
[0181] Step 1: Determine whether the tag value of the pipeline computing unit is greater than zero.
[0182] If so, proceed to Step 2; if not, proceed to Step 5.
[0183] If the tag value of the pipeline computing unit is zero, it indicates that the data of the current pipeline computing unit needs to be sent to the main processor. If the tag value of the pipeline computing unit is greater than zero, it indicates that the data of the current pipeline computing unit has not been calculated yet.
[0184] Step 2: Determine whether the predicted data path condition of the next-level pipeline computing unit of the pipeline computing unit is predicted to be unblocked.
[0185] If so, proceed to Step 3; if not, proceed to Step 4.
[0186] If the predicted data path condition of the next-level pipeline computing unit of the pipeline computing unit is predicted to be unblocked, it indicates that the data of the pipeline computing unit can be input into the next-level pipeline computing unit for continued calculation.
[0187] If the predicted data path condition of the next-level pipeline computing unit of the pipeline computing unit is predicted to be blocked, it indicates that the data of the pipeline computing unit cannot be input into the next-level pipeline computing unit for continued calculation.
[0188] Step 3: Determine that the predicted data path condition of the last-level pipeline computing unit in the next clock cycle is predicted to be unblocked.
[0189] In the next clock cycle, input the result output by the pipeline computing unit into the next-level pipeline computing unit of the pipeline computing unit.
[0190] Step 4: Determine that the predicted data path condition of the pipeline computing unit is predicted to have backpressure.
[0191] In the next clock cycle, store the result output by the pipeline computing unit in the cache module in the pipeline computing unit.
[0192] Step 5: Determine whether the output arbitration result of the pipeline computing unit indicates arbitration success.
[0193] If so, proceed to Step 6; if not, proceed to Step 9.
[0194] If the output arbitration result of the pipeline computing unit indicates arbitration success, then in the next clock cycle, the data of the current pipeline computing unit needs to be sent to the main processor.
[0195] If the output arbitration result of the pipeline computing unit indicates arbitration success, then in the next clock cycle, the data of the current pipeline computing unit cannot be sent to the main processor, and prediction backpressure or prediction bypass is performed.
[0196] Step 6, determine whether the current main processor currently indicates receiving the calculation result.
[0197] If so, continue to execute Step 7; if not, continue to execute Step 8.
[0198] If the current main processor indicates receiving the calculation result, then in the next clock cycle, the data of this pipeline computing unit can be sent to the main processor.
[0199] If the current main processor indicates not receiving the calculation result and the current arbitration is successful, indicating that there is no valid data in the follow-up, then it can be determined that this pipeline computing unit is a prediction bypass, and in the next clock cycle, the data of this pipeline computing unit is stored in the next-level cache module.
[0200] Step 7, determine that the predicted data path situation of the pipeline computing unit is predicted not to be blocked.
[0201] In this case, it can also be called a predicted output in the present disclosure. In the next clock cycle, the result output by the pipeline computing unit is used as the calculation result corresponding to the calculation instruction. The calculation result is sent to the main processor through the sending module.
[0202] Step 8, determine that the predicted data path situation of the pipeline computing unit is a prediction bypass.
[0203] In the next clock cycle, the result output by the pipeline computing unit is stored in the cache module in the next-level pipeline computing unit of the pipeline computing unit.
[0204] Step 9, determine whether the next-level pipeline computing unit is blocked.
[0205] If so, continue to execute Step 10; if not, continue to execute Step 11.
[0206] Step 10, determine that the predicted data path situation of the pipeline computing unit is a prediction backpressure.
[0207] In the next clock cycle, the result output by the pipeline computing unit is stored in the cache module in the pipeline computing unit.
[0208] Step 11, determine that the predicted data path situation of the pipeline computing unit is a prediction bypass.
[0209] In the next clock cycle, the result output by the pipeline computing unit is stored in the cache module in the next-level pipeline computing unit of the pipeline computing unit.
[0210] In the above embodiments, the prediction and behavior control of the pipeline computing unit in different situations are described, so that the behavior of the pipeline computing unit is determined in the current clock cycle. Thus, when the next clock cycle arrives, the behavior of the pipeline computing unit is controlled to make the pipeline computing proceed in an orderly manner.
[0211] Taking the co-processor with a 4-stage pipeline computing unit as an example below, the tag values and valid bits mentioned above in the present disclosure are described.
[0212] For simplicity of description below, the pipeline computing unit is abbreviated as the pipeline.
[0213] In practical applications, users have different requirements for computing accuracy. Therefore, the number of computing cycles for corresponding computing instructions to execute can be different. In the present disclosure, the computing instruction is abbreviated as the instruction. Therefore, the number of cycles for different instructions to execute can be 1, 2, 3, or 4. In the case where the instruction cycle changes, there will be a situation where the results of two instructions are calculated and completed simultaneously, which will involve arbitration of the output results, and the result of the instruction that comes first is preferentially output.
[0214] For example: If the number of execution cycles of instruction A in the first clock cycle is 4 and the number of execution cycles of instruction B in the second clock cycle is 3, then these two instructions will be calculated and completed simultaneously in the fifth clock cycle. Therefore, the result of instruction B needs to be blocked for 1 clock cycle before being output.
[0215] However, if the result of instruction B is blocked in the third-stage pipeline computing unit, the data operations in the front of the pipeline will also be blocked accordingly, which will result in loss of pipeline efficiency. Therefore, the result of instruction B needs to be bypassed to the fourth-stage pipeline computing unit, so as to free up the space in the cache module in the third-stage pipeline computing unit for the data operations in the front. This also ensures that the data in the pipeline advances at least 1 pipeline computing unit per cycle. Therefore, a tag value is set for each stage of the pipeline. The tag value is used to record how many more times the data in this stage of the pipeline needs to be calculated before it can be output. Among them, the tag value of the first-stage pipeline is determined by the number of computing cycles of the instruction executed therein, and the tag values of other stages of the pipeline are assigned by the previous pipeline. Please refer to Table 1, which gives an explanation of the changes in the tag values of each stage of the pipeline.
[0216] Table 1. Updates of Pipeline Behavior, Tags, and Valid Bits
[0217]
[0218] For valid bits, since not every piece of data in the pipeline is valid data. For example, the situation of not being valid data can include empty data, or the data is about to be output. A valid bit can be set for each stage of the pipeline. When the pipeline is back-pressured, if the data at a certain stage is invalid, the data at the previous stage is allowed to perform forward operations to reach this stage. Please continue to refer to Table 1 above, which gives an explanation of the changes in the valid bits for each stage of the pipeline.
[0219] When the tag values of multiple stages of the pipeline reach 0 simultaneously, that is, the data on the pipeline all meet the output conditions. At this time, arbitration needs to be performed on these output data to determine the data that can be output in the next clock cycle. When performing arbitration on the pipeline, the arbitration result can be quickly determined according to the valid bits. If there is no valid data after the data with the last position, then the arbitration of the data output of this stage of the pipeline is successful; otherwise, the arbitration fails.
[0220] It should be noted that when multiple pipeline data participate in arbitration simultaneously, a situation where all arbitrations fail may occur.
[0221] Exemplarily, please refer to Figure 4 , Figure 4 which is an arbitration schematic diagram of a pipeline computing unit provided by an embodiment of the present disclosure. Figure 4 Exemplarily shown on the left is the execution situation of 3 instructions corresponding to each clock cycle. Among them, the cycle number represents the stage at which the instruction is executed at the corresponding clock, and the information in the parentheses after the cycle number represents the operation performed after the execution of this cycle. The number of cycles required for each instruction to complete the calculation is shown in the parentheses after the instruction. Taking instruction A as an example, instruction A needs to calculate for 4 cycles to complete the calculation.
[0222] The following introduces the execution situations of 3 instructions in different cycles.
[0223] At clock 1: Instruction A is obtained.
[0224] At clock 2: Instruction A is input into the first stage of the pipeline for processing, and the first calculation cycle of instruction A is executed. Instruction B is obtained.
[0225] At clock 3: Instruction A is input into the second stage of the pipeline for processing, and the second calculation cycle of instruction A is executed. Instruction B is input into the first stage of the pipeline for processing, and the first calculation cycle of instruction B is executed. Instruction C is obtained.
[0226] At clock 4: Instruction A is input to the third-stage pipeline for processing, and the third computing cycle of Instruction A is executed. Instruction B is input to the second-stage pipeline for processing, and the second computing cycle of Instruction B is executed. Instruction C is input to the first-stage pipeline for processing, and the first computing cycle of Instruction C is executed. After this cycle is completed, both Instruction B and Instruction C are computed, but Instruction A, which was received first, has not been output yet. To ensure order, at this time, Instruction B and Instruction C need to be bypassed. After the result of Instruction A is output, the result of Instruction B is output, and finally the result of Instruction C is output.
[0227] It can be seen that the two data in the first and second stages of the pipeline have been computed and are waiting to be output, but the valid data corresponding to the instruction that arrived first in the third stage has not been computed yet. That is, there is still valid data (in the third-stage pipeline) after the last data (in the second-stage pipeline). At this time, to ensure order, it is necessary to wait until the data that arrived first is computed, and then output these results in order.
[0228] When the control unit performs pipeline behavior control, for each stage of the pipeline, it needs to determine the action taken by this stage of the pipeline based on the action determined in the next cycle of the next stage of the pipeline. This determination is initiated first by the last stage of the pipeline and then recursively from back to front. Therefore, the conditions required for each stage of the pipeline to initiate an action can be clarified. Please refer to Table 2, which shows the conditions corresponding to the pipeline behavior, that is, when the required conditions are met, the behavior of the corresponding pipeline is controlled.
[0229] Table 2. Conditions Corresponding to Different Pipeline Behaviors
[0230]
[0231]
[0232] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a data processing method provided by an embodiment of the present disclosure. The method provided in this embodiment is applied to a coprocessor, and the coprocessor includes: a control unit and an N-stage pipeline computing unit, where N is an integer greater than 1. Each stage of the pipeline computing unit includes a pipeline module and a cache module connected in sequence; the cache modules are connected between every two adjacent ones. The coprocessor can be the coprocessor in any of the above embodiments. The method provided in this embodiment includes the following steps.
[0233] Step 501: The control unit receives a computing instruction sent by the main processor.
[0234] Further, the computing instruction can be established based on the R-type instruction in the Reduced Instruction Set Computer (RISC-V).
[0235] Step 502: The control unit parses the calculation instruction to obtain the mathematical operation and operands corresponding to the calculation instruction; in the order of receiving the calculation instructions, it inputs the mathematical operation and operands corresponding to the calculation instruction into the N-stage pipelined computing unit.
[0236] Step 503: After receiving the calculation instruction, the N-stage pipelined computing unit performs pipelined iterative calculation processing on the calculation instruction through the N-stage pipeline module to obtain the calculation result corresponding to the calculation instruction.
[0237] Among them, the pipeline module is established based on the Coordinate Rotation Digital Computer algorithm (CORDIC).
[0238] Step 504: During the process of the N-stage pipelined computing unit executing the calculation instruction, when the result prediction output by the pipeline module is blocked, the control unit: stores the result output by the pipeline module with the prediction blocked in the cache module corresponding to the pipeline module with the prediction blocked; or stores the result output by the pipeline module with the prediction blocked in the cache module corresponding to the next-stage pipeline module of the pipeline module with the prediction blocked.
[0239] Step 505: The control unit sends the calculation result corresponding to the calculation instruction to the main processor.
[0240] In some embodiments, the method further includes the following steps 505-506.
[0241] Step 506: The control unit determines, in the order from the back to the front of the pipelined computing unit, the predicted data path conditions of each pipelined computing unit in the next clock cycle; the predicted data path conditions include that the pipeline module is predicted to be blocked or not blocked; being predicted to be blocked includes: predicted bypass or predicted backpressure.
[0242] Step 507: In the next clock cycle, the control unit controls the behavior of each pipelined computing unit based on the predicted data path conditions of each pipelined computing unit; the behavior of the pipelined computing unit refers to the processing method of the result output by the pipeline module. Among them, when the predicted data path condition of the pipeline module is predicted bypass, the behavior of the pipelined computing unit includes storing the result output by the pipelined computing unit in the cache module of the next-stage pipelined computing unit of the pipelined computing unit; when the pipelined computing unit predicts backpressure, the behavior of the pipelined computing unit includes storing the result output by the pipelined computing unit in the cache module of the pipelined computing unit.
[0243] In some embodiments, based on the above embodiments, further, step 506 may include the following steps.
[0244] Step 5061: The control unit determines the predicted data path condition of the last pipeline calculation unit in the next clock cycle based on the tag value of the last pipeline calculation unit in the N-level pipeline calculation unit and the current situation of the main processor receiving the calculation result.
[0245] Among them, the tag value of the pipeline calculation unit refers to the number of clock cycles required to complete the calculation instructions in the pipeline calculation unit.
[0246] Step 5062: In the order from back to front, starting from the (N - 1)-th pipeline calculation unit, each pipeline calculation unit in the N-level pipeline calculation unit is processed as follows: Determine the predicted data path condition of the pipeline calculation unit in the next clock cycle according to the condition information of the pipeline calculation unit.
[0247] Among them, the condition information of the pipeline calculation unit includes at least one of the following: the tag value of the pipeline calculation unit, the data path condition of the next pipeline calculation unit of the pipeline calculation unit in the next clock cycle, the arbitration result of the pipeline calculation unit, and the information of the main processor currently receiving the calculation result; the arbitration result of the pipeline calculation unit is used to indicate whether the result output by the pipeline calculation unit is sent to the main processor in the next clock cycle.
[0248] In some embodiments, based on the above embodiments, further, step 5062 may include the following step 50621, and correspondingly, step 507 may be implemented through the following step 5071.
[0249] Step 50621: For the first (N - 1) pipeline calculation units, if the tag value of the pipeline calculation unit is zero, the arbitration result of the pipeline calculation unit indicates that the arbitration of the pipeline calculation unit is successful, and the main processor currently indicates not to receive the calculation result, then determine the predicted data path condition of the pipeline calculation unit as a predicted bypass.
[0250] Step 5071: In the next clock cycle, the control unit stores the result output by the pipeline calculation unit in the cache module of the next pipeline calculation unit of the pipeline calculation unit.
[0251] In some embodiments, based on the above embodiments, further, step 5062 may include the following step 50622, and correspondingly, step 507 may be implemented through the following step 5072.
[0252] Step 50622: For the first N-1 levels of pipeline computing units, if the tag value of a pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit fails, and the predicted data path condition of the next-level pipeline computing unit of the pipeline computing unit is not predicted backpressure, then determine that the predicted data path condition of the pipeline computing unit is predicted bypass.
[0253] Step 5072: In the next clock cycle, the control unit stores the result output by the pipeline computing unit in the cache module of the next-level pipeline computing unit of the pipeline module.
[0254] In some embodiments, based on the above embodiments, further, step 5062 may include the following step 50623, and correspondingly, step 507 may be implemented through the following step 5073.
[0255] Step 50623: For the first N-1 levels of pipeline computing units, if the tag value of a pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the arbitration of the pipeline computing unit fails, and the predicted data path condition of the next-level pipeline computing unit of the pipeline computing unit is predicted backpressure, then determine that the predicted data path condition of the pipeline computing unit is predicted backpressure.
[0256] Step 5073: In the next clock cycle, the control unit stores the result output by the pipeline computing unit in the cache module of the pipeline computing unit.
[0257] In some embodiments, based on the above embodiments, further, step 5062 may include the following step 50624, and correspondingly, step 507 may be implemented through the following step 5074.
[0258] Step 50624: For the first N-1 levels of pipeline computing units, if the tag value of a pipeline computing unit is greater than zero, and the predicted data path condition of the next-level pipeline computing unit of the pipeline computing unit is predicted backpressure, then determine that the predicted data path condition of the pipeline computing unit is predicted backpressure.
[0259] Step 5074: In the next clock cycle, the control unit stores the result output by the pipeline computing unit in the cache module of the pipeline computing unit.
[0260] In some embodiments, based on the above embodiments, further, step 5062 may include the following step 50625, and correspondingly, step 507 may be implemented through the following step 5075.
[0261] Step 50625: For the first N - 1 levels of pipeline computing units, if the tag value of a pipeline computing unit is greater than zero and the next - level pipeline computing unit of the pipeline computing unit is predicted to be unblocked, the control unit determines that the predicted data path condition of the pipeline computing unit is predicted to be unblocked.
[0262] Step 5075: In the next clock cycle, the control unit inputs the result output by the pipeline computing unit to the next - level pipeline computing unit of the pipeline computing unit.
[0263] In some embodiments, based on the above - mentioned embodiments, the method provided in this embodiment further includes, in step 5062, the following step 50626. Correspondingly, step 507 can be implemented through the following step 5076.
[0264] Step 50626: For the first N - 1 levels of pipeline computing units, if the tag value of a pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates successful arbitration of the pipeline module, and the main processor currently indicates receiving the calculation result, the control unit determines that the predicted data path condition of the pipeline computing unit is predicted to be unblocked.
[0265] Step 5076: In the next clock cycle, the control unit takes the result output by the pipeline computing unit as the calculation result corresponding to the calculation instruction.
[0266] In some embodiments, based on the above - mentioned embodiments, the method provided in this embodiment further includes, in step 5061, the following step 50611. Correspondingly, step 507 can be implemented through the following step 5077.
[0267] Step 50611: For the last - level pipeline computing unit in the N - level pipeline computing unit, if the tag value of the last - level pipeline computing unit is zero and the main processor currently indicates that it can receive the calculation result, the control unit determines that the predicted data path condition of the next clock cycle of the last - level pipeline computing unit is predicted to be unblocked.
[0268] Step 5076: In the next clock cycle, the control unit takes the result output by the last - level pipeline computing unit as the calculation result corresponding to the calculation instruction.
[0269] In some embodiments, based on the above - mentioned embodiments, the method provided in this embodiment further includes, in step 5061, the following step 50612. Correspondingly, step 507 can be implemented through the following step 5078.
[0270] Step 50612: For the last pipeline computing unit in the N-stage pipeline module, if the tag value of the last pipeline computing unit is zero and the main processor currently indicates that it cannot receive the calculation result, the predicted data path situation for the next clock cycle of the last pipeline computing unit is determined to be predicted backpressure.
[0271] Step 5078: In the next clock cycle, the control unit stores the result output by the last pipeline computing unit in the cache module in the last pipeline computing unit.
[0272] In some embodiments, the method provided in this embodiment further includes:
[0273] Step 508: If there is a pipeline computing unit with a tag value of zero in the N-stage pipeline computing unit, the control unit obtains the last pipeline computing unit among the pipeline computing units with a tag value of zero as the first pipeline computing unit.
[0274] Step 509: The control unit determines that the arbitration result of other pipeline computing units in the N-stage pipeline computing unit except the first pipeline computing unit is arbitration failure.
[0275] Step 510: If there is no valid data after the first pipeline computing unit in the pipeline computing unit, the control unit determines that the arbitration result of the first pipeline computing unit is arbitration success.
[0276] Step 511: If there is valid data after the first pipeline computing unit in the pipeline computing unit, the control unit determines that the arbitration result of the first pipeline computing unit is arbitration failure.
[0277] The implementation principle and beneficial effects of the method in this embodiment are similar to those in the above embodiments, and will not be elaborated here.
[0278] Based on the data processing method described in any of the above embodiments, the present disclosure embodiment also provides a computer-readable storage medium. For example, a non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. Computer instructions are stored on the storage medium for executing the data processing method described in any of the above embodiments, which will not be elaborated here.
[0279] The present disclosure embodiment provides a processor system, including a coprocessor and a main processor as described in any of the above embodiments, and the main processor is communicatively connected to the coprocessor.
[0280] The system of this embodiment has a similar implementation principle and beneficial effects to those of the above embodiment, which will not be elaborated here.
[0281] Those skilled in the art will readily conceive of other implementations of the present disclosure after considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
Claims
1. A coprocessor, characterized in that: include: A control unit and N-stage pipeline computing units connected in sequence, where N is an integer greater than 1; Each level of pipeline computing unit includes a pipeline module and a cache module connected in sequence; Every two adjacent cache modules are connected; The control unit is used to receive the calculation instruction sent by the main processor, and parse the calculation instruction to obtain the mathematical operation and operand corresponding to the calculation instruction; Inputting the mathematical operations and operands corresponding to the computing instructions into the N-stage pipeline computing unit in the order in which the computing instructions are received; The N-stage pipeline calculation unit is used to perform pipeline iterative calculation processing on the calculation instruction through the N-stage pipeline module to obtain the calculation result corresponding to the calculation instruction; The control unit is also used for, during the process of the N-stage pipeline computing unit executing a computing instruction, when the result output by the pipeline computing unit is predicted to be blocked: causing the result output by the pipeline computing unit predicted to be blocked to be stored in a cache module in the pipeline computing unit predicted to be blocked; or causing the result output by the pipeline computing unit predicted to be blocked to be stored in a cache module in the pipeline computing unit of the next stage of the pipeline computing unit predicted to be blocked; and sending the computing result corresponding to the computing instruction to the main processor.
2. The coprocessor according to claim 1, characterized in that: The control unit is also used for: According to the order of the pipeline computing units from back to front, respectively determine the predicted data path status of each pipeline computing unit in the next clock cycle; the predicted data path status refers to whether the pipeline module is predicted to be blocked or predicted not blocked; the predicted blocking includes: predicted bypass or predicted back pressure; In the next clock cycle, the behavior of each pipeline computing unit is controlled separately based on the predicted data path status of each pipeline computing unit; the behavior of the pipeline computing unit refers to the processing method of the result output by the pipeline module, wherein, when the predicted data path status of the pipeline module is a predicted bypass, the behavior of the pipeline computing unit includes storing the result output by the pipeline computing unit in a cache module in a pipeline computing unit of the next level of the pipeline computing unit; when the pipeline computing unit predicts back pressure, the behavior of the pipeline computing unit includes storing the result output by the pipeline computing unit in a cache module in the pipeline computing unit.
3. The coprocessor according to claim 2, characterized in that: The control unit is specifically used for: Determine, according to the label value of the last-stage pipeline computing unit in the N-stage pipeline computing units and the current receiving of the computing result by the main processor, the predicted data path status of the last-stage pipeline computing unit in the next clock cycle; the label value of the pipeline computing unit refers to the number of clock cycles required to complete the calculation of the computing instruction in the pipeline computing unit; In order from back to front, starting from the N-1th pipeline computing unit, each pipeline computing unit in the N-stage pipeline computing units is processed as follows: according to the condition information of the pipeline computing unit, the predicted data path condition of the pipeline computing unit in the next clock cycle is determined; the condition information of the pipeline computing unit includes at least one of the following: the label value of the pipeline computing unit, the data path condition of the next-stage pipeline computing unit of the pipeline computing unit in the next clock cycle, the arbitration result of the pipeline computing unit, and the information that the main processor currently receives the calculation result; The arbitration result of the pipeline calculation unit is used to indicate whether the result output by the pipeline calculation unit is sent to the main processor in the next clock cycle.
4. The coprocessor according to claim 3, characterized in that: The control unit is specifically used for: For the first N-1 stages of pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the pipeline computing unit has successfully arbitrated, and the main processor currently indicates that it does not receive the calculation result, then determine that the predicted data path status of the pipeline computing unit is a predicted bypass; in the next clock cycle, store the result output by the pipeline computing unit in a cache module in the pipeline computing unit of the next stage of the pipeline computing unit; For the first N-1 stages of pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the pipeline computing unit has failed arbitration, and the predicted data path condition of the pipeline computing unit at the next stage of the pipeline computing unit is not predicted back pressure, then the predicted data path condition of the pipeline computing unit is determined to be predicted bypass; in the next clock cycle, the result output by the pipeline computing unit is stored in a cache module in the pipeline computing unit at the next stage of the pipeline module; For the first N-1 stages of pipeline computing units, if the tag value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the pipeline computing unit has failed arbitration, and the predicted data path condition of the pipeline computing unit at the next stage of the pipeline computing unit is predicted back pressure, then determine that the predicted data path condition of the pipeline computing unit is predicted back pressure; in the next clock cycle, store the result output by the pipeline computing unit in a cache module in the pipeline computing unit; For the first N-1 levels of pipeline computing units, if the label value of the pipeline computing unit is greater than zero, and the predicted data path condition of the next level of pipeline computing unit of the pipeline computing unit is predicted back pressure, then the predicted data path condition of the pipeline computing unit is determined to be predicted back pressure; in the next clock cycle, the result output by the pipeline computing unit is stored in a cache module in the pipeline computing unit.
5. The coprocessor according to claim 3, characterized in that: The control unit is specifically used for: For the first N-1 pipeline computing units, if the tag value of the pipeline computing unit is greater than zero, and the pipeline computing unit at the next stage of the pipeline computing unit is predicted to be not blocked, then the predicted data path status of the pipeline computing unit is determined to be predicted to be not blocked; in the next clock cycle, the result output by the pipeline computing unit is input to the pipeline computing unit at the next stage of the pipeline computing unit; For the first N-1 stages of pipeline computing units, if the label value of the pipeline computing unit is zero, the arbitration result of the pipeline computing unit indicates that the pipeline module arbitration is successful, and the main processor currently instructs to receive the calculation result, then it is determined that the predicted data path status of the pipeline computing unit is predicted to be not blocked; in the next clock cycle, the result output by the pipeline computing unit is used as the calculation result corresponding to the calculation instruction.
6. The coprocessor according to claim 3, characterized in that: The control unit is specifically used for: For a last-stage pipeline computing unit in the N-stage pipeline computing units, if the tag value of the last-stage pipeline computing unit is zero and the main processor currently indicates that it is capable of receiving a computing result, determining that a predicted data path condition of a next clock cycle of the last-stage pipeline computing unit is predicted to be not blocked; For the last-stage pipeline computing unit in the N-stage pipeline module, if the tag value of the last-stage pipeline computing unit is zero and the main processor currently indicates that the calculation result cannot be received, the predicted data path condition of the next clock cycle of the last-stage pipeline computing unit is determined to be predicted back pressure.
7. The coprocessor according to any one of claims 3 to 6, characterized in that: The control module is also used for: If there is a pipeline computing unit with a label value of zero in the N-stage pipeline computing units, the pipeline computing unit at the rear of the pipeline computing units with a label value of zero is obtained as the first pipeline computing unit; and the arbitration results of the other pipeline computing units in the N-stage pipeline computing units except the first pipeline computing unit are determined to be arbitration failure; If in the pipeline computing unit, there is no valid data after the first pipeline computing unit, determining that the arbitration result of the first pipeline computing unit is arbitration success; If valid data exists after the first pipeline computing unit in the pipeline computing unit, it is determined that the arbitration result of the first pipeline computing unit is arbitration failure.
8. The coprocessor according to claim 1, characterized in that: The computing instructions are established based on R-type instructions in the reduced instruction set RISC-V.
9. A data processing method, characterized in that: Applied to a coprocessor, the coprocessor comprises: a control unit and N-level pipeline computing units connected in sequence, N being an integer greater than 1; each level of pipeline computing unit comprises a pipeline module and a cache module connected in sequence; and each adjacent two cache modules are connected; The control unit receives a calculation instruction sent by the main processor; The control unit parses the calculation instruction to obtain mathematical operations and operands corresponding to the calculation instruction; and inputs the mathematical operations and operands corresponding to the calculation instruction into the N-stage pipeline calculation unit in the order in which the calculation instructions are received; The N-stage pipeline calculation unit performs pipeline iterative calculation processing on the calculation instruction through the N-stage pipeline module to obtain a calculation result corresponding to the calculation instruction; The control unit, during the process of the N-stage pipeline computing unit executing the computing instruction, when the result output by the pipeline computing unit is predicted to be blocked: causes the result output by the pipeline computing unit predicted to be blocked to be stored in a cache module in the pipeline computing unit predicted to be blocked; or causes the result output by the pipeline computing unit predicted to be blocked to be stored in a cache module in a pipeline computing unit at a next stage of the pipeline computing unit predicted to be blocked; The control unit sends a calculation result corresponding to the calculation instruction to the main processor.
10. A processor system, characterized in that: The invention comprises a coprocessor as claimed in any one of claims 1 to 8 and a main processor, wherein the main processor is communicatively connected with the coprocessor.
11. An electronic device, characterized in that: Comprising a processor system as claimed in claim 10 above.