Pipeline control method, arithmetic module and related products
By adopting multiple pipelines and multiple exit ports in the processor and using the digital beat controller to manage instructions to execute, the problems of calculation delay and low efficiency in traditional processor design are solved, and more efficient calculations and orderly output results are achieved.
Patent Information
- Application Number
- CN202010654552.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-12-19
AI Technical Summary
In traditional processor design, the design of a single pipeline causes different types of computing instructions to pass through all pipeline stages, resulting in computing delays and reduced computing efficiency.
The design of multiple pipelines and multiple exit entrances is adopted. The instruction execution of the pipeline is managed through the digital beat controller, allowing the execution of instructions to be flexibly configured at different entrances and exits of the pipeline to prevent the output results from being out of order.
Improves computing efficiency, reduces computing delays, and prevents out-of-order output results from the pipeline output.
Smart Images

Figure CN113918220B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a pipeline control method, an arithmetic module, and related products. Background Art
[0002] In traditional processor designs, a single pipeline is used to implement multiple arithmetic functions in the same arithmetic module. This pipeline generally has only 1 inlet and 1 outlet, uses the same arithmetic path, and data and control signals are transmitted between the pipeline stages through a handshake protocol. However, such a design causes different types of arithmetic instructions to pass through all the pipeline stages. Even if no operation is performed in some pipeline stages, but only data and control information are passed as they are, all the pipeline stages need to be passed through, resulting in calculation latency and reducing the calculation efficiency. Summary of the Invention
[0003] Embodiments of this application provide a pipeline control method, an arithmetic module, and related products, which can improve the calculation efficiency based on the design of multiple inlets and outlets of multiple pipelines, and can prevent the output results from being out of order at the output end of the pipeline.
[0004] In a first aspect, embodiments of this application further provide an arithmetic module, which includes K pipelines, P pipeline inlets, Q pipeline outlets, and a cycle counter. K, P, and Q are all integers greater than or equal to 2; the cycle counter stores the remaining cycles of the first instruction with the longest required cycles among the instructions being executed in the K pipelines;
[0005] The arithmetic module is used for:
[0006] Obtain the remaining cycles of the first instruction;
[0007] Obtain the execution cycles of the second instruction to be issued;
[0008] When the execution cycles of the second instruction are less than the remaining cycles of the first instruction, allow the second instruction to enter the corresponding pipeline.
[0009] In a second aspect, embodiments of this application further provide a pipeline control method, which is applied to an arithmetic module. The arithmetic module includes K pipelines, P pipeline inlets, Q pipeline outlets, and a cycle counter. K, P, and Q are all integers greater than or equal to 2; the cycle counter stores the remaining cycles of the first instruction with the longest required cycles among the instructions being executed in the K pipelines; the method includes:
[0010] Obtain the remaining cycles of the first instruction;
[0011] Obtain the number of execution cycles of the second instruction to be issued;
[0012] When the number of execution cycles of the second instruction is less than the remaining number of cycles of the first instruction, allow the second instruction to enter the corresponding pipeline.
[0013] In a third aspect, an embodiment of the present application further provides a neural network chip, which includes the operation module described in any aspect of the first aspect, or is used to execute the method described in the second aspect.
[0014] In a fourth aspect, an embodiment of the present application further provides a board, which includes: a storage device, an interface device, a control device, and the neural network chip described in the third aspect;
[0015] wherein, the neural network chip is respectively connected to the storage device, the control device, and the interface device;
[0016] The storage device is used to store data;
[0017] The interface device is used to implement data transmission between the chip and external devices;
[0018] The control device is used to monitor the state of the chip.
[0019] In a fifth aspect, an embodiment of the present application further provides an electronic device, which includes the operation module described in any item of the first aspect, or the electronic device is used to execute the method described in the second aspect, or the electronic device includes the neural network chip described in the third aspect, or the electronic device includes the board described in the fourth aspect.
[0020] In a sixth aspect, an embodiment of the present application further provides a pipeline control device, which is applied to an operation module. The operation module includes K pipelines, P pipeline entrances, Q pipeline exits, and a cycle counter. The K, the P, and the Q are all integers greater than or equal to 2; the cycle counter stores the remaining number of cycles of the first instruction with the longest required number of cycles among the instructions being executed in the K pipelines; the device includes: an acquisition unit and an execution unit, where
[0021] The acquisition unit is used to acquire the remaining number of cycles of the first instruction;
[0022] The acquisition unit is further used to acquire the number of execution cycles of the second instruction to be issued;
[0023] The execution unit is used to allow the second instruction to enter the corresponding pipeline when the number of execution cycles of the second instruction is less than the remaining number of cycles of the first instruction.
[0024] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for electronic data exchange, and the computer program causes a computer to execute some or all of the steps described in the second aspect of the embodiments of the present application.
[0025] In an eighth aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps described in the second aspect of the embodiments of the present application. The computer program product may be a software installation package.
[0026] Adopting the embodiments of the present application has the following beneficial effects:
[0027] It can be seen that the pipeline control method, arithmetic module and related products described in the embodiments of the present application are applied to the arithmetic module. The arithmetic module includes K pipelines, P pipeline entrances, Q pipeline exits and a cycle counter. K, P, and Q are all integers greater than or equal to 2. The cycle counter stores the remaining cycles of the first instruction with the longest required cycles among the instructions being executed in the K pipelines. The arithmetic module obtains the remaining cycles of the first instruction and obtains the execution cycles of the second instruction to be issued. When the execution cycles are less than the remaining cycles, the second instruction is allowed to enter the corresponding pipeline. On the one hand, it can improve the computing efficiency based on the design of multiple pipelines with multiple entrances and exits. On the other hand, it can prevent the output results from being out of order at the output end of the pipeline.
[0028] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0030] Figure 1A It is a schematic structural diagram of an arithmetic module provided by an embodiment of the present application;
[0031] Figure 1B It is a demonstration schematic diagram of a pipeline provided by an embodiment of the present application;
[0032] Figure 2 It is a schematic flowchart of a pipeline control method provided by an embodiment of the present application;
[0033] Figure 3 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application;
[0034] Figure 4 Block diagram of the functional units of a pipeline control device provided by an embodiment of the present application;
[0035] Figure 5 Block diagram of the functional units of a combined processing device provided by an embodiment of the present application;
[0036] Figure 6 Block diagram of the functional units of a board card provided by an embodiment of the present application. Detailed implementation manners
[0037] The following will be described in detail respectively.
[0038] Terms such as "first", "second", "third", and "fourth" in the specification and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0039] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0040] The electronic device may include various handheld devices, vehicle-mounted devices, wireless headsets, computing devices, or other processing devices connected to a wireless modem that have wireless communication functions, as well as various forms of user equipment (UE), mobile station (MS), terminal device, etc. The electronic device may be, for example, a smart phone, a tablet computer, a headphone case, etc. For convenience of description, the devices mentioned above are collectively referred to as electronic devices.
[0041] The above electronic devices can be applied to the following (including but not limited to) scenarios: various electronic products such as data processing, robots, computers, printers, scanners, telephones, tablet computers, smart terminals, mobile phones, driving recorders, navigators, sensors, cameras, cloud servers, cameras, video cameras, projectors, watches, earphones, mobile storage, wearable devices, etc.; various means of transportation such as airplanes, ships, vehicles, etc.; various household appliances such as televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, range hoods, etc.; and various medical devices including nuclear magnetic resonance instruments, B-ultrasounds, electrocardiographs, etc.
[0042] In the related art, in order to ensure that the output of instructions can conform to the input order, on the premise of using a single-entry and single-exit pipeline to implement the arithmetic function, operations of different types, different latencies, and different complexities all enter from the first stage at the front of the pipeline, flow through all pipeline stages of the pipeline, and are output from the last stage of the pipeline. Therefore, there are the following technical defects:
[0043] 1. Resource waste. If a certain instruction does not perform any effective logical operations in some pipeline stages of the pipeline and just flows to the next stage as it is. Then it can be considered that for this instruction, there is resource waste in some pipeline stages because the effective logical operations existing in these pipeline stages are not utilized.
[0044] 2. Performance loss. If a certain instruction has completed all effective operations and obtained the final result in a certain pipeline stage in the middle of the pipeline, but still needs to flow the result through the remaining pipeline stages and then be output from the last stage. It can be considered that for this instruction, the actual execution time is longer than the execution time of its effective part, that is, there is performance loss.
[0045] Based on the above analysis, in the embodiment of the present application, in the solution of multiple pipeline multiple exits and entrances, if multiple instructions processed in parallel by multiple pipelines cannot be ordered, instruction out-of-order will occur. At this time, the number-of-clocks protocol of the embodiment of the present application can be used to order instructions in multiple pipelines.
[0046] Specifically, in a processor including multiple pipelines, at the instruction input ends of multiple pipelines, a number-of-clocks controller is set. The number-of-clocks controller stores the instruction with the longest number of clocks required among the instructions being executed in the current pipeline, and is updated by counting down according to the clock signal. For example, if the longest number of clocks among the instructions currently being executed in the current pipeline is 3 clocks, then an instruction whose number of clocks to be issued is less than 3 can enter the pipeline, otherwise, if it is greater than or equal to 3, it cannot enter the pipeline, preventing the output result from being out-of-order at the output end of the pipeline.
[0047] The following will be introduced in detail.
[0048] Please refer to Figure 1A , Figure 1A which is a schematic diagram of an operation module provided by an embodiment of the present application. The operation module includes K pipelines, P pipeline entrances, Q pipeline exits, and a cycle counter. K, P, and Q are all integers greater than or equal to 2. The cycle counter stores the remaining cycles of the first instruction with the longest required cycles among the instructions being executed in the K pipelines;
[0049] The operation module is used to perform the following steps:
[0050] A1. Obtain the remaining cycles of the first instruction;
[0051] A2. Obtain the execution cycles of the second instruction to be issued;
[0052] A3. When the execution cycles of the second instruction are less than the remaining cycles of the first instruction, allow the second instruction to enter the corresponding pipeline.
[0053] Among them, a pipeline may include multiple pipeline stages, and each pipeline stage may correspond to at least one arithmetic circuit. The arithmetic circuit can support operations of multiple data types, and the corresponding arithmetic circuit is selected according to the instruction requirements to complete the corresponding operation. For example, the data type can be 16-bit fixed-point data or 32-bit floating-point data, etc. For example, if the instruction is matrix addition, an adder is selected; if the instruction is matrix multiplication, a multiplier and an adder are selected; if the instruction is a 16-bit fixed-point operation instruction, the 16-bit fixed-point operation is received, and so on. Specifically, the arithmetic circuit can be an arithmetic logic circuit and / or a logic operation circuit. The arithmetic logic circuit can be at least one of the following: a multiplier unit, an adder unit. The logic operation circuit can be at least one of the following: a combinational logic circuit, a sequential logic circuit, etc., which are not limited herein. The combinational logic circuit is composed of the most basic logic gate circuits, such as adders, decoders, encoders, data selectors, etc. The sequential logic circuit is a circuit composed of the most basic logic gate circuits plus a feedback logic circuit (output to input) or devices, such as flip-flops, latches, counters, shift registers, memories, etc. The above computing resources can be understood as the arithmetic circuits that need to be called.
[0054] Among them, the data processed by the pipeline can be at least one of the following: neuron data, weight data, and bias data. The data that the pipeline can process can be at least one of the following data types: fixed-point data, integer data, discrete data, continuous data, power data, floating-point data. The length of the data representation can be 32-bit floating-point data, 16-bit fixed-point data, 16-bit floating-point data, 8-bit fixed-point data, 4-bit fixed-point data, and so on. The data can include at least one of the following: input neuron data, weight data, and bias data.
[0055] Both the first instruction and the second instruction can be vector instructions. The vector instructions can be at least one of the following: vector addition instruction, vector plus scalar instruction, vector subtraction instruction, vector multiplication instruction, vector multiply scalar instruction, vector division instruction, scalar divide vector instruction, vector AND instruction, vector IN AND instruction, vector OR instruction, vector IN OR instruction, vector exponentiation instruction, vector logarithm instruction, vector greater than determination instruction, vector equal to determination instruction, vector NOT instruction, vector select merge instruction, vector maximum value instruction, scalar extension instruction, scalar replace vector instruction, vector replace scalar instruction, vector retrieval instruction, vector dot product instruction, random vector instruction, circular shift instruction, vector load instruction, vector store instruction, vector transfer instruction, matrix multiply vector instruction, vector multiply matrix instruction, matrix multiply scalar instruction, tensor operation instruction, matrix addition instruction, matrix subtraction instruction, matrix retrieval instruction, matrix load instruction, matrix store instruction, matrix transfer instruction.
[0056] Both the first instruction and the second instruction can be single-function instructions or can be combined-function instructions. The single-function instructions can be used to implement single functions, and the combined-function instructions can be used to implement several functions (combined functions). The single-function instructions or combined instructions can be executed using one or more pipelines.
[0057] In the embodiments of the present application, the operation module may include K pipelines, as well as P pipeline entrances and Q pipeline exits. K is a positive integer, and both P and Q are integers greater than or equal to 2. Each of the K pipelines includes multiple pipeline stages, and each pipeline stage corresponds to at least one operation circuit. Any pipeline can include a pipeline entrance, a pipeline exit, and t intermediate pipeline stages. t is a positive integer. For example, as Figure 1B shown, the pipeline entrance, intermediate pipeline stage 1, intermediate pipeline stage 2, and pipeline exit 2 can form a pipeline. There can be at least one register between adjacent pipeline stages.
[0058] In specific implementation, since different computing resources are distributed in different pipeline stages of the pipeline. For example, multiplier units are distributed in the 3rd and 4th pipeline stages, while adder units are distributed in the 1st and 2nd pipeline stages. Different instructions require different computing resources for execution, and not all computing resources are necessarily used. At the same time, some instructions do not need to flow through to the last pipeline stage to output results. Therefore, according to the usage of computing resources by the instructions, the entry and exit of the instructions into the pipeline can be flexibly configured. For example, a single multiplication instruction can enter from the 3rd pipeline stage, complete the calculation in the 3rd and 4th stages, and output the result from the 4th stage; for another example, a single addition instruction can enter from the 5th pipeline stage, complete the calculation in the 5th and 6th stages, and output the result from the 6th stage. The advantage of this is that all instructions can complete the calculation in the pipeline with the minimum calculation delay and using the least amount of computing resources.
[0059] In the embodiments of the present application, the operation module adopts the above-mentioned multi-pipeline design strategy with multiple entrances and exits, which solves the physical realizability problem while saving area. Taking 2 pipelines as an example, it can divide the instructions to be implemented by this module into 2 categories according to functions, reusability, and balance principles, and implement them in these 2 pipelines respectively.
[0060] In specific implementation, the operation module can obtain the remaining beats of the first instruction through a beat counter, and obtain the execution beats of the second instruction to be issued. And when the execution beats are less than the remaining beats, it allows the second instruction to enter the corresponding pipeline, that is, allows the second instruction to enter the pipeline entrance of the corresponding pipeline. In this way, it can prevent the output results from being out of order at the output end of the pipeline.
[0061] In a possible example, the operation module is further specifically configured to:
[0062] When the execution beats of the second instruction are greater than or equal to the remaining beats of the first instruction, wait for the first instruction to complete execution, and after the first instruction completes execution, allow the second instruction to enter the corresponding pipeline, and update the remaining beats of the first instruction in the beat counter.
[0063] In specific implementation, through the above method, it can prevent the output results from being out of order at the output end of the pipeline.
[0064] In the embodiments of the present application, the beat counter is arranged at the instruction input end of the K pipelines.
[0065] In specific implementation, arranging the beat counter at the instruction input end of the K pipelines can be used to count the remaining beats of the instructions in the currently executing pipeline.
[0066] In a possible example, the pipeline includes a plurality of pipeline stages, each pipeline stage corresponding to at least one arithmetic circuit. A resource table is also stored in the cycle counter. The arithmetic module is further configured to:
[0067] Update the remaining cycles of the first instruction according to the resource table.
[0068] In a specific implementation, the cycle counter may further include a resource table, which can be used to record the current or subsequent occupancy of each arithmetic circuit corresponding to each instruction being executed in the current pipeline, that is, the resource table stores the arithmetic circuits occupied by each instruction being executed. Based on this resource table, the cycle counter can also be used to query the resource table. For example, update the remaining cycles of the first instruction according to the resource table. For example, through the resource table, the occupancy of the current arithmetic circuit can be seen, and this occupancy reflects the execution position of the instruction. Furthermore, the execution position of the first instruction can be known, and based on this execution position, the remaining cycles of the first instruction can be known. In this way, the arithmetic module can better manage the instruction sequencing based on the cycle counting.
[0069] The cycle counter can maintain a resource table. The resource table can be used to record the occupancy of each arithmetic circuit. For example, for pipeline 1, a resource table including at least six flag bits can be set, and each flag bit can be used to indicate whether the corresponding arithmetic circuit is in use. By querying the resource table, the cycle counter can better manage the instruction sequencing based on the cycle counting and maximize the instruction out-of-order execution, avoiding the situation where the reuse rate of the arithmetic circuit is low due to only cycle-based restrictions.
[0070] In a possible example, the arithmetic module is further configured to:
[0071] Determine whether to allow the second instruction to enter the corresponding pipeline according to the resource table and the arithmetic circuits required by the second instruction.
[0072] In a specific implementation, even if the cycle count does not meet the requirement (the execution cycles of the second instruction are greater than or equal to the remaining cycles of the first instruction), as long as the arithmetic circuits occupied by the second instruction do not conflict with the arithmetic circuits required by the remaining cycles of the first instruction, the second instruction can also enter the corresponding pipeline. For example, if the execution cycles of instruction A are greater than the remaining cycles of instruction B being executed in the pipeline, when the arithmetic circuits occupied by instruction A conflict with the arithmetic circuits required by the remaining cycles of instruction B, then instruction B is not allowed to enter the corresponding pipeline. On the contrary, when the arithmetic circuits occupied by instruction A do not conflict with the arithmetic circuits required by the remaining cycles of instruction B, then instruction B is allowed to enter the corresponding pipeline.
[0073] In a possible example, the operation module is further specifically configured to implement the following functions:
[0074] B1. Obtain the resource table in the currently executing pipeline. The resource table includes multiple flag bits, and each flag bit is used to indicate whether the corresponding arithmetic circuit is occupied;
[0075] B2. Determine the occupancy of the arithmetic circuits in the currently executing pipeline based on the multiple flag bits.
[0076] In specific implementation, the operation module can be used to obtain the resource table in the currently executing pipeline. The resource table includes multiple flag bits, and each flag bit is used to indicate whether the corresponding arithmetic circuit is occupied. For example, the flag bit can be represented by Flag. Flag = 1 can indicate that the arithmetic circuit is occupied, and Flag = 0 indicates that the arithmetic circuit is not occupied. Furthermore, the occupancy of the arithmetic circuits in the currently executing pipeline can be determined based on the multiple flag bits, as shown in the following table:
[0077] n arithmetic circuits of the pipeline flag bit Arithmetic Circuit 1 Flag1 Arithmetic Circuit 2 Flag2 Arithmetic Circuit 3 Flag3 ... ... Arithmetic Circuit n Flagn
[0078] For example, taking the n arithmetic circuits in the currently executing pipeline as an example, each arithmetic circuit corresponds to a flag bit, and each flag bit indicates the occupancy of the corresponding arithmetic circuit.
[0079] In a possible example, the resource table is further used to record the execution duration of each currently executing instruction. The operation module obtains the execution duration through the cycle counter and determines the remaining occupancy duration of the corresponding arithmetic circuit based on the execution duration.
[0080] In specific implementation, the resource table can also be used to record the execution duration of each currently executing instruction. Since an instruction is composed of codes and the execution duration of each code is fixed, the execution position of the codes of each instruction can be calculated, and based on this, the execution duration and the remaining execution duration of each instruction can be determined. The operation module can obtain the execution duration through the cycle counter and determine the remaining occupancy duration of the corresponding arithmetic circuit, so that when the arithmetic circuit is idle, it can be allocated to other instructions or pipelines for use. In this way, on the basis of the cycle, the instructions can be better managed in order and the instruction out-of-order can be maximally achieved. Avoiding the situation of low reuse rate of functional circuits caused only by the cycle limit.
[0081] In a possible example, when pipeline i includes 6 arithmetic circuits, the 6 arithmetic circuits in the order from top to bottom according to pipeline i are respectively: a first-stage adder, a second-stage adder, a first-stage multiplier, a second-stage multiplier, the first-stage adder, and the second-stage adder, and pipeline i is any one of the K pipelines.
[0082] In specific implementation, taking pipeline i as an example, pipeline i is any one of the K pipelines. When pipeline i includes 6 arithmetic circuits, the 6 arithmetic circuits in the order from top to bottom according to pipeline i are respectively: a first-stage adder, a second-stage adder, a first-stage multiplier, a second-stage multiplier, the first-stage adder, and the second-stage adder. Of course, the specific setting of the pipeline can be determined according to the actual situation (operation resource requirements), and the pipeline can be freely configured to improve the operation efficiency of the operation module.
[0083] For example, taking two pipelines as an example, the first pipeline is LINE1 and the second pipeline is LINE2. In LINE1, six functional modules are arranged in the order from top to bottom of the pipeline. They are respectively: stage1-1-adder 1 (first-stage adder), stage1-2-adder 2 (second-stage adder), stage1-3-multiplier 1 (first-stage multiplier), stage1-4-multiplier 2 (second-stage multiplier), stage1-5-adder 1 (first-stage adder), stage1-6-adder 2 (second-stage adder). The first-stage adder and the second-stage adder are used in cooperation to complete the two-stage operation of addition. Of course, the same applies to the multiplier. It is also possible to only set a first-stage adder or multiplier in the pipeline.
[0084] In a possible example, pipeline i supports a single instruction multiple data (SIMD) hardened instruction.
[0085] In specific implementation, pipeline i supports a single instruction multiple data (SIMD) hardened instruction. Taking LINE1 as an example above, the single instruction multiple data (SIMD) hardened instructions supported by LINE1 can include: add + multiply + add (used for all 1-6), add + multiply (not used for 5 and 6), multiply + add (not used for 1 and 2), add + add (not used for 3 and 4). Specifically, add + multiply + add means that the first-stage adder, the second-stage adder, the first-stage multiplier, the second-stage multiplier, the first-stage adder, and the second-stage adder are all used, that is, all 6 arithmetic circuits are used.
[0086] In a possible example, at least two of the K pipelines support SIMD hardware instructions. Each of the at least two pipelines includes a plurality of arithmetic circuits, and the arithmetic circuit is at least one of the following: a random arithmetic circuit, an adder, a search arithmetic circuit, a configuration arithmetic circuit, a multiplier, a pooling arithmetic circuit, a comparison arithmetic circuit, a logic arithmetic circuit, an extreme value arithmetic circuit, a filtering arithmetic circuit, and an interference cancellation circuit.
[0087] In specific implementation, for example, SIMD may include at least two pipelines, and each pipeline may include 6 functional modules. Taking LINE1 as an example, for instance, add + multiply + add (used for all 1 - 6), which is six beats; or, add + multiply (not used for 5 and 6), which is four beats (add and multiply are each completed by two arithmetic circuits).
[0088] In specific implementation, if there are existing ADD + MUL + ADD (six beats) and POOL (one beat) instructions in the pipeline, then in the beat counter, record the current remaining beats of ADD + MUL + ADD (six beats). When the beats of a newly entered instruction are less than its current remaining beats, it can enter. If the beats of the newly entered instruction are greater than 6, then after it finishes execution, enter the new instruction and update the current remaining beats of the instruction with the longest beats in the beat counter.
[0089] In specific implementation, when the conditions of the beat counter are not met, the arithmetic module can control the beat counter not to receive instructions. When the conditions are met, the arithmetic module can control the beat counter to receive instructions and send the instructions to the pipeline.
[0090] It can be seen that the flow arithmetic module described in the embodiments of the present application includes K pipelines, P pipeline entrances, Q pipeline exits, and a beat counter. K, P, and Q are all integers greater than or equal to 2. The remaining beats of the first instruction with the longest required beats among the instructions being executed in the K pipelines are saved in the beat counter. The arithmetic module obtains the remaining beats of the first instruction, obtains the execution beats of the second instruction to be launched, and when the execution beats are less than the remaining beats, allows the second instruction to enter the corresponding pipeline. On the one hand, it can improve the computing efficiency based on the design of multiple pipelines with multiple entrances and exits. On the other hand, it can prevent the disorder of the output results at the output end of the pipeline.
[0091] The above Figure 1A As shown in the consistent embodiments, please refer to Figure 2 , Figure 2 is a schematic flowchart of a pipeline control method provided by an embodiment of the present application, which is applied to, such as Figure 1AThe operation module shown, the operation module includes K pipelines, P pipeline inlets, Q pipeline outlets and a cycle counter, where K, P, and Q are all integers greater than or equal to 2; the cycle counter stores the remaining cycles of the first instruction with the longest required cycles among the instructions being executed in the K pipelines; as Figure 2 shown, the pipeline control method includes:
[0092] 201. Obtain the remaining cycles of the first instruction.
[0093] 202. Obtain the execution cycles of the second instruction to be issued.
[0094] 203. When the execution cycles of the second instruction are less than the remaining cycles of the first instruction, allow the second instruction to enter the corresponding pipeline.
[0095] In specific implementation, the specific descriptions of the above steps 201 - 203 can refer to the above description and will not be elaborated here.
[0096] It can be seen that the pipeline control method described in the embodiments of the present application is applied to an operation module. The operation module includes K pipelines, P pipeline inlets, Q pipeline outlets and a cycle counter, where K, P, and Q are all integers greater than or equal to 2; the cycle counter stores the remaining cycles of the first instruction with the longest required cycles among the instructions being executed in the K pipelines. By obtaining the remaining cycles of the first instruction through the operation module, obtaining the execution cycles of the second instruction to be issued, and allowing the second instruction to enter the corresponding pipeline when the execution cycles are less than the remaining cycles, on the one hand, it can improve the computing efficiency based on the design of multiple pipelines with multiple inlets and outlets, and on the other hand, it can prevent the disorder of the output results at the output end of the pipeline.
[0097] Consistent with the above embodiments, please refer to Figure 3 , Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As shown in the figure, the electronic device includes a processor, a memory, a communication interface, an operation module, and one or more programs. The operation module includes K pipelines, P pipeline inlets, Q pipeline outlets and a cycle counter, where K, P, and Q are all integers greater than or equal to 2; the cycle counter stores the remaining cycles of the first instruction with the longest required cycles among the instructions being executed in the K pipelines; wherein, the above one or more programs are stored in the above memory and are configured to be executed by the above processor. In the embodiments of the present application, the above programs include instructions for performing the following steps:
[0098] Obtain the remaining cycles of the first instruction;
[0099] Obtain the number of execution cycles of the second instruction to be issued;
[0100] When the number of execution cycles of the second instruction is less than the remaining number of cycles of the first instruction, allow the second instruction to enter the corresponding pipeline.
[0101] It can be seen that the electronic device described in the embodiments of the present application includes an arithmetic module, the arithmetic module includes K pipelines, P pipeline entrances, Q pipeline exits, and a cycle counter. K, P, and Q are all integers greater than or equal to 2. The cycle counter stores the remaining number of cycles of the first instruction, which requires the longest number of cycles among the instructions being executed in the K pipelines. Obtain the remaining number of cycles of the first instruction, obtain the number of execution cycles of the second instruction to be issued, and when the number of execution cycles is less than the remaining number of cycles, allow the second instruction to enter the corresponding pipeline. On the one hand, it can improve the computing efficiency based on the design of multiple pipelines with multiple entrances and exits, and on the other hand, it can prevent the disorder of the output results at the output end of the pipeline.
[0102] In a possible example, the above program further includes instructions for performing the following steps:
[0103] When the number of execution cycles of the second instruction is greater than or equal to the remaining number of cycles of the first instruction, wait for the first instruction to complete execution, and after the first instruction completes execution, allow the second instruction to enter the corresponding pipeline and update the remaining number of cycles of the first instruction in the cycle counter.
[0104] In a possible example, the pipeline includes multiple pipeline stages, each pipeline stage corresponds to at least one arithmetic circuit, and the cycle counter also stores a resource table, which stores the arithmetic circuits occupied by each instruction being executed. The above program further includes instructions for performing the following steps:
[0105] Update the remaining number of cycles of the first instruction according to the resource table.
[0106] In a possible example, the above program further includes instructions for performing the following steps:
[0107] Determine whether to allow the second instruction to enter the corresponding pipeline according to the resource table and the arithmetic circuits required by the second instruction.
[0108] In a possible example, the resource table is further used to record the execution duration of each instruction being executed. The above program further includes instructions for performing the following steps:
[0109] Obtain the execution duration through the cycle counter and determine the remaining occupation duration of the corresponding arithmetic circuit according to the execution duration.
[0110] In a possible example, when pipeline i includes six arithmetic circuits, the six arithmetic circuits in the order from top to bottom according to pipeline i are: a first-stage adder, a second-stage adder, a first-stage multiplier, a second-stage multiplier, the first-stage adder, and the second-stage adder, and pipeline i is any one of the K pipelines.
[0111] In a possible example, pipeline i supports single instruction multiple data stream (SIMD) hardened instructions.
[0112] In a possible example, at least two of the K pipelines support SIMD hardware instructions. Each of the at least two pipelines includes a plurality of arithmetic circuits, and the arithmetic circuit is at least one of the following: a random arithmetic circuit, an adder, a search arithmetic circuit, a configuration arithmetic circuit, a multiplier, a pooling arithmetic circuit, a comparison arithmetic circuit, a logic arithmetic circuit, an extreme value extraction arithmetic circuit, a filtering arithmetic circuit, and an interference cancellation circuit.
[0113] The above mainly introduces the solution of the embodiment of the present application from the perspective of the execution process on the method side. It can be understood that in order for an electronic device to implement the above functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments provided in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0114] The embodiment of the present application can divide the functional units of the electronic device according to the above method examples. For example, each functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. It should be noted that the division of units in the embodiment of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0115] Please refer to Figure 4 , Figure 4It is a schematic structural diagram of a pipeline control device provided by this embodiment. The pipeline control device is applied to an electronic device, which includes an operation module. The operation module includes K pipelines, P pipeline entrances, Q pipeline exits, and a clock counter. K, P, and Q are all integers greater than or equal to 2; the clock counter stores the remaining clock cycles of the first instruction with the longest required clock cycles among the instructions being executed in the K pipelines; the pipeline control device may include: an acquisition unit 401 and an execution unit 402, where,
[0116] The acquisition unit 401 is configured to acquire the remaining clock cycles of the first instruction;
[0117] The acquisition unit 401 is further configured to acquire the execution clock cycles of the second instruction to be issued;
[0118] The execution unit 402 is configured to allow the second instruction to enter the corresponding pipeline when the execution clock cycles of the second instruction are less than the remaining clock cycles of the first instruction.
[0119] It can be seen that the pipeline control device described in the embodiment of the present application is applied to an operation module, which includes K pipelines, P pipeline entrances, Q pipeline exits, and a clock counter. K, P, and Q are all integers greater than or equal to 2; the clock counter stores the remaining clock cycles of the first instruction with the longest required clock cycles among the instructions being executed in the K pipelines. By acquiring the remaining clock cycles of the first instruction, acquiring the execution clock cycles of the second instruction to be issued, and allowing the second instruction to enter the corresponding pipeline when the execution clock cycles are less than the remaining clock cycles, on the one hand, it can improve the computing efficiency based on the design of multiple pipelines with multiple entrances and exits, and on the other hand, it can prevent the output results from being out of order at the output end of the pipeline.
[0120] In a possible example, the execution unit 402 is specifically configured to:
[0121] When the execution clock cycles of the second instruction are greater than or equal to the remaining clock cycles of the first instruction, wait for the first instruction to complete execution, and after the first instruction completes execution, allow the second instruction to enter the corresponding pipeline and update the remaining clock cycles of the first instruction in the clock counter.
[0122] In a possible example, the pipeline includes multiple pipeline stages, each pipeline stage corresponds to at least one arithmetic circuit, and a resource table is further stored in the clock counter. The resource table stores the arithmetic circuits occupied by each instruction being executed. The execution unit 402 is specifically configured to:
[0123] Update the remaining clock cycles of the first instruction according to the resource table.
[0124] In a possible example, the execution unit 402 is specifically configured to:
[0125] Determine whether to allow the second instruction to enter the corresponding pipeline according to the resource table and the arithmetic circuits required by the second instruction.
[0126] In a possible example, when pipeline i includes 6 arithmetic circuits, the 6 arithmetic circuits in the order from top to bottom of pipeline i are: a first-stage adder, a second-stage adder, a first-stage multiplier, a second-stage multiplier, the first-stage adder, and the second-stage adder, and pipeline i is any one of the K pipelines.
[0127] In a possible example, pipeline i supports single instruction multiple data stream (SIMD) hardened instructions.
[0128] In a possible example, at least two of the K pipelines support SIMD hardware instructions. Each of the at least two pipelines includes multiple arithmetic circuits, and the arithmetic circuit is at least one of the following: a random arithmetic circuit, an adder, a search arithmetic circuit, a configuration arithmetic circuit, a multiplier, a pooling arithmetic circuit, a comparison arithmetic circuit, a logic arithmetic circuit, an extreme value extraction arithmetic circuit, a filtering arithmetic circuit, and an interference cancellation circuit.
[0129] It can be understood that the functions of the program modules of the pipeline control device in this embodiment can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions of the above method embodiments, which will not be elaborated here.
[0130] Figure 5 FIG. is a structural diagram showing a combined processing device 500 according to an embodiment of the present disclosure. As Figure 5 shown therein, the combined processing device 500 includes an electronic device 502, an interface device 504, other processing devices 506, and a storage device 508. According to different application scenarios, the electronic device may include one or more computing devices 510, and the computing device may be configured to execute the operations described herein in connection with the appended Figure 1A-4 description.
[0131] In different embodiments, the electronic device of the present disclosure can be configured to perform operations specified by a user. In an exemplary application, the electronic device can be implemented as a single-core artificial intelligence processor or a multi-core artificial intelligence processor. Similarly, one or more computing devices included in the electronic device can be implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core. When multiple computing devices are implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core, the electronic device of the present disclosure can be regarded as having a single-core structure or a homogeneous multi-core structure.
[0132] In an exemplary operation, the electronic device of the present disclosure can interact with other processing devices through an interface device to jointly complete the operations specified by the user. Depending on the implementation, the other processing devices of the present disclosure can include one or more types of processors such as a central processing unit (CPU), a graphics processing unit (GPU), an artificial intelligence processor, etc., which are general-purpose and / or dedicated processors. These processors can include, but are not limited to, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and the number thereof can be determined according to actual needs. As mentioned above, only with respect to the electronic device of the present disclosure, it can be regarded as having a single-core structure or a homogeneous multi-core structure. However, when the electronic device and other processing devices are considered together, the two can be regarded as forming a heterogeneous multi-core structure.
[0133] In one or more embodiments, the other processing device can serve as an interface for external data and control of the electronic device of the present disclosure (which can be specifically implemented as an operation device related to artificial intelligence, such as neural network operations), and perform basic controls including but not limited to data transfer, start-up and / or shutdown of computing devices. In other embodiments, the other processing device can also cooperate with the electronic device to jointly complete computing tasks.
[0134] In one or more embodiments, the interface device can be used to transfer data and control instructions between an electronic device and other processing devices. For example, the electronic device can obtain input data from other processing devices via the interface device and write it into the storage device (or memory) on the chip of the electronic device. Further, the electronic device can obtain control instructions from other processing devices via the interface device and write them into the control cache on the chip of the electronic device. Alternatively or optionally, the interface device can also read the data in the storage device of the electronic device and transfer it to other processing devices.
[0135] Additionally or optionally, the combined processing device of the present disclosure may further include a storage device. As shown in the figure, the storage device is respectively connected to the electronic device and the other processing device. In one or more embodiments, the storage device can be used to store the data of the electronic device and / or the other processing device. For example, the data may be data that cannot be fully stored in the internal or on-chip storage device of the electronic device or other processing devices.
[0136] In some embodiments, the present disclosure also discloses a chip (such as Figure 6 the chip 602 shown in Figure 5 ). In one implementation, the chip is a System on Chip (SoC) and integrates one or more combined processing devices as shown in Figure 6 . The chip can be connected to other related components through an external interface device (such as the external interface device 606 shown in Figure 6 ). The related components can be, for example, a camera, a display, a mouse, a keyboard, a network card, or a wifi interface. In some application scenarios, other processing units (such as a video codec) and / or interface modules (such as a DRAM interface) etc. can be integrated on the chip. In some embodiments, the present disclosure also discloses a chip package structure that includes the above chip. In some embodiments, the present disclosure also discloses a board that includes the above chip package structure. The following will describe the board in detail with reference to Figure 6 .
[0137] Figure 6 is a schematic structural diagram of a board 600 according to an embodiment of the present disclosure. As shown in Figure 6As shown in the figure, the board includes a storage device 604 for storing data, which includes one or more storage units 610. The storage device can be connected to the control device 608 and the chip 602 described above and perform data transmission through means such as a bus. Further, the board also includes an external interface device 606, which is configured to perform data relay or transfer functions between the chip (or the chip in the chip package structure) and an external device 612 (such as a server or a computer, etc.). For example, the data to be processed can be transmitted from the external device to the chip through the external interface device. Another example is that the calculation result of the chip can be transmitted back to the external device via the external interface device. According to different application scenarios, the external interface device can have different interface forms. For example, it can adopt a standard PCIE interface, etc.
[0138] In one or more embodiments, the control device in the disclosed board can be configured to regulate the state of the chip. For this purpose, in one application scenario, the control device can include a microcontroller unit (MCU) for regulating the working state of the chip.
[0139] According to the above combination Figure 5 and Figure 6 the description of, those skilled in the art can understand that the present disclosure also discloses an electronic device or apparatus, which can include one or more of the above boards, one or more of the above chips, and / or one or more of the above combined processing devices.
[0140] According to different application scenarios, the electronic devices or apparatuses of the present disclosure may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, Internet of Things terminals, mobile terminals, mobile phones, dash cams, navigators, sensors, cameras, cameras, video cameras, projectors, watches, earphones, mobile storage, wearable devices, vision terminals, autonomous driving terminals, transportation means, household appliances, and / or medical devices. The transportation means includes airplanes, ships, and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, range hoods; the medical devices include nuclear magnetic resonance imagers, B-ultrasound devices, and / or electrocardiographs. The electronic devices or apparatuses of the present disclosure can also be applied to fields such as the Internet, Internet of Things, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and medical care. Further, the electronic devices or apparatuses of the present disclosure can also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as the cloud, edge, and terminal. In one or more embodiments, the electronic devices or apparatuses with high computing power according to the solution of the present disclosure can be applied to cloud devices (such as cloud servers), while the electronic devices or apparatuses with low power consumption can be applied to terminal devices and / or edge devices (such as smart phones or cameras). In one or more embodiments, the hardware information of cloud devices is compatible with the hardware information of terminal devices and / or edge devices, so that appropriate hardware resources can be matched from the hardware resources of cloud devices according to the hardware information of terminal devices and / or edge devices to simulate the hardware resources of terminal devices and / or edge devices, so as to complete unified management, scheduling, and collaborative work of end-cloud integration or cloud-edge-end integration.
[0141] It should be noted that, for the purpose of simplicity, some methods and their embodiments of the present disclosure are expressed as a series of actions and their combinations. However, those skilled in the art can understand that the solution of the present disclosure is not limited by the order of the described actions. Therefore, based on the disclosure or teachings of the present disclosure, those skilled in the art can understand that some of the steps can be executed in other orders or simultaneously. Further, those skilled in the art can understand that the embodiments described in the present disclosure can be regarded as optional embodiments, that is, the actions or modules involved are not necessarily required for the implementation of a certain or certain solutions of the present disclosure. In addition, according to different solutions, the present disclosure focuses on different descriptions of some embodiments. In view of this, those skilled in the art can understand the parts not detailed in a certain embodiment of the present disclosure by referring to the relevant descriptions of other embodiments.
[0142] In terms of specific implementation, based on the disclosure and teachings of this disclosure, those skilled in the art can understand that several embodiments disclosed in this disclosure can also be implemented in other ways not disclosed herein. For example, regarding each unit in the foregoing embodiments of the electronic device or apparatus, it is divided herein based on consideration of logical functions, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions of the units or components can be selectively disabled. Regarding the connection relationship between different units or components, the connections discussed above in conjunction with the drawings can be direct or indirect couplings between the units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection using an interface, where the communication interface can support signal transmission in electrical, optical, acoustic, magnetic, or other forms.
[0143] In this disclosure, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. The aforementioned components or units may be located at the same position or distributed across multiple network units. Additionally, according to actual needs, some or all of the units can be selected to achieve the purpose of the solution described in the embodiments of this disclosure. Further, in some scenarios, multiple units in the embodiments of this disclosure can be integrated into one unit or each unit physically exists separately.
[0144] In some implementation scenarios, the above-mentioned integrated units can be implemented in the form of software program modules. If implemented in the form of software program modules and sold or used as an independent product, the integrated units can be stored in a computer-readable memory. Based on this, when the solution of this disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in the memory, which may include several instructions for causing a computer device (such as a personal computer, a server, or a network device, etc.) to execute some or all of the steps of the method described in the embodiments of this disclosure. The aforementioned memory may include, but is not limited to, various media such as USB flash drives, flash memory drives, read-only memory (ROM), random access memory (RAM), external hard drives, magnetic disks, or optical discs that can store program code.
[0145] In some other implementation scenarios, the above integrated units can also be implemented in the form of hardware, i.e., specific hardware circuits, which can include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit can include, but is not limited to, physical devices, and the physical devices can include, but are not limited to, devices such as transistors or memristors. In view of this, various devices described in this article (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as CPUs, GPUs, FPGAs, DSPs, and ASICs, etc. Further, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), ROM, and RAM, etc.
[0146] An embodiment of this application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program causes the computer to execute some or all of the steps of any one of the methods described in the above method embodiments. The above computer includes an electronic device.
[0147] An embodiment of this application also provides a computer program product. The above computer program product includes a non-transitory computer-readable storage medium storing a computer program. The above computer program is operable to cause the computer to execute some or all of the steps of any one of the methods described in the above method embodiments. The computer program product can be a software installation package, and the above computer includes an electronic device.
[0148] Although multiple embodiments of this disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many changes, variations, and alternative ways can be conceived by those skilled in the art without departing from the spirit and scope of this disclosure. It should be understood that various alternatives to the embodiments of this disclosure described herein can be employed in practicing this disclosure. The appended claims are intended to define the scope of protection of this disclosure and thus cover equivalents or alternatives within the scope of these claims.
Claims
1. An operation module, characterized in that, The operation module includes K pipelines, P pipeline inlets, Q pipeline outlets, and a cycle counter. K, P, and Q are all integers greater than or equal to 2. The cycle counter stores the remaining cycles of the first instruction with the longest required cycle count among the instructions being executed in the K pipelines. The operation module is used for: Obtaining the remaining cycles of the first instruction; Obtaining the execution cycles of the second instruction to be issued; When the execution cycles of the second instruction are less than the remaining cycles of the first instruction, allowing the second instruction to enter the corresponding pipeline.
2. The arithmetic module according to claim 1, characterized in that The operation module is further specifically used for: When the execution cycles of the second instruction are greater than or equal to the remaining cycles of the first instruction, waiting for the first instruction to complete execution, and after the first instruction completes execution, allowing the second instruction to enter the corresponding pipeline and updating the remaining cycles of the first instruction in the cycle counter.
3. The arithmetic module according to claim 1 or 2, characterized in that, Each pipeline includes multiple pipeline stages, and each pipeline stage corresponds to at least one arithmetic circuit. The cycle counter also stores a resource table. The operation module is further used for: Updating the remaining cycles of the first instruction according to the resource table.
4. The arithmetic module according to claim 3, wherein The operation module is further used for: Determining whether to allow the second instruction to enter the corresponding pipeline according to the resource table and the arithmetic circuits required by the second instruction.
5. The arithmetic module according to claim 3, wherein The resource table is further used to record the execution duration of each instruction being executed. The operation module is further used to obtain the execution duration through the cycle counter and determine the remaining occupation duration of the corresponding arithmetic circuit according to the execution duration.
6. The arithmetic module according to claim 1 or 2, characterized in that At least two of the K pipelines support SIMD hardware instructions. Each of the at least two pipelines includes multiple arithmetic circuits, and the arithmetic circuit is at least one of the following: random arithmetic circuit, adder, search arithmetic circuit, configuration arithmetic circuit, multiplier, pooling arithmetic circuit, comparison arithmetic circuit, logic arithmetic circuit, extreme value extraction arithmetic circuit, filtering arithmetic circuit, and interference cancellation circuit.
7. A pipeline control method, characterized in that, Applied to an operation module, the operation module includes K pipelines, P pipeline inlets, Q pipeline outlets, and a cycle counter. K, P, and Q are all integers greater than or equal to 2. The cycle counter stores the remaining cycles of the first instruction with the longest required cycle count among the instructions being executed in the K pipelines. The method includes: Obtaining the remaining cycles of the first instruction; Obtaining the execution cycles of the second instruction to be issued; When the execution cycles of the second instruction are less than the remaining cycles of the first instruction, allowing the second instruction to enter the corresponding pipeline.
8. A neural network chip, characterized in that, The neural network chip includes the operation module according to any one of claims 1-6, or is used to execute the method according to claim 7.
9. A board card, characterized in that, The board includes: a storage device, an interface device, a control device, and the neural network chip according to claim 8; Wherein, the neural network chip is respectively connected to the storage device, the control device, and the interface device; The storage device is used for storing data; The interface device is used to realize data transmission between the chip and an external device; The control device is used to monitor the state of the chip.
10. An electronic device, characterized in that, The electronic device is configured to execute the method according to claim 7, or the electronic device includes an arithmetic module according to any one of claims 1-6, the electronic device includes a neural network chip according to claim 8, or the electronic device includes a board according to claim 9.
11. A computer-readable storage medium, characterized in that, Store a computer program for electronic data exchange, wherein the computer program causes a computer to execute the method according to claim 7.
Citation Information
Patent Citations
Order execution method and sequence processor
CN105446700A
Information processing method and terminal device
CN109997154A