Arithmetic logic unit parallel processing system and method and electronic equipment
By utilizing the parallel processing system of arithmetic logic units (ALUs), and through the collaborative work of the instruction issuing unit, write-back unit, and ALU group, the problem of low utilization of vector arithmetic logic units is solved, and efficient vector processing instruction execution is achieved.
Patent Information
- Application Number
- CN202511340992.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-19
AI Technical Summary
In the existing technology, the utilization rate of vector arithmetic logic unit is low. Increasing the number of units will reduce the utilization rate of the units, and the reduction instruction cannot make full use of the computing power.
A parallel processing system using arithmetic logic units is adopted. Through the collaborative work of instruction issue unit, write-back unit and multiple ALU groups, and by utilizing internal and external registers and multiplexers, the system implements reduction instructions and vector arithmetic logic operations, thereby reducing the number and area of arithmetic units and improving utilization.
It improves the utilization rate of the vector arithmetic logic unit, and enhances the execution efficiency of vector processing instructions and system throughput.
Smart Images

Figure CN120832124A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of chip architecture processing, and in particular to an arithmetic logic unit parallel processing system, method and electronic device. BACKGROUND
[0002] Vector arithmetic logic operation components usually contain a plurality of scalar operation components with the same number of elements in the vector operand, so that one arithmetic logic operation instruction can be sent to the operation component for operation every clock cycle. On this basis, there are two common schemes for vector reduction instructions: 1. Add one or several separate scalar logic operation components to perform reduction calculation; 2. Hardware converts the reduction instruction into a plurality of vector arithmetic logic instructions to reuse the arithmetic logic operation component.
[0003] However, in actual scenarios, scheme 1 increases the number of the same operation components, reducing the utilization rate of the operation components; and scheme 2 cannot fully utilize the computing power of the operation components due to the natural data dependency of the reduction instruction, and increases the number of instructions in the pipeline, causing the utilization rate of other operation components to also decrease. SUMMARY
[0004] Therefore, the purpose of the present application is to provide an arithmetic logic unit parallel processing system, method and electronic device, which can use the same arithmetic logic operation component to implement reduction instructions and vector arithmetic logic operations, reduce the number and area of vector arithmetic logic operation components, and improve the utilization rate of vector arithmetic logic operation components.
[0005] In a first aspect, an embodiment of the present application provides an arithmetic logic unit parallel processing system, which comprises an instruction sending unit, a write-back unit and a plurality of ALU groups comprising a plurality of arithmetic logic units; The input end of each ALU group is connected to the input end of each arithmetic logic unit in the current ALU group; and the output end of each ALU group is connected to the output end of each arithmetic logic unit in the current ALU group. The instruction sending unit is configured to send a vector processing instruction to a corresponding ALU group; The ALU group is configured to call a corresponding arithmetic logic unit according to the corresponding operand of the received vector processing instruction to process the vector processing instruction; The write-back unit is configured to aggregate the processing results output by each ALU group and call back the aggregated results to a target unit corresponding to the vector processing instruction. The write-back unit is configured to aggregate the processing results output by each ALU group and call back the aggregated results to a target unit corresponding to the vector processing instruction. The write-back unit is configured to aggregate the processing results output by each ALU group and call back the aggregated results to a target unit corresponding to the vector processing instruction.
[0006] Optionally, the arithmetic logic unit parallel processing system further comprises: an intra-group register; the intra-group register is arranged in each ALU group; The input end of the intra-group register is connected with the output end of all arithmetic logic units in the current ALU group respectively. The input end of the intra-group register is further connected with the input end of the ALU group. The output end of the intra-group register is connected with the input end of all arithmetic logic units in the current ALU group respectively. The intra-group register is used for storing the calculation result corresponding to the arithmetic logic unit in the ALU group.
[0007] Optionally, the arithmetic logic unit parallel processing system further comprises: an intra-group multiplexer; the intra-group multiplexer is arranged in each ALU group; each intra-group multiplexer is connected with the corresponding arithmetic logic unit; The output end of each intra-group multiplexer is connected with the input end of the corresponding arithmetic logic unit respectively; the input end of each intra-group multiplexer is connected with the output end of the intra-group register. The intra-group multiplexer is used for determining the arithmetic logic unit to be called in the current ALU group according to the calculation result stored in the intra-group register.
[0008] Optionally, the arithmetic logic unit parallel processing system further comprises: an inter-group multiplexer; the inter-group multiplexer is arranged between the write-back unit and the ALU group; The input end of the inter-group multiplexer is connected with the output end of each ALU group; the output end of the inter-group multiplexer is connected with the input end of the write-back unit. The inter-group multiplexer is used for summarizing the processing result output by each ALU group and determining the target unit corresponding to the summarized result.
[0009] Optionally, the arithmetic logic unit parallel processing system further comprises: an inter-group register; the inter-group register is arranged between the instruction issuing unit and the ALU group; The input end of the inter-group register is connected with the output end of the instruction issuing unit; the output end of the inter-group register is connected with the input end of each ALU group respectively. The inter-group register is used for determining the corresponding operand according to the received vector processing instruction.
[0010] Optionally, the instruction issuing unit is provided with a delay issuing module; The output end of the delay issuing module is connected with the input end of the inter-group register; the delay issuing module is used for sending the vector processing instruction to the inter-group register according to the delay strategy determined according to the type of the vector processing instruction; wherein, the vector processing instruction at least includes the vector arithmetic logic instruction and the vector reduction instruction.
[0011] In a second aspect, the present application provides an arithmetic logic unit parallel processing method, which is applied to the arithmetic logic unit parallel processing system mentioned in the first aspect; the arithmetic logic unit parallel processing system at least comprises an instruction emission unit, a write-back unit and a plurality of ALU groups comprising a plurality of arithmetic logic units; The method comprises: obtaining a type parameter of a corresponding vector processing instruction in the instruction emission unit, and determining a delay processing strategy corresponding to the vector processing instruction based on the type parameter; controlling the instruction emission unit to send the vector processing instruction to the corresponding ALU group by using the delay processing strategy, and controlling the arithmetic logic units in the ALU group to respond to the vector processing instruction; controlling the write-back unit to aggregate the processing results output by each ALU group, and calling back the aggregation results to the target unit corresponding to the vector processing instruction.
[0012] Optionally, the determining of the delay processing strategy corresponding to the vector processing instruction based on the type parameter comprises: if the type parameter corresponds to a vector arithmetic logic instruction, calculating a delay time of the arithmetic logic unit according to an element number s of an operand corresponding to the vector arithmetic logic instruction, a number m of the arithmetic logic units in each ALU group and a total number n of the arithmetic logic units; wherein the delay time is ceil(s / m) / (n / m); and ceil is a rounding up calculation; determining the delay processing strategy corresponding to the vector processing instruction based on the delay time.
[0013] Optionally, the determining of the delay processing strategy corresponding to the vector processing instruction based on the type parameter comprises: if the type parameter corresponds to a vector reduction instruction, calculating a delay time of the arithmetic logic unit according to an element number s of an operand corresponding to the vector reduction instruction and a total number n of the arithmetic logic units; wherein the delay time is ceil(s / n); and ceil is a rounding up calculation; determining the delay processing strategy corresponding to the vector processing instruction based on the delay time.
[0014] In a third aspect, the present application provides an electronic device, which comprises a processor and a memory; the memory stores computer executable instructions capable of being executed by the processor; and the processor executes the computer executable instructions to implement the steps of the arithmetic logic unit parallel processing method provided in the second aspect.
[0015] In a fourth aspect, the present application also provides a storage medium storing computer executable instructions, which, when invoked and executed by a processor, cause the processor to implement the steps of the arithmetic logic unit parallel processing method of the second aspect.
[0016] The arithmetic logic unit parallel processing system, method and electronic device provided by the embodiments of the present application comprise an instruction transmitting unit, a write-back unit and a plurality of ALU groups comprising a plurality of arithmetic logic units. The instruction transmitting unit is connected to the input end of each ALU group. The output end of each ALU group is connected to the write-back unit. The input end of each ALU group is connected to the input end of all the arithmetic logic units in the current ALU group. The output end of each ALU group is connected to the output end of all the arithmetic logic units in the current ALU group. The instruction transmitting unit is configured to send a vector processing instruction to the corresponding ALU group. The ALU group is configured to invoke the corresponding arithmetic logic unit to process the vector processing instruction according to the corresponding operand of the received vector processing instruction. The write-back unit is configured to aggregate the processing results output by each ALU group and call the aggregated results back to the target unit corresponding to the vector processing instruction. The scheme can use the same arithmetic logic operation component to implement the reduction instruction and the vector arithmetic logic operation, reduce the number and area of the vector arithmetic logic operation component, and improve the utilization rate of the vector arithmetic logic operation component.
[0017] Other features and advantages of the present application will be set forth in the descriptions below, and in part will become apparent to those skilled in the art upon examination of the following or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and attained by the structure particularly pointed out in the description, claims and drawings.
[0018] In order to make the above-mentioned objects, features and advantages of the present application more apparent, the following will describe a preferred embodiment in detail, and the accompanying drawings will be described as follows. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or the prior art description. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without any creative effort.
[0020] Figure 1 A flowchart of the interaction between the reduction instruction and the arithmetic logic instruction in the prior art provided by the embodiments of the present application; Figure 2A structural schematic diagram of an arithmetic logic unit parallel processing system provided by an embodiment of the present application is provided. Figure 3 A structural schematic diagram of another arithmetic logic unit parallel processing system provided by an embodiment of the present application is provided. Figure 4 A flow chart of an arithmetic logic unit parallel processing method provided by an embodiment of the present application is provided. Figure 5 A flow chart of a step of determining a delay processing strategy corresponding to a vector processing instruction based on a type parameter in an arithmetic logic unit parallel processing method provided by an embodiment of the present application is provided. Figure 6 A flow chart of a step of determining a delay processing strategy corresponding to a vector processing instruction based on a type parameter in another arithmetic logic unit parallel processing method provided by an embodiment of the present application is provided. Figure 7 A structural schematic diagram of an electronic device provided by an embodiment of the present application is provided.
[0021] Icon: 100 - instruction emitting unit; 200 - write back unit; 300 - ALU group; 400 - in-group register; 500 - in-group multiplexer; 600 - out-group multiplexer; 700 - out-group register; 800 - delay emitting module. 101 - processor; 102 - memory; 103 - bus; 104 - communication interface. DETAILED DESCRIPTION
[0022] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be described below in connection with the embodiments. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of the present application.
[0023] A vector arithmetic logic operation component usually contains a plurality of scalar operation components with the same number of elements in a vector operand, so that one arithmetic logic operation instruction can be emitted to the operation component for operation every clock cycle. On this basis, there are two common schemes for vector reduction instructions: 1. Adding one or several scalar logic operation components to perform reduction calculation; 2. Hardware converts the reduction instruction into a plurality of vector arithmetic logic instructions to multiplex the arithmetic logic operation component.
[0024] But in the actual scene, scheme 1 increases the number of the same operation components and reduces the utilization rate of the operation components; scheme 2 cannot fully utilize the computing capacity of the operation components due to the natural data dependency of the reduction instruction, and increases the number of instructions on the pipeline, thereby causing the utilization rate of other operation components to also decrease.
[0025] There is a reduction instruction in the V Extension of RISC-V, which reduces the element 0 of the source operand 1 and all elements of the source operand 2 through an arithmetic logic operation, and finally writes an element to the element 0 of the destination operand. Figure 1 As shown in the figure, assuming that each source operand of the instruction has 4 elements, there will be an arithmetic logic operation component in the pipeline for the calculation of the normal vector arithmetic logic instruction. Converting the reduction instruction of the addition operation into the vector arithmetic logic instruction will produce 3, due to the dependency between the operation data, the first instruction can only calculate the partial sum psum0 and psum1, that is, only 2 arithmetic logic operation components in 4 can be used. The second instruction and the third instruction calculate the partial sum psum2 and the final result sum respectively, and they can only use one arithmetic logic operation component. Psum is the abbreviation of path sum, which is a node statistics process of a binary tree in a data structure.
[0026] Based on this, the present application provides an arithmetic logic unit parallel processing system, method and electronic equipment, which can implement the reduction instruction and the vector arithmetic logic operation using the same arithmetic logic operation component, reduce the number and area of the vector arithmetic logic operation component, and improve the utilization rate of the vector arithmetic logic operation component.
[0027] In order to facilitate the understanding of the present embodiment, first, a kind of arithmetic logic unit parallel processing system disclosed in the present application is introduced in detail, as shown in the figure, the arithmetic logic unit parallel processing system includes: instruction emission unit 100, write back unit 200 and multiple ALU groups 300 comprising several arithmetic logic units. Figure 2
[0028] Among them, the instruction emission unit 100 is connected with the input end of each ALU group 300 respectively;The output end of each ALU group 300 is connected with the write back unit;The input end of each ALU group 300 is connected with the input end of all arithmetic logic units in the current ALU group 300 respectively;The output end of each ALU group 300 is connected with the output end of all arithmetic logic units in the current ALU group 300 respectively.
[0029] The instruction transmitting unit 100 is used for sending the vector processing instruction to the corresponding ALU group 300; the ALU group 300 is used for calling the corresponding arithmetic logic unit to process the vector processing instruction according to the corresponding operand of the received vector processing instruction; and the write-back unit 200 is used for collecting the processing results output by each ALU group 300 and calling the collected results back to the corresponding target unit of the vector processing instruction.
[0030] Specifically, the arithmetic logic unit parallel processing system is a hardware architecture oriented to efficient vector operation, and the core design thereof greatly improves the execution efficiency of the vector processing instruction through modular grouping and parallel cooperation mechanism. The system mainly consists of three key modules: the instruction transmitting unit 100, the write-back unit 200, and a plurality of ALU groups 300 containing a plurality of arithmetic logic units (ALUs), and the modules form a closed loop for cooperative work through a specific connection relationship.
[0031] From the hardware connection relationship, the instruction transmitting unit 100 as the instruction center of the system is directly connected with the input end of each ALU group 300, and this design ensures that the instruction can be accurately and efficiently distributed to the target processing unit. Meanwhile, the output end of each ALU group 300 is connected with the write-back unit 200, which provides a channel for the collection and return of the processing results. Inside the ALU group 300, the input end is connected with the input ends of all arithmetic logic units in the group, and the output end is connected with the output signals of all ALUs in the group. This internal interconnection design enables a single ALU group to serve as a cooperative processing unit, which can flexibly schedule the resources in the group to respond to instructions.
[0032] In terms of function implementation, the modules have clear division of labor and close cooperation. The core role of the instruction transmitting unit 100 is instruction distribution and scheduling. It receives the vector processing instruction transmitted from the upper system (such instructions usually need to perform the same or similar operations on a group of data, such as vector addition, multiplication, etc.), and according to the operation type, data size, etc. of the instruction, it determines which ALU group should process the instruction, and then accurately sends the instruction to the corresponding ALU group 300. In this process, it may also be responsible for the preliminary decoding of the instruction, to clearly determine the operation type, operand source, etc. key information, to provide a basis for the processing of the ALU group.
[0033] The ALU group 300 serves as the system's computational core, handling specific vector operations. Upon receiving a vector processing instruction, it dispatches the corresponding arithmetic logic unit (ALU) within the group to perform the operation based on the operands specified in the instruction (which may come from registers, memory, or other storage units). For example, if the instruction requires the addition of an eight-element vector, all eight ALUs within the group can be activated simultaneously, each processing one element in the vector, achieving data-level parallelism. This parallel processing mode significantly reduces the total time required for vector operations and improves system throughput.
[0034] The write-back unit 200 plays the role of result aggregation and return. It collects the calculation results output by each ALU group 300 (these results may be the values after processing each element of the vector), summarizes and integrates them according to the format required by the instruction (such as reorganizing them into complete vector results), and finally returns the summarized results to the target unit specified by the vector processing instruction (which may be a register group, main memory or other calculation unit that needs the result) to ensure that the calculation results are correctly reused.
[0035] like Figure 3 The structural diagram of another arithmetic logic unit parallel processing system is shown in FIG. Figure 3 The ALU in FPGA corresponds to arithmetic and logic unit (ALU). Figure 3 The n scalar operation units are divided into n / m groups (s, n, m are usually powers of 2). Each group can independently calculate a vector reduction instruction or a vector arithmetic logic instruction. Registers are added to the vector arithmetic logic unit to store uncalculated source operands and partial results.
[0036] Optionally, the arithmetic logic unit parallel processing system further includes: an intra-group register 400; the intra-group register 400 is disposed in each ALU group 300. The input end of the intra-group register 400 is respectively connected to the output end of all arithmetic logic units (ALUs) in the current ALU group 300; the input end of the intra-group register 400 is also connected to the input end of the ALU group 300; the output end of the intra-group register 400 is respectively connected to the input end of all arithmetic logic units (ALUs) in the current ALU group 300; and the intra-group register 400 is used to store calculation results corresponding to the arithmetic logic units (ALUs) in the ALU group 300.
[0037] The in-group register 400 is a key storage component inside each ALU group 300, and its connection relationship embodies the data closed-loop design: on the one hand, its input end is directly connected to the output end of all arithmetic logic units (ALUs) in the group, and can store the calculation results of each ALU in real time; on the other hand, its input end is also connected to the total input end of the ALU group 300, and can receive the initial data or intermediate parameters transmitted from the outside; at the same time, its output end is connected to the input end of all ALUs in the group, providing data support for subsequent operations. The core value of this design lies in reducing data access delay and repeated calculation. In multi-step vector operations (such as vector multiplication followed by addition), the intermediate results of the ALU do not need to be transmitted back and forth through external storage (such as the main memory), but can be directly reused through the in-group register 400, greatly shortening the data flow path. For example, when the ALU group executes a composite instruction of vector multiplication followed by addition, the result of the multiplication ALU can be temporarily stored in the in-group register, and the addition ALU directly reads data from the register for operation, avoiding the bandwidth occupation of the external bus and improving the continuity of the in-group operation.
[0038] Optionally, the arithmetic logic unit parallel processing system further comprises: an in-group multiplexer 500; the in-group multiplexer 500 is arranged in each ALU group; each in-group multiplexer 500 is connected to the corresponding arithmetic logic unit ALU. Wherein, the output end of each in-group multiplexer 500 is connected to the input end of the corresponding arithmetic logic unit ALU; the input end of each in-group multiplexer 500 is connected to the output end of the in-group register 400; the in-group multiplexer 500 is used for determining the arithmetic logic unit to be called in the current ALU group according to the calculation result stored in the in-group register 400.
[0039] The in-group multiplexer 500 provides precise resource switching for the ALU in each ALU group 300, and its input end is connected to the output end of the in-group register 400, and its output end is connected to the input end of the corresponding ALU. Its core function is to dynamically select the ALU to be called according to the calculation result stored in the in-group register 400. In actual operation, the processing of a vector instruction may involve multiple iterations or conditional branches (such as selecting addition or subtraction operation according to the intermediate result). For example, when a certain intermediate result stored in the in-group register is negative, the in-group multiplexer can trigger the absolute value ALU according to this result, while other elements with positive results continue to use the conventional addition ALU. This dynamic scheduling mechanism based on data enables the ALU group 300 to flexibly allocate resources according to the real-time calculation state, avoids the waste of resources caused by fixed ALU allocation, and significantly improves the adaptability of in-group parallel processing.
[0040] Optionally, the arithmetic logic unit parallel processing system further comprises: an out-of-group multiplexer 600; the out-of-group multiplexer 600 is arranged between the write-back unit 200 and the ALU group 300. The input end of the out-of-group multiplexer 600 is connected with the output end of each ALU group 300; the output end of the out-of-group multiplexer 600 is connected with the input end of the write-back unit 200; the out-of-group multiplexer 600 is used for aggregating the processing results output by each ALU group 300 and determining the target unit corresponding to the aggregated results.
[0041] The out-of-group multiplexer 600 is arranged between the write-back unit 200 and the ALU groups 300 and is a result coordination hub connecting the outputs of the multiple ALU groups and the write-back unit. The input end of the out-of-group multiplexer 600 aggregates the output results of all the ALU groups 300, and the output end is connected with the input end of the write-back unit 200. The core function of the out-of-group multiplexer 600 is to solve the conflict and orientation problems of the results of the multiple ALU groups. When multiple ALU groups process different vector instructions in parallel, the output time of the results may overlap, and the target units may be different (for example, the result of one ALU group needs to be written into register R1, and the result of another ALU group needs to be written into memory address 0x100). The out-of-group multiplexer 600 first classifies and aggregates all the results, and then accurately directs the corresponding results to the specified channel of the write-back unit 200 according to the target information (such as register number and memory address) preset in the instruction, so as to ensure that the results are not misaligned and the result return efficiency is improved when multiple groups are processed in parallel.
[0042] Optionally, the arithmetic logic unit parallel processing system further comprises: an out-of-group register 700; the out-of-group register 700 is arranged between the instruction issuing unit 100 and the ALU group 300. The input end of the out-of-group register 700 is connected with the output end of the instruction issuing unit 100; the output end of the out-of-group register 700 is connected with the input end of each ALU group 300; and the out-of-group register 700 is used for determining the corresponding operands according to the received vector processing instruction.
[0043] The out-of-group register 700 is located between the instruction issue unit 100 and each ALU group 300, forming a transition buffer for instruction-operand: its input end receives the vector processing instruction issued by the instruction issue unit 100, and its output end distributes the processed operands to each ALU group 300. Its core function is to analyze the instruction in advance and prepare the operands, breaking the waiting barrier between instruction issue and ALU execution. For example, when the instruction issue unit sends a vector addition instruction, the out-of-group register will first analyze the operand source in the instruction (such as reading from the register group or memory), load each element of the vector into the internal cache in advance, and group them according to the parallel processing requirements of the ALU group (such as 8 ALUs corresponding to 8 elements), and provide the operands directly when the ALU group is ready. This process reduces the idle time of the ALU group due to waiting for operands, making the instruction execution pipeline smoother.
[0044] Optionally, the instruction issue unit 100 is provided with a delay issue module 800; the output end of the delay issue module 800 is connected to the input end of the out-of-group register 700, and the delay issue module 800 is used to send the vector processing instruction to the out-of-group register 700 according to the delay strategy determined according to the type of the vector processing instruction; wherein the vector processing instruction at least includes vector arithmetic logic instructions and vector reduction instructions.
[0045] The delay issue module 800 is integrated in the instruction issue unit 100, and its output end is connected to the out-of-group register 700, which is responsible for dynamically adjusting the issue time according to the type of the vector processing instruction, and the core target is to avoid data conflict and resource competition. Vector processing instructions are mainly divided into two categories: Vector arithmetic logic instructions (such as vector addition and subtraction, multiplication and division): usually can be directly parallel processed, each element operation is independent, and there is no dependency; Vector reduction instructions (such as vector sum and maximum value): need to gradually summarize the intermediate results (such as first sum of the first two elements, and then add the third element), and there is a strong dependency.
[0046] The delay issue module will develop a delay strategy for different instruction types: for example, for reduction instructions, the subsequent instructions will be delayed according to the completion time of the previous step operation to avoid the ALU group idling due to data not being ready; for arithmetic logic instructions, they can be quickly issued according to the maximum parallelism, fully utilizing the ALU resources. This on-demand delay mechanism optimizes the timing of instruction issue, ensures that each ALU group obtains the correct instruction and data at the correct time, and greatly reduces the probability of pipeline blocking.
[0047] These optional components utilize a multi-layered design encompassing intra-group optimization (registers + multiplexers), inter-group coordination (extra-group registers + multiplexers), and instruction control (delayed-issue modules). This allows for more refined system features in data reuse, resource scheduling, and timing control. Working in conjunction with core modules (instruction issue, ALU group, and write-back unit), they enable the system to efficiently handle independent and parallel vector operations while flexibly responding to complex, dependent instructions. This significantly improves the adaptability of parallel processing and overall performance.
[0048] It can be seen from the arithmetic logic unit parallel processing system in the above embodiment that the system can use the same arithmetic logic operation unit to implement reduction instructions and vector arithmetic logic operations, reduce the number and area of vector arithmetic logic operation units, and improve the utilization rate of vector arithmetic logic operation units.
[0049] An embodiment of the present invention further provides an arithmetic logic unit parallel processing method, which is applied to the arithmetic logic unit parallel processing system mentioned in the above embodiment; the arithmetic logic unit parallel processing system at least includes: an instruction issuing unit, a write-back unit, and multiple ALU groups including multiple arithmetic logic units; like Figure 4 As shown, the method includes: Step S401: obtaining a type parameter of a corresponding vector processing instruction in an instruction issuing unit, and determining a delay processing strategy corresponding to the vector processing instruction based on the type parameter.
[0050] First, the system extracts the type parameters of the vector processing instruction to be processed from the instruction emission unit (for example, whether the instruction is a vector arithmetic logic instruction, a vector reduction instruction, or another specific type). These type parameters directly determine the operational characteristics of the instruction. For example, arithmetic logic instructions may support full parallel processing, while reduction instructions may have data dependencies. Based on these characteristics, the system determines the corresponding delay processing strategy: for example, an immediate emission strategy is adopted for arithmetic logic instructions without dependencies, while a delay rule is set for reduction instructions with dependencies to wait for the previous result to avoid operation conflicts.
[0051] Step S402: Using the delayed processing strategy to control the instruction issuing unit to send the vector processing instruction to the corresponding ALU group, and controlling the arithmetic logic unit in the ALU group to respond to the vector processing instruction.
[0052] According to the delay processing strategy determined in step S401, the instruction emitting unit will accurately send the vector processing instruction to the corresponding ALU group at the appropriate time (e.g. after the delay condition is met). The ALU group receiving the instruction will immediately respond: according to the required operands of the instruction (possibly from the in-group register or external storage), the corresponding arithmetic logic unit (ALU) in the group is scheduled to start parallel processing. For example, for a vector instruction containing 8 elements, the 8 ALUs in the group can simultaneously perform the same operation on different elements, realizing data-level parallelism.
[0053] Step S403: The control write-back unit uses the processing results output by each ALU group to summarize and call back the summarized results to the target unit corresponding to the vector processing instruction.
[0054] After each ALU group completes the operation, it will output the processing result to the write-back unit. The write-back unit will integrate the output results of all ALU groups (such as reorganizing into complete vector results, sorting processing results, etc.), and according to the target information (such as target register number, memory address, etc.) preset in the vector processing instruction, accurately call back the final result after summarization to the corresponding target unit, ensuring that the operation result is correctly reused or stored.
[0055] Optionally, the delay processing strategy corresponding to the vector processing instruction is determined based on the type parameter, as shown in Figure 5 , which includes: Step S501: If the type parameter corresponds to a vector arithmetic logic instruction, the delay time of the arithmetic logic unit is calculated according to the number of elements s of the operand corresponding to the vector arithmetic logic instruction, the number m of arithmetic logic units in each ALU group, and the total number n of arithmetic logic units; wherein the delay time is ceil(s / m) / (n / m); ceil is the upward rounding calculation; Step S502: Determine the delay processing strategy corresponding to the vector processing instruction based on the delay time.
[0056] For a vector arithmetic logic instruction, if the number of elements of the operand is s, then each arithmetic logic operation group can receive a new vector arithmetic logic instruction after ceil(s / m) times; there are n / m groups in total, and if all the instructions emitted are vector arithmetic logic instructions, then the vector arithmetic logic operation component as a whole can receive a new vector arithmetic logic instruction every ceil(s / m) / (n / m) times. When s, n, and m are powers of 2, the throughput is consistent with that of non-grouping, both of which are ceil(s / n).
[0057] Optionally, the delay processing strategy corresponding to the vector processing instruction is determined based on the type parameter, as shown in Figure 6 , which includes: Step S601: If the type parameter corresponds to a vector reduction instruction, calculating the delay time of the arithmetic logic unit according to the element number s of the operation number corresponding to the vector reduction instruction and the total number n of the arithmetic logic units; wherein the delay time is ceil(s / n); ceil is the ceiling calculation; Step S602: determining the delay processing strategy corresponding to the vector processing instruction based on the delay time.
[0058] For the vector reduction instruction, each arithmetic logic operation group will be different according to the number of m and the type of instruction, and each group t can emit a new vector reduction instruction; there are n / m groups in total, and if the emitted instructions are all vector reduction instructions, then the vector arithmetic logic operation component can receive a new vector reduction instruction every ceil(s / n).
[0059] From the arithmetic logic unit parallel processing method in the above embodiment, it can be known that the method can use the same arithmetic logic operation component to implement the reduction instruction and the vector arithmetic logic operation, reduce the number and area of the vector arithmetic logic operation component, and improve the utilization rate of the vector arithmetic logic operation component.
[0060] The arithmetic logic unit parallel processing method provided by the embodiment of the present application has the same implementation principle and technical effects as the arithmetic logic unit parallel processing system embodiment, and for brevity of description, the part not mentioned in the system embodiment can refer to the corresponding content in the arithmetic logic unit parallel processing method embodiment.
[0061] The embodiment also provides an electronic device, and a structure diagram of the electronic device is shown in Figure 7 The device includes a processor 101 and a memory 102; wherein the memory 102 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the steps of the above arithmetic logic unit parallel processing method.
[0062] Figure 7 The electronic device shown in the figure also includes a bus 103 and a communication interface 104, and the processor 101, the communication interface 104 and the memory 102 are connected through the bus 103.
[0063] The memory 102 can include a high-speed random access memory (RAM, Random Access Memory), and can also include a non-volatile memory, for example, at least one disk memory. The bus 103 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 7Only one bidirectional arrow is used to represent multiple buses or multiple types of buses.
[0064] The communication interface 104 is configured to connect with at least one user terminal and other network units through a network interface, and transmit the encapsulated IPv4 packet or the IPv4 packet to the user terminal through the network interface.
[0065] The processor 101 can be an integrated circuit chip with processing capability. In the implementation process, the steps of the above method can be completed by the integrated logic circuit or the instruction of the software form in the processor 101. The processor 101 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as a hardware code processor to execute, or be executed by a combination of hardware and software modules in the code processor. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory 102, and the processor 101 reads the information in the memory 102, and combines the hardware to complete the steps of the method of the foregoing embodiments.
[0066] The embodiment of the present application further provides a storage medium, and the storage medium stores a computer program. When the computer program is run by a processor, the steps of the arithmetic logic unit parallel processing method in the foregoing embodiment are executed.
[0067] In several embodiments provided in the present application, it should be understood that the disclosed system, device, apparatus and method can be implemented in other manners. The above described system embodiments are merely illustrative. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, or a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, or the among different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0068] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. In actual implementation, some or all of the units can be selected according to the actual needs to achieve the purposes of the embodiments of the present application.
[0069] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit.
[0070] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0071] Finally, it should be noted that the above-described embodiments are merely specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, but not to limit the same. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that any person skilled in the art can still modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some of the technical features, within the technical scope disclosed by the present application. The modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An arithmetic logic unit parallel processing system, characterized by, The arithmetic logic unit parallel processing system includes: an instruction issuing unit, a write-back unit, and a plurality of ALU groups including a plurality of arithmetic logic units; The instruction issuing unit is connected to the input end of each of the ALU groups respectively; the output end of each of the ALU groups is connected to the write-back unit; The input end of each of the ALU groups is respectively connected to the input ends of all the arithmetic logic units in the current ALU group; the output end of each of the ALU groups is respectively connected to the output ends of all the arithmetic logic units in the current ALU group; The instruction issuing unit is used to send the vector processing instruction to the corresponding ALU group; The ALU group is used to call the corresponding arithmetic logic unit according to the operand corresponding to the received vector processing instruction to process the vector processing instruction; The write-back unit is used to summarize the processing results output by each of the ALU groups and call back the summary results to the target unit corresponding to the vector processing instruction.
2. The system of claim 1, wherein, The arithmetic logic unit parallel processing system further includes: an intra-group register; the intra-group register is arranged in each of the ALU groups; The input terminals of the registers in the group are respectively connected to the output terminals of all the arithmetic logic units in the current ALU group; The input end of the register in the group is also connected to the input end of the ALU group; The output ends of the intra-group registers are respectively connected to the input ends of all the arithmetic logic units in the current ALU group; The intra-group register is used to store calculation results corresponding to the arithmetic logic units in the ALU group.
3. The system of claim 2, wherein, The arithmetic logic unit parallel processing system further includes: an intra-group multiplexer; the intra-group multiplexer is arranged in each of the ALU groups; each of the intra-group multiplexers is connected to the corresponding arithmetic logic unit; The output end of each of the multiplexers within the group is connected to the input end of the corresponding arithmetic logic unit; the input end of each of the multiplexers within the group is connected to the output end of the register within the group; The intra-group multiplexer is used to determine the arithmetic logic unit to be called in the current ALU group according to the calculation result stored in the intra-group register.
4. The system of claim 1, wherein, The arithmetic logic unit parallel processing system further includes: an external multiplexer; the external multiplexer is arranged between the write-back unit and the ALU group; The input end of the external multiplexer is connected to the output end of each of the ALU groups; the output end of the external multiplexer is connected to the input end of the write-back unit; The outer-group multiplexer is used to summarize the processing results output by each of the ALU groups and determine the target unit corresponding to the summary result.
5. The system of claim 1, wherein, The arithmetic logic unit parallel processing system further includes: an external register; the external register is arranged between the instruction issuing unit and the ALU group; The input end of the external register is connected to the output end of the instruction issuing unit; the output end of the external register is respectively connected to the input end of each of the ALU groups; The out-of-group register is used to determine the corresponding operation number according to the received vector processing instruction.
6. The system of claim 5, wherein, The instruction emission unit is provided with a delay emission module; The output end of the delay emission module is connected with the input end of the out-of-group register, and the delay emission module is used to send the vector processing instruction to the out-of-group register according to the delay strategy determined according to the type of the vector processing instruction; wherein the vector processing instruction at least includes vector arithmetic logic instruction and vector reduction instruction above-mentioned two types.
7. An arithmetic logic unit parallel processing method, characterized by, The method is applied to the arithmetic logic unit parallel processing system in any one of claims 1 to 6; The arithmetic logic unit parallel processing system at least includes: instruction emission unit, write back unit and multiple ALU groups containing several arithmetic logic units; The method includes: Obtaining the type parameter of the corresponding vector processing instruction in the instruction emission unit, and determining the delay processing strategy corresponding to the vector processing instruction based on the type parameter; Using the delay processing strategy to control the instruction emission unit to send the vector processing instruction to the corresponding ALU group, and controlling the arithmetic logic unit in the ALU group to respond to the vector processing instruction; Controlling the write back unit to summarize the processing results output by each ALU group, and calling back the summary results to the target unit corresponding to the vector processing instruction.
8. The arithmetic logic unit parallel processing method of claim 7, wherein, Based on the type parameter, the delay processing strategy corresponding to the vector processing instruction is determined, including: If the type parameter corresponds to vector arithmetic logic instruction, the delay time of the arithmetic logic unit is calculated according to the element number s of the corresponding operation number of the vector arithmetic logic instruction, the number m of arithmetic logic units in each ALU group and the total number n of the arithmetic logic units; wherein the delay time is ceil(s / m) / (n / m); ceil is the upward rounding calculation; Based on the delay time, the delay processing strategy corresponding to the vector processing instruction is determined.
9. The arithmetic logic unit parallel processing method of claim 7, wherein, Based on the type parameter, the delay processing strategy corresponding to the vector processing instruction is determined, including: If the type parameter corresponds to vector reduction instruction, the delay time of the arithmetic logic unit is calculated according to the element number s of the corresponding operation number of the vector reduction instruction and the total number n of the arithmetic logic units; wherein the delay time is ceil(s / n); ceil is the upward rounding calculation; Based on the delay time, the delay processing strategy corresponding to the vector processing instruction is determined.
10. An electronic device, comprising: The processor and the memory are included, the memory stores computer executable instructions capable of being executed by the processor, and the processor executes the computer executable instructions to realize the steps of the arithmetic logic unit parallel processing method in any one of claims 7 to 9.
Citation Information
Patent Citations
System and method for an asynchronous processor with pipelined arithmetic and logic unit
CN105393211A
Deep neural network hardware accelerator device
CN116451752A
ALU array, processor based on ALU array and message processing method
CN120215876A
Single instruction multiple data processor including scalar arithmetic lotgic unit
CN1519704A
Processing unit, device comprising two processing units, method for testing a processing unit and a device comprising two processing units
US20100257343A1
Cited By
Vector reduction instruction execution system and method and storage medium
CN121957912A
Vector reduction instruction execution system, method, and storage medium
CN121957912B