ALU array, processor based on ALU array and message processing method
By splitting the ALU into computing units of multiple computing types and building an ALU array with the switching unit, the problems of low message processing efficiency and ALU idleness in the prior art are solved, and more efficient ALU resource utilization and message processing efficiency are achieved.
Patent Information
- Application Number
- CN202311831899.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has low efficiency in message processing, resulting in low utilization of ALU logical resource, and only one ALU is working, while other ALUs are idle.
By splitting the ALU into computing units of multiple computing types, an ALU array is constructed with multiple switching units, each switching unit receives and sends it to the corresponding computing unit for processing, thereby achieving simultaneous processing of multiple pending sub-instructions.
The utilization rate of ALU logical resources and the efficiency of packet processing are improved, the problem of ALU idleness is solved, and the parallel processing of multiple pending sub-instructions is realized.
Smart Images

Figure CN120215876A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to an ALU array, a processor based on the ALU array, and a packet processing method. Background Art
[0002] In a programmable packet processing pipeline, dozens or even hundreds of Arithmetic and Logic Units (ALUs) provide its large-scale parallel processing performance. Among them, each ALU supports instruction functions such as arithmetic operations, comparison operations, bit operations, and logical operations internally, and usually needs to support the entire instruction set.
[0003] During the packet forwarding process, according to the lookup result of the Match lookup table engine, it is determined how to process the packet header, that is, Action. Action is mapped to several ALUs through a compiler, and different instructions are respectively configured to operate on different packet header fields and update the results to the target packet header fields.
[0004] However, the prior art has low efficiency in packet processing. Summary of the Invention
[0005] This application provides an ALU array, a processor based on the ALU array, and a packet processing method to solve the problem of low efficiency of the prior art in packet processing.
[0006] In a first aspect, an embodiment of this application provides an ALU array, including:
[0007] Multiple switching units and computing units of multiple operation types;
[0008] Each switching unit is connected to at least one of the remaining switching units, and each switching unit is also connected to computing units of at least one operation type;
[0009] Each switching unit is configured to receive the sub-instructions to be processed mapped by an instruction decoder, and send the sub-instructions to the computing unit that is connected to the switching unit and corresponds to the operation type required by the sub-instructions to be processed.
[0010] In a possible design, the operation type includes at least one of the following: arithmetic operation, bit operation, comparison operation, and logical operation.
[0011] In a possible design, each switching unit is connected to the remaining switching units according to a topological structure, and the topological structure includes any one of the following: linear topology, triangular topology, or polygonal topology.
[0012] In a possible design, there is at least one target computing unit among the computing units included in the ALU array, and each target computing unit is connected to at least two switching units.
[0013] In a possible design, each switching unit is connected to four computing units of different operation types.
[0014] In a second aspect, an embodiment of the present application provides a processor based on an ALU array, including:
[0015] An instruction decoder and an arithmetic logic unit ALU array as shown in the first aspect and each possible design;
[0016] The instruction decoder is connected to the ALU array;
[0017] The instruction decoder is configured to map the to-be-processed sub-instruction to the switching unit corresponding to the position according to the position of each to-be-processed sub-instruction in the to-be-processed instruction, and the to-be-processed instruction is generated by a compiler compiling all actions Action corresponding to the to-be-processed message.
[0018] In a possible design, the processor further includes:
[0019] A first memory, the first memory is connected to the ALU array, and is used to store the to-be-processed instruction and the processing result of each to-be-processed sub-instruction.
[0020] In a possible design, the processor further includes:
[0021] A second memory, the second memory is connected to the ALU array, and the second memory is used to store the operands required for processing the to-be-processed instruction.
[0022] In a third aspect, an embodiment of the present application provides a message processing method, which is applied to the processor based on an ALU array as described in the second aspect and each possible design. The method includes:
[0023] The compiler compiles all actions Action corresponding to the to-be-processed message to generate a to-be-processed instruction, and the to-be-processed instruction includes a plurality of to-be-processed sub-instructions;
[0024] The instruction decoder maps the to-be-processed sub-instruction to the switching unit corresponding to the position according to the position of each to-be-processed sub-instruction in the to-be-processed instruction;
[0025] For each switching unit, the to-be-processed sub-instruction obtained is processed by the target computing unit connected to the switching unit to obtain the processing result of the to-be-processed sub-instruction, and the operation type of the target computing unit is the operation type required by the to-be-processed sub-instruction.
[0026] In a possible design, for each switch unit, the target calculation unit connected to the switch unit processes the obtained sub-instructions to be processed to obtain the processing result of the sub-instructions to be processed, including:
[0027] For each switch unit, obtain the target data that the sub-instructions to be processed need to process;
[0028] The switch unit sends the sub-instructions to be processed and the target data to the corresponding target calculation unit;
[0029] The target calculation unit processes the target data according to the sub-instructions to be processed to generate the processing result of the sub-instructions to be processed;
[0030] The switch unit obtains the processing result of the sub-instructions to be processed from the target calculation unit.
[0031] In a possible design, the method further includes:
[0032] The switch unit sends the processing result of the sub-instructions to be processed to the first memory;
[0033] Or,
[0034] The switch unit sends the processing result of the sub-instructions to be processed to the first switch unit, and the first switch unit is used to reprocess the processing result of the switch unit;
[0035] Or,
[0036] The switch unit sends the processing result of the sub-instructions to be processed to the second switch unit, and the second switch unit is used to send the processing result of the switch unit to the first memory.
[0037] In a possible design, for each switch unit, obtaining the target data that the sub-instructions to be processed need to process includes:
[0038] The switch unit obtains the operands that the sub-instructions to be processed need to process from the first memory and / or the second memory;
[0039] Or,
[0040] If the sub-instructions to be processed indicate reprocessing the processing result of the third switch unit, the switch unit obtains the processing result of the third switch unit from the third switch unit;
[0041] The target data is the sub-data to be processed or the processing result of the third switch unit.
[0042] In a possible design, the compiler compiles all Actions corresponding to the packet to be processed to generate an instruction to be processed, including:
[0043] The compiler compiles all Actions corresponding to the packet to be processed to obtain a sub-instruction to be processed corresponding to each Action;
[0044] For each sub-instruction to be processed, the compiler determines a switch unit capable of processing the sub-instruction according to the operation type of the computing unit connected to each switch unit;
[0045] The compiler determines the position of the sub-instruction to be processed in the instruction to be processed according to the mapping relationship and the switch unit corresponding to each sub-instruction to be processed, where the mapping relationship is used to illustrate the correspondence between the position in the instruction to be processed and the switch unit;
[0046] The compiler generates the instruction to be processed according to the position of each sub-instruction to be processed in the instruction to be processed.
[0047] The ALU array, the processor based on the ALU array, and the packet processing method provided by the embodiments of the present application. The ALU array includes a plurality of switch units and computing units of multiple operation types. Among them, each switch unit is connected to at least one of the remaining switch units, and each switch unit is also connected to at least one computing unit of one operation type. Each switch unit is used to receive the sub-instruction to be processed mapped by the instruction decoder and send the sub-instruction to the computing unit connected to the switch unit and corresponding to the operation type required by the sub-instruction to be processed. By splitting the ALU into computing units of multiple operation types in the embodiments of the present application, multiple computing units can process multiple sub-instructions to be processed simultaneously, solving the problem in the prior art that only one ALU is working and other ALUs are idle, and improving the utilization rate of ALU logic resources and the efficiency of packet processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings here are incorporated into the description and constitute a part of this description, showing embodiments consistent with the present application and used together with the description to explain the principles of the present application.
[0049] Figure 1 It is a schematic structural diagram of the first embodiment of the ALU array provided by the embodiment of the present application;
[0050] Figure 2 It is a schematic structural diagram of the second embodiment of the ALU array provided by the embodiment of the present application;
[0051] Figure 3 It is a schematic structural diagram of the third embodiment of the ALU array provided by the embodiment of the present application;
[0052] Figure 4 Schematic diagram of the first embodiment of the processor based on the ALU array provided by the embodiment of the present application;
[0053] Figure 5 Schematic diagram of the ALU array structure in mode one provided by the embodiment of the present application;
[0054] Figure 6 Schematic diagram of the ALU array structure in mode two provided by the embodiment of the present application;
[0055] Figure 7 Schematic diagram of the ALU array structure in multiple modes provided by the embodiment of the present application;
[0056] Figure 8 Schematic diagram of the second embodiment of the processor based on the ALU array provided by the embodiment of the present application;
[0057] Figure 9 Schematic diagram of the third embodiment of the processor based on the ALU array provided by the embodiment of the present application;
[0058] Figure 10 Schematic diagram of the process of the first embodiment of the message processing method provided by the embodiment of the present application;
[0059] Figure 11 Schematic diagram of the process of the first embodiment of the message processing method provided by the embodiment of the present application.
[0060] Through the above-mentioned drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0062] Before introducing the embodiments of the present application, the terms related to the embodiments of the present application will be explained first.
[0063] Protocol-Independent Switch Architecture (PISA):
[0064] The programmable switching and routing architecture is a technology for building high-performance network devices. It features flexibility and scalability to meet the growing demands of network traffic and complexity. In this context, an important architecture is the programmable switch architecture, and PISA is a widely adopted architecture.
[0065] The PISA architecture provides a programmable data plane processing method, enabling network devices to handle multiple protocols and functions, not limited to specific protocols or applications. Traditional network devices usually focus on specific protocol processing. The PISA architecture decouples the packet processing logic from the device's hardware and transfers it to programmable software or hardware modules, thus achieving greater flexibility and scalability.
[0066] Reconfigurable Match Tables (RMT):
[0067] The programmable switching and routing architecture is a technology for building high-performance network devices. It has flexibility and scalability and can meet the growing demands of network traffic and complexity. In this context, an important architecture is the RMT architecture, which is widely used in the Barefoot Tofino chip.
[0068] The RMT architecture is a data plane processing architecture based on programmable match tables, used to implement high-performance and highly customizable network devices. In traditional network devices, packet processing is usually completed by fixed-function hardware logic. The RMT architecture realizes flexible packet processing by transferring the packet processing logic to the programmable match table. The match table is a key data structure that stores the rules for network devices to match and operate on packets.
[0069] In the RMT architecture, the match table consists of multiple entries, each entry containing a match field and an associated set of actions. When a packet enters the network device, its key fields (such as source IP address, destination IP address, port number, etc.) are matched against the entries in the match table. Once a match is successful, the relevant actions will be triggered and executed, such as modifying the packet header, redirecting traffic, updating counters, etc.
[0070] The Match-Action-Unit (MAU) is a programmable processing unit in the RMT architecture, responsible for processing the actions of the entries in the match table. Each MAU typically contains a set of processing logics that can be customized according to specific requirements. For example, one MAU can implement the header parsing and modification of data packets, and another MAU can be responsible for traffic classification and routing. By combining these MAUs together, complex network devices can be constructed to flexibly adapt to different network requirements.
[0071] Very Long Instruction Word (VLIW):
[0072] VLIW is a parallel computing architecture designed to improve the performance and efficiency of processors. The core idea of VLIW technology is to include multiple operation instructions in one instruction, and these instructions can be executed in parallel. Compared with the traditional Single Instruction, Single Data (SISD) architecture, VLIW can execute multiple operations simultaneously in one clock cycle, thus improving the computing throughput.
[0073] In the Tofino chip, VLIW is used for the processing of the network data plane. Network data plane processing involves a large number of data packet parsing, classification, and operations, and these operations can be executed in parallel in one clock cycle. The VLIW architecture realizes a high degree of instruction-level parallelism by packing multiple operation instructions into one very long instruction word and sending it to the underlying processing unit.
[0074] One of the benefits of using VLIW technology is that it can effectively utilize hardware resources. The processing unit in the Tofino chip can execute multiple operations simultaneously without replicating multiple individual processing units. This can reduce the hardware complexity and cost and provide a higher performance density.
[0075] Another advantage is that the VLIW architecture has a lower instruction latency. Since multiple operation instructions are executed in parallel in the same instruction, the waiting time between instructions is reduced, thus reducing the overall instruction latency. This is crucial for high-performance network devices and can provide faster data packet processing speed and lower latency.
[0076] VLIW ALU Array:
[0077] In the Barefoot Tofino chip, VLIW instructions generate a 5440-bit instruction set, which contains the instructions of 224 ALUs. These instructions are sent to the underlying processing unit, which consists of an array of 224 ALUs.
[0078] Each ALU is an independent processing unit responsible for performing the arithmetic and logical operations specified in the instructions. These ALUs are arranged in an array in the underlying processing unit and operate in parallel. By executing the instructions of multiple ALUs in parallel, the underlying processing unit can achieve a high degree of parallel computing power.
[0079] Each instruction in the VLIW instruction contains the operation code and operands of the corresponding ALU. When the instruction reaches the underlying processing unit, the instruction is decoded and the corresponding operation is sent to the corresponding ALU for execution. Each ALU performs the corresponding arithmetic or logical operation according to the instruction and returns the result to the data path.
[0080] The design of the array composed of 224 ALUs in the underlying processing unit can make full use of hardware resources and provide a high degree of computing power. This design enables the Barefoot Tofino chip to process network packets with high performance and efficiency, supporting complex network processing and data plane operations.
[0081] Next, the application background involved in this application will be explained.
[0082] Currently, after receiving a packet, a network device matches the key fields of the packet (such as source IP address, destination IP address, port number, etc.) with the entries in the match table to determine the action (i.e., action) that needs to be performed on the packet header. Specifically, the match table consists of multiple entries, and each entry includes a match field and a set of actions associated with the match field. Then, the key fields of the packet are matched with the match fields of each entry in the match table. If the key field of the packet is the same as the match field of any entry, it is considered that the packet matches the entry successfully. Once the match is successful, the action corresponding to the entry will be triggered and executed.
[0083] Furthermore, after determining the Action that needs to be performed on the packet header, different instructions need to be configured through the compiler and mapped to several ALUs to operate on different packet header fields and update the results to the target packet header fields.
[0084] However, at each moment, only one instruction in the instruction set is executed by the ALU, and the logic of other ALUs is in an idle state. This inevitably causes waste of ALU logic resources and results in low efficiency in packet processing.
[0085] Based on the above technical problems, the present application proposes an ALU array. In this ALU array, the ALU is split into multiple computing units according to the operation type, and together with multiple switch units, an ALU array is constructed. In this way, when the compiler maps the sub-instructions to be processed to the corresponding switch units through the instruction decoder, each switch unit can simultaneously process the sub-instructions to be processed through the computing units, thereby achieving the purpose of simultaneously processing multiple sub-instructions to be processed, improving the utilization rate of ALU logic resources and the efficiency of message processing.
[0086] Next, the technical solution of the present application will be described in detail through specific embodiments.
[0087] It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0088] First, an explanation of the ALU array provided by the present application will be given.
[0089] The ALU array provided by the present application includes multiple switch units and computing units of multiple operation types.
[0090] In the embodiments of the present application, the ALU is divided into computing units of multiple operation types according to the instruction function, and together with multiple switch units (English: Switch), a computing resource pool array (i.e., the ALU array) is constructed to perform corresponding Actions on the message header fields. That is to say, the ALU array includes multiple switch units and computing units of multiple operation types, each switch unit is connected to at least one of the remaining switch units, and each switch unit is also connected to at least one computing unit.
[0091] Among them, data can be transmitted between the connected switch units.
[0092] Optionally, when each switch unit is connected to multiple computing units, the operation types of these multiple computing units are different. For example, assuming that a switch unit is connected to 4 computing units, these 4 computing units belong to different operation types.
[0093] Optionally, the operation type includes at least one of the following: arithmetic operation, bit operation, comparison operation, and logical operation.
[0094] Among them, arithmetic operations are used to implement addition or subtraction, etc., bit operations are used to implement displacement or rotation, etc., comparison operations are used to compare the magnitude relationship between two values, and logical operations are used to implement AND, OR, NOT, XOR, etc.
[0095] Optionally, each switch unit is connected to the remaining switch units according to a topological structure, and the topological structure includes any one of the following: linear topology, triangular topology, or polygonal topology.
[0096] Exemplarily, the polygon topology is a quadrilateral, pentagon, hexagon, and so on.
[0097] Optionally, there is at least one target computing unit among the computing units included in the ALU array, and each target computing unit is connected to at least two switching units. In this way of connection, multiple switching units can share the same computing unit, making the overall utilization rate of the computing unit higher, that is, the vacancy rate lower, so that the number of computing units can be reduced in hardware, thereby greatly reducing the hardware resource overhead.
[0098] Figure 1 This is a schematic structural diagram of the first embodiment of the ALU array provided by the embodiment of the present application. As Figure 1 shown, Figure 1 in the "S" is a switching unit, "A", "B", "C", "D" are four computing units with different operation types, and the operation types of "A", "B", "C", "D" are arithmetic operation, comparison operation, bit operation, and logical operation respectively. In the Figure 1 shown ALU array, each switching unit is connected to the remaining switching units according to the quadrilateral topology, and each switching unit is connected to 4 adjacent switching units. At the same time, each switching unit is respectively connected to 4 computing units with different operation types, and is responsible for sending the sub-instructions and operands to be processed to the computing units it is connected to, and recycling the processing results of the computing units.
[0099] Figure 2 This is a schematic structural diagram of the second embodiment of the ALU array provided by the embodiment of the present application. As Figure 2 shown, each switching unit is connected to the remaining switching units according to the triangular topology, and each switching unit is connected to 6 adjacent switching units. At the same time, each switching unit is connected to at least one computing unit.
[0100] Figure 3 This is a schematic structural diagram of the third embodiment of the ALU array provided by the embodiment of the present application. As Figure 3 shown, each switching unit is connected to the remaining switching units according to the pentagonal topology, and each switching unit is connected to 3 adjacent switching units. At the same time, each switching unit is connected to 3 computing units.
[0101] It should be understood that in the ALU array structures shown in the above Figure 1 , Figure 2 and Figure 3 each computing unit can be shared by adjacent switching units, effectively improving the overall utilization rate of the computing unit, reducing the vacancy rate of the computing unit, reducing the number of computing units in hardware, and thus greatly reducing the hardware resource overhead.
[0102] It should be understood that there may be other connection methods between the switch units, and there are also other connection methods between the switch units and the computing units, which can be determined according to the actual situation and are not specifically limited herein.
[0103] The ALU array provided by the embodiments of the present application includes multiple switch units and computing units of multiple operation types. Among them, each switch unit is connected to at least one of the remaining switch units, and each switch unit is also connected to at least one computing unit of one operation type. Each switch unit is used to receive the sub-instructions to be processed mapped by the instruction decoder, and send the sub-instructions to be processed to the computing unit that is connected to the switch unit and corresponds to the operation type required by the sub-instructions to be processed. By splitting the ALU into computing units of multiple operation types, multiple computing units can simultaneously process multiple sub-instructions to be processed, solving the problem in the prior art that only one ALU is working and the other ALUs are idle, and improving the utilization rate of ALU logic resources and the efficiency of message processing.
[0104] Next, the processor based on the ALU array provided by the embodiments of the present application will be explained.
[0105] Figure 4 It is a schematic structural diagram of Embodiment 1 of the processor based on the ALU array provided by the embodiments of the present application. As Figure 4 shown, the processor 10 based on the ALU array includes: an instruction decoder 11 and an ALU array 12 connected to the instruction decoder 11.
[0106] It should be understood that the ALU array 12 can refer to the content of the embodiment shown in Figures 1 to 3 and will not be elaborated herein.
[0107] Among them, the instruction decoder 11 is used to map the sub-instructions to be processed to the switch unit corresponding to this position according to the position of each sub-instruction to be processed in the instruction to be processed, and the instruction to be processed is generated by the compiler compiling all Actions corresponding to the message to be processed.
[0108] In practical applications, after obtaining the message to be processed, the key fields of the message to be processed are matched with the entries in the matching table, so as to determine the Actions that need to be executed on the message header fields of the message to be processed. The compiler compiles each Action corresponding to the message to be processed to generate the corresponding sub-instructions to be processed. Then, the switch unit capable of processing the sub-instructions to be processed is determined from the ALU array 12, so as to determine the position of the sub-instructions to be processed in the instruction to be processed. Among them, there is a mapping relationship between the position of the instruction to be processed and the switch position, and this mapping relationship is pre-configured.
[0109] Furthermore, the allocation algorithm deployed in the decoder can ensure that the computing units are not simultaneously issued with sub-instructions to be processed by multiple switching units, preventing unpredictable conflicts. However, the processing results of the computing units can be used by multiple switching units.
[0110] Exemplarily, sub-instruction 1 to be processed and sub-instruction 2 to be processed are respectively used to perform corresponding Actions on different packet headers, and the corresponding operation types are both arithmetic operations. Switching unit 1 and switching unit 2 share a computing unit 1 with an arithmetic operation type. Switching unit 3 and switching unit 4 share a computing unit 2 with an arithmetic operation type. Then, when the compiler generates sub-instruction 1 to be processed and sub-instruction 2 to be processed, if it is determined that switching unit 1 processes sub-instruction 1 to be processed, it is necessary to determine that switching unit 3 or switching unit 4 processes sub-instruction 2 to be processed, and it cannot be determined that switching unit 2 processes sub-instruction 2 to be processed, so as to ensure that computing unit 1 is not simultaneously called by switching unit 1 and switching unit 2.
[0111] Optionally, when the compiler compiles each Action corresponding to the packet to be processed, the ALU array 12 has the following two configuration modes:
[0112] Mode 1
[0113] Suppose the Action indicates to perform a calculation on a packet header field. The compiler compiles this Action to obtain a corresponding sub-instruction to be processed. In this case, each switching unit in the ALU array 12 can form a pipelined system with a 1-unit cycle delay, that is, different switching units are not interconnected, and only their own gating functions and the computing units connected to them can be used. Each switching unit receives one data item and outputs one data item at each clock (abbreviation: clk).
[0114] Figure 5 Schematic diagram of the ALU array structure in Mode 1 provided by the embodiment of the present application. As Figure 5 shown, each switching unit enclosed by a circle is a pipelined system with a 1-unit cycle delay. The switching unit enclosed by the upper left circle can directly obtain the operands, issue the sub-instruction to be processed and the operands to the computing unit connected to it, recycle the processing results of the computing unit, and output the processing results. The other switching unit enclosed by a circle cannot directly obtain the operands and needs to obtain the operands through the upper switching unit. Then, it issues the sub-instruction to be processed and the operands to the computing unit connected to it, recycles the processing results of the computing unit, and outputs the processing results through the upper switching unit.
[0115] Mode 2
[0116] Suppose the Action indicates a calculation on a packet header field. The compiler compiles the Action to obtain multiple to-be-processed sub-instructions. The compiler configures multiple Ss in the ALU array 12 to form a system with multiple unit cycle delays, which is used to perform multi-level complex calculations on the corresponding packet header field for the Action.
[0117] Taking the system composed of 2 Ss with 2 unit cycle delays as an example for illustrative description, Figure 6 This is a schematic diagram of the ALU array structure in Mode 2 provided by the embodiment of the present application. As Figure 6 shown, the two switch units enclosed by the circle form a system with 2 unit cycle delays. The upper switch unit in this system obtains operands, issues the to-be-processed sub-instructions and operands to the computing unit it is connected to, and retrieves the processing result of the computing unit. Then, it transmits this processing result to the lower switch unit. The lower switch unit issues the to-be-processed sub-instructions and the processing result of the first switch unit to the computing unit it is connected to, so that the computing unit can further calculate the processing result of the upper switch unit. After the processing is completed, the second switch unit retrieves the processing result of this computing unit, and the retrieved processing result is the processing result obtained by performing 2-level calculations on a packet header field.
[0118] It should be understood that in an extreme case, the compiler can connect all the switch units (assuming there are M in total) in series in any path to form a pipelined system with M unit cycle delays, which can perform M operations on a packet header field.
[0119] It should be understood that the above Mode 1 and Mode 2 can exist simultaneously. That is to say, when processing the to-be-processed packets, there can be a pipelined system with 1 unit cycle delay and a pipelined system with multiple unit cycle delays in the ALU array 12 at the same time.
[0120] Exemplarily, Figure 7 This is a schematic diagram of the ALU array structure in multiple modes provided by the embodiment of the present application. As Figure 7 shown, the switch unit in Mode 1 can call the adjacent computing unit to perform operations on the input data and output. The switch unit in Mode 2 can obtain operands from the upper switch unit, select and pass them through the interconnection between switch units for interaction. In this way, data can be transferred between different groups of Packet Header Vectors (PHVs), or complex multi-expression composite operations can be supported. The switch unit in Mode 3 selects the input data, does not perform other calculations and outputs.
[0121] The processor based on the ALU array provided by the embodiment of the present application includes an instruction decoder and an arithmetic logic unit (ALU) array connected to the instruction decoder. The ALU array includes a plurality of switch units and computing units of multiple operation types. Each switch unit is connected to at least one of the remaining switch units, and each switch unit is also connected to at least one computing unit. The instruction decoder is configured to map the to-be-processed sub-instructions to the switch units corresponding to the positions according to the positions of the to-be-processed sub-instructions in the to-be-processed instruction. The to-be-processed instruction is generated by a compiler compiling all Actions corresponding to the to-be-processed message. By splitting the ALU into computing units of multiple operation types in the embodiment of the present application, multiple computing units can simultaneously process multiple to-be-processed sub-instructions, solving the problem in the prior art that only one ALU is working and the other ALUs are idle, and improving the utilization rate of ALU logic resources and the efficiency of message processing.
[0122] Since the demand for ALU in message processing is a large number of field operations, the calculation needs to be very flexible, but the amount of calculation is not large. The number of ALUs in the traditional solution is generally dozens or hundreds. In the embodiment of the present application, M / N computing link pipelines can be flexibly configured, where M is the total number of switch units in the ALU array 12, and N is the number of switch units required for a single link. The computing logic scale that the original M ALUs can support (assuming that the computing logic required by the service is evenly distributed among the computing units of 4 operation types) only requires at least M computing units to support, and the logic resources can be reduced by 3 / 4 (assuming there are 4 types of computing units in this example). Considering that the demand for computing logic in the service may not be uniform, several types of computing units with different numbers can be allocated according to the service characteristics and requirements.
[0123] The prior art is only applicable to a single and large-number scenario. When encountering complex service requirements, multiple stages are required to complete, resulting in waste of other resources such as look-up table units within the stage. In the embodiment of the present application, the switch units can form a computing pipeline link along any path, adapting to both the scenario of single and large-number computing and the scenario of complex and small-number computing, enabling the computing to be processed within a single stage, and the resources of the pipeline can be used more reasonably, further reducing the latency of processing this service and greatly improving the flexibility of service deployment.
[0124] Based on the processor 10 based on the ALU array shown in the above embodiment, Figure 8 This is a schematic structural diagram of the second embodiment of the processor based on the ALU array provided by the embodiment of the present application. As Figure 8 shown, the processor 10 based on the ALU array may further include:
[0125] The first memory 13 is connected to the ALU array 12 and is used to store the instructions to be processed and the processing results of each sub-instruction to be processed.
[0126] In practical applications, the first memory 13 can be a PHV.
[0127] Based on the processor 10 with an ALU array shown in any of the above embodiments, Figure 9 This is a schematic structural diagram of the third embodiment of the processor with an ALU array provided by the embodiments of the present application. As Figure 9 shown, the processor 10 with an ALU array may further include:
[0128] A second memory 14 is connected to the ALU array 12, and the second memory 14 is used to store the operands required for processing the sub-instructions to be processed.
[0129] In practical applications, the first memory 13 can be an ACTION DATA BUS.
[0130] The switch unit is capable of obtaining operands from a limited number of grouped PHVs and ACTION DATA BUSs, and is capable of writing the processing results back to a PHV.
[0131] Based on Figure 9 the processor 10 with a coarsely reconfigurable ALU array shown, for the outside of the ALU array 12, each switch unit can obtain the operands in the PHV and ACTION DATA BUS, and can also obtain the INSTRUCTION (sub-instruction to be processed) mapped by the instruction decoder. Each switch unit can also output the processing results to the PHV. For the inside of the ALU array 12, each switch unit can obtain the processing results of several adjacent computing units around it and the output results of several adjacent switch units around it; each switch unit can also provide operands and sub-instructions to be processed to several adjacent computing units, and bypass the results to other switch units.
[0132] In practical applications, for the 3D topology under multi-layer circuits, it can be implemented such that the first layer is the ALU array 12 shown in any of the embodiments, the second layer is the interconnection layer, and the third layer is also the ALU array 12 shown in any of the embodiments, and it is stacked in this way.
[0133] In addition to the prior art involved in the background art, there are also the following three prior arts for processing messages.
[0134] Prior Art 1. A message is processed by a coarse-grained reconfigurable array system for deep learning. The system includes: a controller for determining input information to be input to at least one processing unit, where the input information includes weights, input data, status instructions, and operation instructions. The status instructions are used to determine the execution status of the operation instructions, and the operation instructions are used for at least one processing unit to calculate the weights and input data; an input bus for inputting the weights and input data to at least one processing unit; a configuration bus for inputting the status instructions and operation instructions to at least one processing unit; a processing unit group including multiple processing units, and the multiple processing units form a reconfigurable array. Each processing unit is used to calculate the weights and input data according to the operation instructions to obtain result data; an output bus for at least one processing unit to output the result data.
[0135] For Prior Art 1, each Provider Edge (PE) can, through an instruction format, enable the status instructions to control the sequential execution, loop execution, and no-operation of the operation instructions; the operation instructions can control the source of operation data per cycle for the PE, the operation type, writing the result data back to the local register sub-unit, and outputting the result data through the output bus. This means that a single PE needs to fully support all operation types, which is too redundant for the application scenario of this application and will cause a large amount of resource waste. In addition, this application requires a large number of input / output units and the calculation needs to be very flexible, but the amount of calculation is not large. The solution represented by Prior Art 1 will cause a large amount of resource waste on each bus when its input / output bus architecture is expanded to hundreds of times the input / output scale, and its fixed output port causes all data to have to flow through all columns of PEs, which does not meet the application scenario requirements of this application.
[0136] Prior Art 2. A message is processed by a SKINNY-128-128 encryption algorithm system based on a coarse-grained reconfigurable computing unit. The system includes: a reconfigurable configuration system, a reconfigurable data path and computing module, a main control microprocessor, and a system bus; the reconfigurable configuration system includes a configuration information initialization interface, a multi-level configuration information storage unit, a configuration information parsing module, and a location information register; the reconfigurable data path and computing module includes a reconfigurable computing array, a register channel, an intermediate result storage unit, an input first-in-first-out register group, and an output first-in-first-out register group. The reconfigurable computing array includes reconfigurable computing unit blocks, and each reconfigurable computing unit block includes multiple rows of operators, a read control module, and a write control module; the output end of the configuration information register is connected to the reconfigurable data path and computing module; among them, the operators include functions of logical operation, arithmetic operation, shift operation, look-up table operation, and permutation operation.
[0137] For the prior art 2, each operator needs to support multiple types. The operators in odd rows include logical operations, arithmetic operations, shift operations, and permutation operations; the operators in even rows include logical operations, arithmetic operations, shift operations, and look-up table operations; among them, the logical operations include pass-through operation and inversion operation of one operand, exclusive OR operation, AND operation, and OR operation of two operands; the arithmetic operations include addition operation of two operators and addition operation with modulo; the shift operations include arithmetic left shift operation, circular left shift operation, arithmetic right shift operation, and circular right shift operation; the look-up table operations include look-up table operations with up to 4-way parallelism, and the data bit width of the look-up table operation ranges from 4 bits to 32 bits; the permutation operation supports arbitrary permutation of 64-bit data. For the application scenario of this application, the prior art 2 is too redundant, resulting in a large amount of resource waste. In addition, the data flow between operators is very single, and the operators in the same row cannot interact, which does not meet the requirements of the application scenario of this application.
[0138] In the prior art 3, the logic classification unit obtains at least first target classification data and second target classification data from the first data received through the data bus, and sends them to the corresponding first ALU and second ALU according to the preset mapping relationship. Classify the target execution information after the TCAM preprocesses the table entries and service data, so that the instruction memory determines the first information and the second information, and sends them to the corresponding first ALU and second ALU. The first ALU and the second ALU send the data obtained by performing operations based on the first target classification data and the first information, the second target classification data and the second information respectively through the data bus.
[0139] For the prior art 3, as the business complexity increases, the width of the data bus continues to grow, the number of ALUs in the above programmable processing architecture increases, and the full interconnection between the data and the ALUs also increases, resulting in a sharp increase in the required number of connections between the ALU and the data bus, and the physical implementation cost is huge. Since the connection method of the architecture shown in the prior art 3 is relatively fixed, new requirements are imposed on the compiler. The compiler may allocate the same packet header field to multiple PHVs. When this packet header field needs to be modified in some cases, multiple PHVs need to be modified simultaneously, so multiple ALUs are required to modify the corresponding PHVs, resulting in resource waste.
[0140] The ALU array 12 based on coarse-grained reconfiguration provided by the embodiments of this application can be applied to a programmable chip, which can classify the data that needs to be processed by the ALU and send it to the corresponding computing unit for operation processing, without interconnecting with all data, which can save connection resources, reduce the connection or connection with the data bus, and thus reduce the physical implementation cost. At the same time, the embodiments of this application can share the processing result of a computing unit through the switching unit, thereby realizing the saving of ALU resources.
[0141] In summary, for the processor based on the ALU array provided in the embodiments of the present application, through the coarsely grained reconfigurable ALU array 12, the idle logic resources are reduced under the same computing power, and the resource utilization rate is improved. Through the serial computing logic, complex operations that originally required multiple stages of deployment can be completed in one stage, saving the deployment resources of the logic table. Indirectly, more logic tables can be deployed in the entire architecture to support more complex business logics, providing the ability to process complex composite instructions in a single stage and effectively solving the problems of resource waste and inflexible deployment existing in the existing architecture. Moreover, the processor based on the ALU array provided in the present application is flexibly programmable and can implement serial and parallel computing logics.
[0142] Based on the processor based on the ALU array shown in any of the above embodiments, the process of implementing the message processing method based on this processor will be explained below.
[0143] Figure 10 It is a schematic flowchart of the first embodiment of the message processing method provided in the embodiments of the present application. As Figure 10 shown, this message processing method is applied to the processor 10 based on the ALU array shown in any of the above embodiments, and this method may include the following steps:
[0144] S101. The compiler compiles all Actions corresponding to the message to be processed to generate instructions to be processed.
[0145] Among them, the instructions to be processed include multiple sub-instructions to be processed.
[0146] In a possible implementation manner, the compiler compiles all Actions corresponding to the message to be processed to obtain the sub-instructions to be processed corresponding to each Action. For each sub-instruction to be processed, according to the operation type of the computing unit connected to each switch unit, the switch unit capable of processing the sub-instruction to be processed is determined. According to the mapping relationship and the switch unit corresponding to each sub-instruction to be processed, the position of the sub-instruction to be processed in the instruction to be processed is determined, and the mapping relationship is used to illustrate the correspondence between the position in the instruction to be processed and the switch unit. According to the position of each sub-instruction to be processed in the instruction to be processed, the instruction to be processed is generated.
[0147] Specifically, after obtaining the message to be processed, the compiler matches the key fields of the message to be processed with the entries in the matching table to determine the Action to be performed on the message header fields of the message to be processed. Subsequently, the compiler compiles each Action to obtain the sub-instructions to be processed corresponding to each Action. Further, the compiler determines the switch unit that can process the sub-instructions to be processed based on the computing units connected to each switch unit. Exemplarily, assuming that the sub-instruction to be processed 1 requires arithmetic operations, the switch unit connected to the computing unit with the arithmetic operation type is determined as the switch unit for processing the sub-instruction to be processed 1. During this process, it is necessary to ensure that the computing unit is not simultaneously called by different switch units. Finally, based on the switch unit corresponding to each sub-instruction to be processed, the position of each sub-instruction to be processed in the instruction to be processed is determined, thereby generating the instruction to be processed. There is a mapping relationship between each position in the sub-instruction to be processed and the switch unit, and this mapping relationship is pre-configured.
[0148] S102. The instruction decoder maps the sub-instructions to be processed to the switch units corresponding to their positions according to the positions of the sub-instructions to be processed in the instruction to be processed.
[0149] After obtaining the instruction to be processed, the instruction decoder analyzes the sub-instructions to be processed included in the instruction to be processed and the positions of each sub-instruction to be processed in the instruction to be processed. Subsequently, based on the mapping relationship between the positions in the instruction to be processed and the switch units, the instruction decoder maps the sub-instructions to be processed to the switch units corresponding to their positions according to the positions of the sub-instructions to be processed in the instruction to be processed.
[0150] S103. For each switch unit, the sub-instructions to be processed obtained are processed by the target computing unit connected to the switch unit to obtain the processing results of the sub-instructions to be processed.
[0151] For each switch unit, after obtaining the sub-instructions to be processed mapped by the instruction decoder, the sub-instructions to be processed are sent to the target computing unit corresponding to the sub-instructions to be processed, and the operation type of the target computing unit is the operation type required by the sub-instructions to be processed. The target computing unit then calculates the sub-instructions to be processed to obtain the processing results, and the switch unit obtains the processing results generated by its calculation from the target computing unit.
[0152] It should be understood that the specific implementation process and principle of S103 will be elaborated in the Figure 10 embodiments shown, and will not be limited here.
[0153] The message processing method provided by the embodiment of the present application compiles all Actions corresponding to the message to be processed by a compiler to generate instructions to be processed, and the instructions to be processed include multiple sub-instructions to be processed. An instruction decoder maps each sub-instruction to be processed to a corresponding switch unit according to the position of the sub-instruction to be processed in the instruction to be processed. For each switch unit, the target calculation unit connected to the switch unit processes the obtained sub-instruction to be processed to obtain the processing result of the sub-instruction to be processed, and the operation type of the target calculation unit is the operation type required by the sub-instruction to be processed. In this technical solution, the compiler compiles all Actions corresponding to the message to be processed to generate multiple sub-instructions to be processed, and the instruction decoder maps each sub-instruction to be processed to a corresponding switch unit, so that the switch unit processes the sub-instruction to be processed through the corresponding calculation unit, thereby achieving the purpose of synchronously processing multiple sub-instructions to be processed, solving the problem that only one ALU is working and other ALUs are idle in the prior art, and improving the utilization rate of ALU logic resources and the efficiency of message processing.
[0154] Based on the above embodiment, the implementation process and principle of S103 will be explained below.
[0155] Figure 11 This is a schematic flowchart of the first embodiment of the message processing method provided by the embodiment of the present application. As Figure 11 shown, S103 may include the following steps:
[0156] S111. For each switch unit, obtain the target data that the sub-instruction to be processed needs to process.
[0157] In a possible implementation manner, the switch unit obtains the operands that the sub-instruction to be processed needs to process from the first memory and / or the second memory.
[0158] In another possible implementation manner, if the sub-instruction to be processed indicates to process the processing result of the third switch unit again, the switch unit obtains the processing result of the third switch unit from the third switch unit.
[0159] That is to say, the processing sub-instruction and the sub-instruction to be processed by the third switch unit are used to perform multi-level processing on the same message header field.
[0160] S112. The switch unit sends the sub-instruction to be processed and the target data to the corresponding target calculation unit.
[0161] S113. The target calculation unit processes the target data according to the sub-instruction to be processed to generate the processing result of the sub-instruction to be processed.
[0162] S114. The switch unit obtains the processing result of the sub-instruction to be processed from the target computing unit.
[0163] Further, the switch unit may send the processing result of the sub-instruction to be processed to the first memory; the switch unit may also send the processing result to the first switch unit, and the first switch unit is used to reprocess the processing result of the switch unit, that is, the first switch unit and the switch unit are used to process the same message header field; the switch unit may also send the processing result to the second switch unit, and the second switch unit is used to send the processing result of the switch unit to the first memory.
[0164] To facilitate the explanation of this solution, the following two cases of processing the message header field once and processing the message header field multiple times are respectively illustrated by examples.
[0165] Case 1: Processing the message header field once
[0166] For message header field 1, the compiler generates sub-instruction 1 to be processed, and sends the sub-instruction 1 to switch unit 1 that can process the sub-instruction 1 through the instruction encoder. Switch unit 1 obtains the operands from the PHV and / or ACTION DATA BUS according to the sub-instruction 1 to be processed, sends the operands and the sub-instruction 1 to the connected computing node, and obtains the processing result from the computing node. Further, switch unit 1 sends the processing result to be stored in the PHV.
[0167] Case 2: Processing the message header field multiple times
[0168] For the convenience of description, this application takes processing the message header field twice as an example for explanation. For message header field 2, the compiler generates two sub-instructions to be processed, namely sub-instruction 2 and sub-instruction 3, and both sub-instruction 2 and sub-instruction 3 are used to process message header field 2. The compiler sends the sub-instruction 2 to switch unit 2 that can process the sub-instruction 2 through the instruction encoder, and sends the sub-instruction 3 to switch unit 3 that can process the sub-instruction 3. Switch unit 2 obtains the operands from the PHV and / or ACTION DATABUS according to the sub-instruction to be processed, sends the operands and the sub-instruction 2 to the connected computing node, and obtains the processing result from the computing node.
[0169] Further, switch unit 2 sends the processing result to switch unit 3, and switch unit 3 sends the processing result of switch unit 2 and the sub-instruction 3 to the connected computing node, and obtains the processing result from the computing node. This processing result is the final result of the secondary calculation processing of the message header field.
[0170] An embodiment of the present application further provides an electronic device, which may include: a processor, a memory, and computer program instructions stored in the memory and executable on the processor. The processor is the processor 10 based on the ALU array described in any of the foregoing embodiments, and is configured to execute the message processing method provided in any of the foregoing embodiments.
[0171] Optionally, the various components of the electronic device may be connected via a system bus.
[0172] The memory may be a separate storage unit or an integrated storage unit in the processor. The number of processors is one or more.
[0173] Optionally, the electronic device may further include a communication interface for interacting with other devices.
[0174] The system bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.
[0175] All or part of the steps of implementing the foregoing method embodiments may be completed by hardware related to program instructions. The foregoing program may be stored in a readable memory. When the program is executed, it executes the steps including the foregoing method embodiments; and the foregoing memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof.
[0176] The electronic device provided by the embodiment of the present application can be used to execute the message processing method provided in any of the foregoing method embodiments. The implementation principle and technical effects are similar and will not be elaborated here.
[0177] An embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, which, when running on a computer, cause the computer to execute the above-mentioned message processing method.
[0178] The above-mentioned computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium may be any available medium accessible by a general-purpose or special-purpose computer.
[0179] Optionally, the readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium may also be part of the processor. The processor and the readable storage medium may be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium may also exist as discrete components in a device.
[0180] An embodiment of the present application further provides a computer program product including a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when the at least one processor executes the computer program, the above-mentioned message processing method can be implemented.
[0181] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. An ALU array, characterized in that, including: a plurality of switch units and calculation units of multiple operation types; each switch unit is connected to at least one of the remaining switch units, and each switch unit is also connected to calculation units of at least one operation type; each switch unit is used to receive the to-be-processed sub-instructions mapped by the instruction decoder, and send the to-be-processed sub-instructions to the calculation unit that is connected to the switch unit and corresponds to the operation type required by the to-be-processed sub-instructions.
2. The arithmetic logic unit ALU array according to claim 1, characterized in that, The operation types include at least one of the following: arithmetic operation, bit operation, comparison operation, and logical operation.
3. The ALU array according to claim 1 or 2, characterized in that, Each switch unit is connected to the remaining switch units according to a topological structure, and the topological structure includes any one of the following: linear topology, triangular topology, or polygonal topology.
4. The ALU array according to claim 3, characterized in that, There is at least one target calculation unit among the calculation units included in the ALU array, and each target calculation unit is connected to at least two switch units.
5. The ALU array according to claim 2, characterized in that, Each switch unit is connected to calculation units of four different operation types.
6. A processor based on an ALU array, characterized in that, including: an instruction decoder and an arithmetic logic unit ALU array as described in any one of claims 1 to 5; the instruction decoder is connected to the ALU array; the instruction decoder is used to map the to-be-processed sub-instructions to the switch unit corresponding to the position according to the position of each to-be-processed sub-instruction in the to-be-processed instruction, and the to-be-processed instruction is generated by the compiler compiling all actions Action corresponding to the to-be-processed message.
7. The processor according to claim 6, wherein, The processor further includes: a first memory, the first memory is connected to the ALU array, and is used to store the to-be-processed instruction and the processing result of each to-be-processed sub-instruction.
8. The processor according to claim 6 or 7, characterized in that, The processor further includes: a second memory, the second memory is connected to the ALU array, and the second memory is used to store the operands that the to-be-processed sub-instructions need to process.
9. A message processing method, characterized in that, Applied to the processor based on the ALU array as described in any one of claims 6 to 8, the method includes: the compiler compiles all actions Action corresponding to the to-be-processed message to generate a to-be-processed instruction, and the to-be-processed instruction includes a plurality of to-be-processed sub-instructions; the instruction decoder maps the to-be-processed sub-instructions to the switch unit corresponding to the position according to the position of each to-be-processed sub-instruction in the to-be-processed instruction; for each switch unit, the to-be-processed sub-instructions obtained are processed by the target calculation unit connected to the switch unit to obtain the processing result of the to-be-processed sub-instructions, and the operation type of the target calculation unit is the operation type required by the to-be-processed sub-instructions.
10. The method according to claim 9, wherein The step of, for each switch unit, processing the to-be-processed sub-instructions obtained by the target calculation unit connected to the switch unit to obtain the processing result of the to-be-processed sub-instructions includes: for each switch unit, obtaining the target data that the to-be-processed sub-instructions need to process; the switch unit sends the to-be-processed sub-instructions and the target data to the corresponding target calculation unit; the target calculation unit processes the target data according to the to-be-processed sub-instructions to generate the processing result of the to-be-processed sub-instructions. The switch unit obtains the processing result of the to-be-processed sub-instruction from the target computing unit.
11. The method according to claim 10, wherein The method further includes: The switch unit sends the processing result of the to-be-processed sub-instruction to the first memory; Or, The switch unit sends the processing result of the to-be-processed sub-instruction to the first switch unit, and the first switch unit is configured to re-process the processing result of the switch unit; Or, The switch unit sends the processing result of the to-be-processed sub-instruction to the second switch unit, and the second switch unit is configured to send the processing result of the switch unit to the first memory.
12. The method according to claim 11, characterized in that, For each switch unit, obtaining the target data that the to-be-processed sub-instruction needs to process includes: The switch unit obtains the operands that the to-be-processed sub-instruction needs to process from the first memory and / or the second memory; Or, If the to-be-processed sub-instruction indicates to re-process the processing result of the third switch unit, the switch unit obtains the processing result of the third switch unit from the third switch unit; The target data is the to-be-processed sub-data or the processing result of the third switch unit.
13. The method according to any one of claims 9 to 12, characterized in that, The compiler compiles all Actions corresponding to the to-be-processed message to generate to-be-processed instructions, including: The compiler compiles all Actions corresponding to the to-be-processed message to obtain the to-be-processed sub-instructions corresponding to each Action; For each to-be-processed sub-instruction, the compiler determines the switch unit capable of processing the processing sub-instruction according to the operation type of the computing unit connected to each switch unit; The compiler determines the position of the to-be-processed sub-instruction in the to-be-processed instruction according to the mapping relationship and the switch unit corresponding to each to-be-processed sub-instruction, and the mapping relationship is used to illustrate the correspondence between the position in the to-be-processed instruction and the switch unit; The compiler generates the to-be-processed instruction according to the position of each to-be-processed sub-instruction in the to-be-processed instruction.
Citation Information
Cited By
Arithmetic logic unit parallel processing system and method and electronic equipment
CN120832124A