Packet network capable of inter-transpose
The packet network addresses the challenges of data transmission in deep learning by using a fetch unit with a network interface and routers to manage data flow and ensure sequential transmission, achieving efficient data transposition and improved network performance.
Patent Information
- Application Number
- PCT/KR2024/016950
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-10-31
- Publication Date
- 2025-06-05
AI Technical Summary
Existing packet networks face challenges in efficiently transmitting and processing data packets in deep learning applications, particularly due to increased node counts leading to larger packet headers and potential queue blocking issues.
The proposed packet network employs a fetch unit with a network interface, fetch network, and fetch network controller to manage data transmission and processing. Each fetch unit includes a fetch buffer, routers with data processing mapping tables, and an arbiter unit to control data flow and ensure sequential transmission.
This solution enables efficient data transposition during the fetching process without additional processing, reducing the need for separate computational processes or data re-storage, thereby enhancing network throughput and reducing costs.
Smart Images

Figure KR2024016950_05062025_PF_FP_ABST
Abstract
Description
Inter-transposable packet network
[0001] The present invention relates to a neural network processor, and more particularly, to a network for transmitting packets.
[0002] This application claims priority to Korean Patent Application No. 10-2023-0169300, filed in the Republic of Korea on November 29, 2023, the entire disclosure of which is incorporated herein by reference.
[0003] An artificial neural network (ANN) is an artificial intelligence that implements artificial intelligence by connecting artificial neurons that are mathematically modeled after the neurons that make up the human brain. One mathematical model for an artificial neuron is as shown in Equation (1) below. Specifically, an artificial neuron receives an input signal x i Enter x i The weight w corresponding to each i Multiply by to get the total. Next, the artificial neuron uses the activation function to get the activation value, and passes this activation value to the next connected artificial neuron.
[0004] y = f(w1* x1+ w2* x2+ .... w n * x n ) = f(Σ w i * x i ), where i = 1......n, n = # input signal - Formula (1)
[0005] A deep neural network (DNN), a type of ANN, has a layered network architecture of artificial neurons (nodes). A DNN consists of an input layer, an output layer, and multiple hidden layers between the input and output layers. The input layer consists of multiple nodes that receive input values, and the nodes in the input layer transmit output values calculated through the mathematical model described above to the nodes in the next hidden layer connected to the input layer. The nodes in the hidden layer receive input values through the mathematical model described above, calculate output values, and transmit the output values to the nodes in the output layer.
[0006] The computational process of deep learning, a form of machine learning performed in DNN, can be categorized into a training process in which a given DNN continuously learns based on learning data to improve the computational ability of the DNN, and a process in which the DNN learned through the training process is used to make inferences about new input data.
[0007] Deep learning's inference process is achieved through forward propagation, where input data is received by nodes in the input layer, and computations are sequentially performed in the hidden and output layers in the order of the layers. Ultimately, the output layer nodes derive the conclusion of the inference process based on the output values of the hidden layers.
[0008] In contrast, deep learning's training process adjusts the weights of nodes to reduce the gap between the conclusion of the inference process and the actual correct answer. Typically, weights are modified using gradient descent. To implement gradient descent, the differential value of the difference between the conclusion of the inference process and the actual correct answer must be calculated for each weight of each node. In this process, the differential value of the weight of the preceding node of the DNN is calculated using the chain rule of the differential value of the weight of the subsequent node of the DNN. Since the chain rule is calculated in the reverse direction of the inference process, the deep learning training process uses a backpropagation method.
[0009] In other words, DNN has a hierarchical structure, and the nodes in each layer receive result values from a number of nodes in the previous layer, perform operations based on the mathematical model of the above-mentioned nodes, output new result values, and pass the new result values to the nodes in the next layer.
[0010] Meanwhile, the computational structure of a DNN can adopt a distributed processing structure that distributes the numerous operations performed by nodes in each layer across multiple computational units. The computations performed by nodes in each layer are distributed across multiple computational units, each of which retrieves the data required for the operation from memory, performs the operation, and stores the results back in memory.
[0011] The purpose of this specification is to provide a packet network capable of inter-transposition.
[0012] This specification is not limited to the above-mentioned tasks, and other tasks not mentioned will be clearly understood by those skilled in the art from the description below.
[0013] In order to solve the above-described problem, an operation processing device according to the present specification is provided. In an operation processing device including a plurality of fetch units that read data required for an operation for performing processing of a neural network from a memory and provide the data to an operation unit, each fetch unit includes: a network interface having a fetch buffer from which data stored in the memory is fetched; a fetch network having a plurality of routers that guarantee sequential transmission and have a data processing mapping table describing a method of processing the input data according to a node ID of the data fetched to the network interface; and a fetch network controller that reconfigures each of the data processing mapping tables of the plurality of routers to form a software topology according to an operation type; wherein the memory includes the same number of data memory slices as the plurality of routers, and the network interface includes:
[0014] It may include an interface controller that assigns a destination node ID to data fetched into each of the above data memory slices.
[0015] According to one embodiment of the present specification, the network interface can assign a number of destination node IDs equal to the number of routers belonging to the same group by the fetching network controller to the data fetched to each data memory slice.
[0016] According to one embodiment of the present specification, the network interface can sequentially and repeatedly assign a destination node ID to data fetched to each data memory slice in the order in which they were fetched.
[0017] According to one embodiment of the present specification, each router may include a main input port for inputting data from the memory, a first forwarding output port for transmitting data to an adjacent first router, a first forwarding input port for inputting data transmitted from the adjacent first router, a second forwarding output port for transmitting data to an adjacent second router, a second forwarding input port for inputting data transmitted from the adjacent second router, and a main output port for providing data to the computational unit.
[0018] According to one embodiment of the present specification, the data processing mapping table can store information on whether blocking, reflection, and output of input data are executed.
[0019] According to one embodiment of the present specification, the fetch network controller can set whether to execute the blocking and reflection depending on the topology to be reconfigured.
[0020] According to one embodiment of the present specification, the fetch network controller can set the blocking in the data processing mapping table of routers belonging to the same group in the reconfigured topology to be identical and output only one node ID.
[0021] The processing device according to the present specification may further include an arbiter unit for controlling an input port according to tail information of input data and changing the tail information according to tail maintenance information of the input data, for each router.
[0022] According to one embodiment of the present specification, the tail information may be either a first logic value (T) meaning the end of a packet or a second logic value (F) meaning not the end of a packet, and the tail maintenance information may be either a first logic value (T) meaning to maintain the tail information or a second logic value (F) meaning to change the tail information.
[0023] According to one embodiment of the present specification, when controlling an input port to input data from the memory, the arbiter unit can change the input port to input data from an adjacent router if the tail information of the data input from the memory is a first logic value (T).
[0024] According to one embodiment of the present specification, the arbiter unit can change the tail information of the input data to the second logic value (F) when the tail maintenance information of the data input from the memory is the second logic value (F).
[0025] According to one embodiment of the present specification, the arbiter unit can control data to be input from the memory when tail information of data input from the adjacent router is a first logic value (T).
[0026] According to one embodiment of the present specification, the network interface may further set the tail information and the tail maintenance information to the fetched data.
[0027] According to one embodiment of the present specification, the network interface may set tail information of the last data of a packet among the fetched data to a first logical value (T), and set tail information of the remaining data to a second logical value (F).
[0028] According to one embodiment of the present specification, the network interface may set tail maintenance information of data having the last node ID among the software topology configurations of the router to a first logical value (T) meaning to maintain the tail information, and may set tail maintenance information of data having the remaining node IDs to a second logical value (F) meaning to change the tail information.
[0029] Other specific details of the present invention are included in the detailed description and drawings.
[0030] According to this specification, transposing data is possible during the data fetching process without additional processing such as a separate computational process or data re-storage.
[0031] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0032] FIG. 1 is a block diagram schematically illustrating the configuration of an operation processing device according to one embodiment of the present invention.
[0033] Figure 2 is a block diagram illustrating each component of the operation processing device according to this specification in more detail.
[0034] FIG. 3 is a block diagram illustrating in more detail the configuration of a fetch unit according to one embodiment of the present specification.
[0035] Figure 4 is a reference diagram for explaining the configuration of a router according to this specification.
[0036] Figure 5 is a software topology according to the first embodiment.
[0037] Figure 6 is a reference diagram of a data processing mapping table according to the first embodiment.
[0038] Figure 7 is a software topology according to the second embodiment.
[0039] Figure 8 is a reference diagram of a data processing mapping table according to the second embodiment.
[0040] Figure 9 is a software topology according to the third embodiment.
[0041] Figure 10 is a reference diagram of a data processing mapping table according to the third embodiment.
[0042] Figure 11 is a software topology according to the fourth embodiment.
[0043] Figure 12 is a reference diagram of a data processing mapping table according to the fourth embodiment.
[0044] FIG. 13 is an example of data input timing according to one embodiment of the present specification.
[0045] Figures 14 to 20 are reference diagrams for processing data by the arbiter unit.
[0046] FIG. 21 is a reference diagram of inter-transpose according to one embodiment of the present specification.
[0047] FIG. 22 is an example of a data processing mapping table for inter-transposition according to one embodiment of the present specification.
[0048] Figure 23 is a reference diagram of the results of inter-transposition according to one embodiment of the present specification.
[0049] The advantages and features of the invention disclosed in this specification, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, this specification is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided only to ensure that the disclosure of this specification is complete and to fully inform those of ordinary skill in the art (hereinafter referred to as "skilled workers") of the scope of this specification, and the scope of rights of this specification is defined only by the scope of the claims.
[0050] The terminology used herein is for the purpose of describing embodiments and is not intended to limit the scope of the present disclosure. In this specification, singular forms also include plural forms, unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components in addition to the components mentioned.
[0051] Throughout the specification, the same reference numerals refer to the same elements, and the term "and / or" includes each and every combination of the elements mentioned. Although terms such as "first," "second," etc. are used to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another. Therefore, it should be understood that a first element mentioned below may also be a second element within the technical scope of the present invention.
[0052] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those skilled in the art to which this specification pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise. Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0053] The data used in deep learning training can take the form of tensors, ranging in size from hundreds of kilobytes to hundreds of megabytes. This data can be stored in multiple memory banks that make up on-chip memory.
[0054] Multiple memory banks and multiple computational units are connected to a network for data transmission. In the case of a network-on-chip, the network can be configured within the chip and may include a router. The router includes a router that can forward data packets received from multiple nodes to multiple nodes. The router can perform at least one of the following actions: i) ensuring that data packets (i.e., traffic) received from various directions are forwarded toward their destination, ii) performing arbitration in the event of contention, and iii) performing flow control to prevent packet loss. The performance and cost of such a router are determined by topology, bandwidth, flow control, buffers, etc., and naturally, routers with high throughput at low cost (cost, area, and energy) are pursued.
[0055] Meanwhile, in deep learning, most traffic patterns involve reusing the same tensor data multiple times to generate multiple output tensor data. Therefore, to reduce the number of memory accesses, the router can read the input tensor data from memory and broadcast or multicast it to multiple computational units. A common method of multicasting is to record the destination for each data (dg, data packet) and then transmit it. However, this method has the problem that the packet header size increases proportionally with the number of nodes (e.g., if the packet header includes a bitmap indicating the destination, at least 64 bits are required for 64 nodes). Buffered flow control typically causes head-of-line blocking, which occurs depending on the buffer size. One approach to address this issue is source throttling, which detects and avoids congestion after it has occurred. Therefore, a network capable of high throughput at low cost is needed, taking into account the routing patterns inherent in deep learning.
[0056] FIG. 1 is a block diagram schematically illustrating the configuration of an operation processing device according to one embodiment of the present invention.
[0057] As illustrated in FIG. 1, the operation processing unit (10) may include a memory (100), a fetch unit (200), an operation unit (300), and a commit unit (400). However, as illustrated in FIG. 1, the operation processing unit (10) does not have to include all of the memory (100), the fetch unit (200), the operation unit (300), and the commit unit (400). For example, the memory (100) and the commit unit (400) may be arranged outside the operation processing unit (10).
[0058] The memory (100) can store at least one of the data described in this specification. For example, the memory (100) can store input data, tensors, output data, filters, operation result data of the operation unit, all data used in the fetch unit, etc. The memory (100) can be formed in the form of a data memory such as SRAM (Static Random Access Memory) or DRAM (Dynamic Random Access Memory), but this is not necessarily the case.
[0059] The fetch unit (200) can read data required for operation from input data stored in the memory (100) and provide it to the operation unit (300). If the input data is a tensor, the fetch unit (200) can read the tensor stored in the memory (100) and feed it to the operation unit (300) according to the form of the operation. The form of the operation may include, for example, a matrix multiplication operation, convolution, or grouped convolution. In this case, the fetch unit (200) can sequentially read a data group having data equal to or greater than the unit data processing capacity of one or more operators provided in the operation unit (300) from the memory (100) and feed it to the operation unit (300).
[0060] The operation unit (300) can process the input data received from the fetch unit (200) to generate output data. The operation unit (300) can be configured to match (correspond to) the type of operation to be performed. For example, the operation unit (300) can process data fed from the fetch unit (200) in a streaming manner, but is not limited thereto. The operation unit (300) can include one or more calculators.
[0061] The commit unit (400) can store the operation result data output (e.g., output in a streaming manner) from the operation unit (300) in the memory (100). When the commit unit (400) performs an operation of storing the operation result data received from the operation unit (300) in the memory (100), the commit unit (400) can store the operation result data in the memory (100) based on the form of the operation to be performed next. For example, the commit unit (400) can transform the operation result data into a preset form or a form suitable for future operation and store it in the memory (100).
[0062] Figure 2 is a block diagram illustrating each component of the operation processing device according to this specification in more detail.
[0063] Referring to FIG. 2, the memory (100), fetch unit (200), operation unit (300), and commit unit (400) described above will be described in more detail.
[0064] The memory (100) may be configured based on a memory address space. For example, the memory address space may be continuous or sequential. In addition, the memory address space may be one-dimensional, but is not limited thereto, and may have a two-dimensional or more array. The internal structure of the memory (100) may be configured with an independently accessible slice structure. For example, the memory (100) may include a plurality of data memory slices (101). In this case, the number of data memory slices (101) may be determined according to the number of dot product engines (310) included in the operation unit (300). For example, the number of slices (101) may be equal to the number of dot product engines (310) included in the operation unit (300). For example, when the input data is a tensor, the tensor may be divided in the channel direction and the height direction, and then stored in the data memory slice (101).
[0065] The fetch unit (200) can read data from the memory (100) and feed the data to the dot product engine (310) of the operation unit (300). For example, the fetch unit (200) can include at least one of a fetch sequencer (210), a network interface (220), a fetch network (230), a feed module (240), and an operation sequencer module (250). The fetch sequencer (210) can control a data fetch operation from the memory (100) to the network interface (220). The network interface (220) is for fetching data stored in the memory (100) and can provide an interface between the memory (100) and the fetch network (230). The fetch network (230) can transmit the fetched data to the feed module (240). The operation sequencer module (250) can control the operation unit (300) to perform a specific operation by controlling the feed module (240) and data input to the feed module (240).
[0066] The fetch network (230) of the fetch unit (200) may have various structures depending on the operation content and the form of data. The fetch network (230) may be configured or reconfigured by software into a topology of a form required by the operation unit (300). In addition, the fetch network (230) may determine the topology depending on the form of input data and the form of the operation. The fetch network (230) may support various communication forms such as Direct, Vertical Multicast, Channel Multicast, and Vertical Nearest Neighbor depending on the operation performed in the operation unit (300), but is not limited thereto.
[0067] For example, in the case of two-dimensional convolution, it is assumed that the values of all input channels must be input to the dot product engine (310) that calculates each output activation. Therefore, the fetch unit (200) can feed the input activation values read sequentially in the channel direction to the dot product engine (310) in a multicast manner. In addition, the fetch unit (200) can use the fetch sequencer (210) to sequentially read data to be input to the operation unit (300) from each data memory slice (101). Each piece of data read from the data memory slice (101) by the fetch sequencer (210) can be transmitted to the operation unit (300) through the fetch network (230) of the fetch unit (200).
[0068] In this way, the fetch unit (200) can read tensor slices from the memory (100) in parallel and feed them to the operation unit (300) in a form that the operation unit (300) can operate on. Here, the fetch network (230) may further include a fetch network controller (not shown in FIG. 2) that configures and manages the fetch network (230) to transmit data read from the memory (100) to the operation unit (300) that requires it.
[0069] The computation unit (300) may include a plurality of dot product engines (310) capable of processing in parallel. For example, the computation unit (300) may include 256 dot product engines (310), but is not limited thereto. Each dot product engine (310) may include one or more calculation units (e.g., 32 MACs). Each dot product engine (310) may perform various calculations depending on the configuration of the calculation units. The dot product engines (310) of the computation unit (300) may also be divided in the channel and height directions to perform calculations to generate output activation.
[0070] The operation unit (300) may further include a register file (not shown) in addition to the dot product engine (310).
[0071] A register file is a storage space that temporarily stores one of the operators that is relatively frequently used or reused when the dot product engine (310) performs calculations. For example, the register file may be configured in the form of SRAM or DRAM, but is not limited thereto.
[0072] For example, when performing computations in a neural network, for a typical convolutional layer with large activations, the weights may be stored in a register file, while the activations may be stored in memory. Furthermore, for a fully connected layer, where the weights are larger than the activations, the weights may be stored in memory, while the activations may be stored in a register file.
[0073] For example, when the operation unit (300) performs a MAC operation, the dot product engine (310) may use input data received from the fetch unit (200) as operands for performing the MAC operation, a register value received from a register file located in the dot product engine (310), and an accumulation value received from an accumulator. Then, the operation result may be stored in the accumulator again or transmitted to the commit unit (400) to be stored in the memory (100) as output data.
[0074] Meanwhile, as described above, the commit unit (400) can transform the output activation calculated by the operation unit (300) into a form required for the next operation and store it in the memory (100).
[0075] For example, in a neural network, the commit unit (400) can store the output activation in memory so that the output activation according to the operation in a specific hierarchical layer can be used for the operation in the next layer. In addition, depending on the data type required for the operation in the next layer, the commit unit (400) can perform a transpose (e.g., tensor manipulation) and transfer the results to the memory (100) through a commit network (not shown) for storage.
[0076] In this way, the commit unit (400) stores the output data in a desired form in the memory (100) after the operation is performed by the operation unit (300). In order to store the output data in a desired form, the commit unit (400) can perform data transpose (Tensor Transpose) using a data transpose module (not shown) and a commit network module (not shown).
[0077] FIG. 3 is a block diagram illustrating in more detail the configuration of a fetch unit according to one embodiment of the present specification.
[0078] Referring to FIG. 3, a network interface (220), a fetch network (230), and a feed module (240) included in a fetch unit (200) according to one embodiment of the present specification can be identified.
[0079] Data stored in each of the above data memory slices (101) can be fetched to a network interface (220). The network interface (220) can include a fetch buffer (222) that stores the fetched data and an interface controller (221) that assigns a node ID corresponding to each data memory slice to the fetched data.
[0080] The fetch network (230) can transmit data fetched to the fetch buffer (222). To this end, the fetch network (230) can include a plurality of routers (232) and a fetch network controller (231).
[0081] Each of the plurality of routers (232) may have a data processing mapping table. The data processing mapping table may indicate a routing / flow control method (e.g., blocking, reflection, or output) of the input data according to the node ID of the input data. The fetch network controller (231) may reconfigure the data processing mapping table. The data processing mapping table may be adaptively reconfigured to the type of operation to be performed on the corresponding data in the future. In this example, the fetch network controller (231) may reconfigure the data processing mapping table of each of the plurality of routers (232) to form a software topology according to the type of operation. The data processing mapping table will be described in more detail later.
[0082] The feed module (240) can provide data received from the fetch network (230) to the operation unit (300). To this end, the feed module (240) can include a feed buffer (242) that stores data output from a plurality of routers (232).
[0083] Meanwhile, the memory (100) may include one or more data memory slices (101). The number of routers (232) may be related to the number of data memory slices (101). For example, the number of routers (232) may be determined based on the number of data memory slices (101), or conversely, the number of data memory slices (101) may be determined based on the number of routers (232). For example, the number of data memory slices (101) may be the same as the number of routers (232). In this case, the routers (232) and the data memory slices (101) may correspond one to one. In this specification, for the convenience of understanding and simplification of the drawing, it will be assumed that there are eight routers (232-1 to 232-8). And since the data stored in the data memory slice (101) can be fetched to the fetch buffer (222) included in the network interface (220), in FIG. 3, eight buffers (222-1 to 222-8) are illustrated as being separated from each other. Accordingly, each fetch buffer (222) corresponds to each data memory slice (101), and data stored in each corresponding data memory slice (101) can be fetched to each fetch buffer (222). However, the example in FIG. 3 is illustrated as being separated from each other for the convenience of understanding, and is not necessarily limited to being physically separated.
[0084] According to one embodiment of the present specification, the plurality of routers (232) can form a hardware 1-dimensional mesh topology. According to another embodiment of the present specification, the plurality of routers (232) can also form a hardware 2-dimensional mesh topology (2D mesh topology). The hardware configuration of the plurality of routers (232) can vary and is not limited to a specific form. Each router (232) can receive data fetched to the fetch buffer (222) and output it to the feed buffer (242), or transmit it to another adjacent router (232). For the convenience of the following description, the routers located at the far left among the plurality of routers will be named 'Router 1 (232-1), Router 2 (232-2), ......, Router 8 (232-8)'.
[0085] Figure 4 is a reference diagram for explaining the configuration of a router according to this specification.
[0086] Referring to Fig. 4, three routers can be identified. Among these three routers, the router (232-Ref) located in the middle will be used as a reference for the configuration of the router according to this specification. Among the two routers adjacent to the reference router (232-Ref), the router located on the left will be named 'the first router (232-F)', and the router located on the right will be named 'the second router (232-S)'. The 'first' and 'second' are merely names to distinguish the two routers adjacent to the reference router (232-Ref) and do not indicate priorities between the routers.
[0087] A router (232) according to the present specification may include a main input port (①), a first forwarding output port (②), a first forwarding input port (③), a second forwarding output port (④), a second forwarding input port (⑤), and a main output port (⑥). The main input port (①) is a port through which data is input from the memory (101), i.e., the fetch buffer (222). The first forwarding output port (②) is a port through which data is transmitted to an adjacent first router (232-F). The first forwarding input port (③) is a port through which data transmitted from an adjacent first router (232-F) is input. The second forwarding output port (④) is a port through which data is transmitted to an adjacent second router (232-S). The second forwarding input port (⑤) is a port through which data transmitted from an adjacent second router (232-S) is input. The main output port (⑥) is a port for providing data to the operation unit (300), i.e., the feed buffer (242).
[0088] Accordingly, data output through the first forwarding output port (②) of the reference router (232-Ref) is input to the second forwarding input port (⑤) of the first router (232-F). Data output through the second forwarding output port (④) of the first router (232-F) is input to the first forwarding input port (③) of the reference router (232-Ref). Data output through the second forwarding output port (④) of the reference router (232-Ref) is input to the first forwarding input port (③) of the second router (232-S). Data output through the first forwarding output port (②) of the second router (232-S) is input to the second forwarding input port (⑤) of the reference router (232-Ref).
[0089] Meanwhile, referring back to FIG. 3, the first forwarding output port (②) and the first forwarding input port (③) are not illustrated for Router 1 (232-1). Since Router 1 (232-1) may be physically or software-wise located at the far left, the first forwarding output port (②) and the first forwarding input port (③) may not physically exist. Alternatively, Router 1 (232-1) may physically have the first forwarding output port (②) and the first forwarding input port (③), but may not use them in software. Router 8 (232-8) also does not have the second forwarding output port (④) and the second forwarding input port (⑤) illustrated for the same reason as Router 1 (232-1).
[0090] Meanwhile, in this specification, it will be assumed that the router (232) transmits data in a counterclockwise direction. Accordingly, it will be assumed that each router (232) transmits data input through the main input port (①) and the second transmission input port (⑤) only through the first transmission output port (②). In addition, it will be assumed that each router (232) transmits data input through the first transmission input port (③) only through the second transmission output port (④). When the input / output ports are set during the data transmission process as described above, duplicate output of data can be prevented. In addition, it should be understood that the counterclockwise transmission does not limit the operation processing device according to this specification, and that if it is changed to a clockwise direction, the relationship between the input / output ports may also change.
[0091] The router (232) according to the present specification can read the node ID of data input through the main input port (①) and process data with the corresponding node ID according to a data processing mapping table. According to one embodiment of the present specification, the data processing mapping table can store information on whether the input data is blocked, reflected, or output. In other words, the router (232) according to the present specification can process, according to the data processing mapping table, whether to block, reflect, or output data without transmitting it to another router based on the node ID. In relation to the data processing mapping table, the router (232) according to the present specification may be set to output data input from one adjacent router to another adjacent router (data transmission) as its default operation, but is not necessarily limited thereto. Therefore, the data processing mapping table can be said to be information on how to process data input from another router.
[0092] In the data processing mapping table, 'blocking' means not transmitting data input through the second transmission input port (⑤) or the first transmission input port (③) through the first transmission output port (②) or the second transmission output port (④). In the data processing mapping table, 'reflection' means outputting data input through the second transmission input port (⑤) through the second transmission output port (④). Alternatively, 'reflection' in the data processing mapping table means processing data that should be output through the first transmission output port (②) as if it were data input through the first transmission input port (③). In the data processing mapping table, 'output' means outputting data input through the first transmission input port (③) through the main output port (⑥).
[0093] Accordingly, depending on what is written in the data processing mapping table, the software topology configured by the plurality of routers (232) can be varied. The fetch network controller (231) can set whether to execute the blocking, reflection, and output according to the topology to be reconstructed, and thus the software topology can be determined by the fetch network controller (231). The data processing mapping table will be described in more detail through various embodiments of FIGS. 5 to 12.
[0094] Figure 5 is a software topology according to the first embodiment.
[0095] Referring to FIG. 5, the first embodiment is an example in which data fetched in each fetch buffer (222) is transferred to one feed buffer (242). That is, data stored in the first fetch buffer (222-1) is transferred only to the first feed buffer (242-1), and data stored in the second fetch buffer (222-2) is transferred only to the second feed buffer (242-2).
[0096] Figure 6 is a reference diagram of a data processing mapping table according to the first embodiment.
[0097] Referring to Figure 6, the data processing mapping table can be seen to be divided into data processing methods (e.g., blocking, reflection, output) based on node ID. Furthermore, entries in the data processing mapping table may indicate whether or not they are to be executed. A "1" indicates that the entry is to be executed, while a "0" indicates that the entry is not to be executed.
[0098] According to the data processing mapping table, looking at Router 1 (232-1) of FIG. 5, data for ID #1 is not blocked, but is reflected and output. Since it is assumed in this specification that the router (232) transmits data counterclockwise, data for ID #1 input into the main input port (①) of Router 1 (232-1) will be transmitted through the first transmission output port (②). At this time, since the data processing mapping table of Router 1 (232-1) is set to reflect the data for ID #1, the data for ID #1, which should be output through the first transmission output port (②), is processed as data input through the first transmission input port (③). And since the data processing mapping table of Router 1 (232-1) is set to output the data for ID #1, the data for ID #1 is output through the main output port (⑥). And the remaining data for IDs #2 to #8 are blocked, not reflected, and not output. Accordingly, data fetched to the first fetch buffer (222-1) of Fig. 5 can only be output to the first feed buffer (242-1). Routers 2 (232-2) to 8 (232-8) of Fig. 5 also operate according to the same principle, so unnecessary repeated explanations will be omitted.
[0099] Figure 7 is a software topology according to the second embodiment.
[0100] Referring to FIG. 7, the second embodiment is an example in which data fetched in each fetch buffer (222) is transferred to two feed buffers (242). That is, data stored in the first fetch buffer (222-1) is transferred to the first feed buffer (242-1) and the second feed buffer (242-2), and data stored in the second fetch buffer (222-2) is transferred to the first feed buffer (242-1) and the second feed buffer (242-2).
[0101] Figure 8 is a reference diagram of a data processing mapping table according to the second embodiment.
[0102] According to the data processing mapping table above, looking at Router 1 (232-1) of FIG. 7, it does not block, reflect, and output ID #1 data. Since it has been described in the first embodiment how Router 1 (232-1) outputs ID #1 data to Feed Buffer 1 (242-1), a repeated explanation will be omitted. In addition, since ID #1 data is processed as data input through the first transmission input port (③), ID #1 data can be transmitted to Router 2 (232-2) through the second transmission output port (④). Looking at Router 2 (232-2) of FIG. 7, it does not block, reflect, and output ID #1 data. Therefore, when ID #1 data is input from Router 1 (232-1), Router 2 (232-2) can output it to Feed Buffer 2 (242-2). Accordingly, data fetched to the first fetch buffer (222-1) of FIG. 7 can be output to the first feed buffer (242-1) and the second feed buffer (242-2).
[0103] Looking at Router 2 (232-2), it does not block or reflect ID #2 data, and outputs it. Since it is assumed in this specification that Router (232) transmits data counterclockwise, ID #2 data input to the main input port (①) of Router 2 (232-2) can be transmitted to Router 1 (232-1) through the first transmission output port (②). And looking at Router 1 (232-1) of Fig. 7, it does not block ID #2 data, but reflects it, and outputs it. ID #2 data input to Router 1 (232-1) can be output to Feed Buffer 1 (242-1) by Router 1 (232-1) in the same way as ID #1 data. And ID #2 data can be transmitted to Feed Buffer (242-2) again. Router #2 (232-2) can output ID #2 data, which is input again from Router #1 (232-1) through the first transmission input port (③), through the main output port (⑥). Accordingly, data fetched to fetch buffer #2 (222-2) of Fig. 7 can be output to feed buffer #1 (242-1) and feed buffer #2 (242-2).
[0104] Meanwhile, in FIG. 7, since the ID #1 data is input through the first transmission input port (③) of the second router (232-2), it is output through the second transmission output port (④) of the second router (232-2). Therefore, the ID #1 data is not input to the first router (232-1) again. Additionally, the third router (232-3) blocks the ID #1 data input through its first transmission input port (③). Furthermore, the router (232-3) also blocks the ID #2 data input through its first transmission input port (③).
[0105] Router 1 (232-1) and router 2 (232-2) of Fig. 7 block, do not reflect, and do not output the remaining ID #3 to #8 data.
[0106] Router 3 (232-3), router 4 (232-4), router 5 (232-5), router 6 (232-6), router 7 (232-7) and router 8 (232-8) of Fig. 7 also operate on the same principle, so unnecessary repeated explanations will be omitted.
[0107] Figure 9 is a software topology according to the third embodiment.
[0108] Referring to FIG. 9, the third embodiment is an example in which data fetched in each fetch buffer (222) is transferred to four feed buffers (242). That is, this is an example in which data stored in fetch buffers 1 to 4 (222-1 to 222-4) are transferred to feed buffers 1 to 4 (242-1 to 242-4), respectively.
[0109] Figure 10 is a reference diagram of a data processing mapping table according to the third embodiment.
[0110] Since the processing of ID #1 data and ID #2 data has been described through the first and second embodiments above, the processing of ID #3 data fetched to the third fetch buffer (222-3) of FIG. 9 will be examined representatively in FIG. 10. First, the ID #3 data fetched to the third fetch buffer (222-3) is input through the main input port (①) of the third router (232-3) and output to the first transmission output port (②) of the third router (232-3).
[0111] Router #2 (232-2) receives ID #3 data through the second transmission input port (⑤) and outputs ID #3 data to the first transmission output port (②) of router #2 (232-2).
[0112] Router 1 (232-1) receives ID #3 data through the second forwarding input port (⑤). According to the data processing mapping table of Router 1 (232-1), Router 1 (232-1) performs reflection and output for ID #3 data, so ID #3 data is output to Feed Buffer 1 (242-1) through the Main Output Port (⑥) and then output to the Second Forwarding Output Port (④) of Router 1 (232-1).
[0113] Router #2 (232-2) receives ID #3 data through the first transmission input port (③). According to the data processing mapping table of Router #2 (232-2), Router #2 (232-2) performs output for ID #3 data, so ID #3 data is output to Feed Buffer #2 (242-2) through the Main Output Port (⑥) and then output to the Second Transmission Output Port (④) of Router #2 (232-2).
[0114] Router #3 (232-3) receives ID #3 data through the first transmission input port (③). According to the data processing mapping table of Router #3 (232-3), Router #3 (232-3) performs output on ID #3 data, so ID #3 data is output to Feed Buffer #3 (242-3) through Main Output Port (⑥) and then output to Second Transmission Output Port (④) of Router #3 (232-3).
[0115] Router #4 (232-4) receives ID #3 data through the first transmission input port (③). According to the data processing mapping table of router #4 (232-4), router #4 (232-4) performs output for ID #3 data, so ID #3 data is output to feed buffer #4 (242-4) through main output port (⑥) and then output to second transmission output port (④) of router #4 (232-4).
[0116] Router #5 (232-5) receives data ID #3 through the first transmission input port (③). According to the data processing mapping table of Router #5 (232-5), Router #5 (232-5) executes blocking on data ID #3, so data ID #3 is no longer output or transmitted.
[0117] Accordingly, the ID #3 data fetched to the third fetch buffer (222-3) of FIG. 9 can be transferred to the first to fourth feed buffers (242-1 to 242-4). In addition, the ID #1 data, the ID #2 data, and the ID #4 data can also be transferred to the first to fourth feed buffers (242-1 to 242-4) in the same manner. Meanwhile, the ID #5 data to the ID #8 data can be transferred to the fifth to eighth feed buffers (242-5 to 242-8) by the same principle.
[0118] Figure 11 is a software topology according to the fourth embodiment.
[0119] Referring to FIG. 11, the fourth embodiment is an example in which data fetched to each fetch buffer (222) is transferred to all feed buffers (242). That is, this is an example in which data stored in fetch buffers 1 to 8 (222-1 to 222-8) are transferred to feed buffers 1 to 8 (242-1 to 242-8), respectively.
[0120] Figure 12 is a reference diagram of a data processing mapping table according to the fourth embodiment.
[0121] As previously described in detail through the first to third embodiments how each router (232) processes input data according to the data processing mapping table, a repetitive explanation will be omitted.
[0122] As can be seen from FIGS. 5 to 12, the fetch network controller (231) can set the blocking and output in the data processing mapping table of the router (232) belonging to the same group in the reconstructed software topology to be identical.
[0123] Meanwhile, we have described how a router (232) processes a single piece of data. However, multiple pieces of data fetched into multiple fetch buffers (222) must be processed simultaneously. Conventional techniques can employ a method of providing a sufficiently large buffer in the router and resolving conflicts when they occur. However, it is self-evident that a sufficiently large buffer can cause various problems, such as increased processor size, increased power consumption, increased heat generation, and processing delays. While it is generally recognized that providing an appropriately sized buffer is optimal, if the buffer in a single router becomes full, data input / output in all connected routers must be delayed. Furthermore, if the delay between routers is variable (e.g., different clock domains), more complex problems can arise. Therefore, a method is needed to appropriately control the timing of inputting the fetched data to each router (232). The processing unit (10) according to the present specification can provide a method for more effectively processing multiple pieces of data.
[0124] FIG. 13 is an example diagram of data input and output control according to one embodiment of the present specification.
[0125] Referring to Fig. 13, it can be confirmed that the software topology of the plurality of routers (232) is the same as the third embodiment illustrated in Fig. 9. Accordingly, when data fetched to the first to fourth fetch buffers (222-1 to 222-4) is input to the first to fourth routers (232-1 to 232-4), it must be output to the first to fourth feed buffers (242-1 to 242-4) without causing a collision. Meanwhile, in this specification, the control of data input and output will be described through the third embodiment, but the operation processing device (10) according to this specification is not limited to the above example.
[0126] In the example illustrated in Fig. 13, one data packet is described as consisting of two flits. In the example illustrated in Fig. 13, the numbers written in the flits indicate the order in which they are input into the feed buffer (242).
[0127] In this specification, data input from memory (100) to router (232) may include tail information and tail maintenance information. For example, data fetched to the fetch buffer (222), i.e., flit, may have a format of header, payload, tail maintenance information (keep_tail), and tail information (tail). The flit illustrated in FIG. 13 is in a format in which 'payload (P)' is followed by 'tail maintenance information (keep_tail)' and 'tail information (tail).'
[0128] The above tail information (tail) may be either a first logic value (T) meaning the end of the packet or a second logic value (F) meaning that it is not the end of the packet, and the tail maintenance information (keep_tail) may be either a first logic value (T) meaning that the tail information is maintained or a second logic value (F) meaning that the tail information is changed.
[0129] According to one embodiment of the present specification, the interface controller (221) can further set the tail information and the tail maintenance information to the fetched data.
[0130] The above interface controller (221) can set the tail information of the last data of the packet among the fetched data to the first logic value (T) and set the tail information of the remaining data to the second logic value (F).
[0131] In addition, the interface controller (221) may set the tail maintenance information of data having the last node ID among the software topology configurations of the router configured by the fetch network controller (231) through the data processing mapping table to a first logic value (T) meaning to maintain the tail information, and may set the tail maintenance information of data having the remaining node IDs to a second logic value (F) meaning to change the tail information.
[0132] According to another embodiment of the present specification, a separate component for setting the tail information and the tail maintenance information to the fetched data may be included in the network interface (220).
[0133] According to the above settings, looking at FIG. 13, it can be confirmed that the tail information of the fleet with an even order among the data fetched in each buffer (222) is 'F' and the tail information of the fleet with an odd order is 'T'. In addition, it can be confirmed that the tail maintenance information of only the 4th buffer (222-4) and the 8th buffer (222-8) with the last node ID of the software topology is 'T'.
[0134] A router (232) according to the present specification may include an arbiter unit (233) that controls an input port according to tail information of input data and changes the tail information according to tail maintenance information of the input data. The operation of the arbiter unit will be described in more detail.
[0135] Figures 14 to 20 are reference diagrams for processing data by the arbiter unit.
[0136] To simplify the drawing, only patch buffers 1 to 4 (222-1 to 222-4) and routers 1 to 4 (232-1 to 232-4) are shown in FIGS. 14 to 20.
[0137] Referring to Fig. 14, the arbiter unit (233) can first set an initial state so that data is input from the memory (100), i.e., from the main input port (①), among the ports related to the input / output of the router (232). Due to the initial state, two flits fetched into each buffer (222) can be sequentially input into each router (232).
[0138] When the above arbiter unit (233) controls the input port so that data is input from the memory (100), if the tail information of the data input from the memory (100), i.e., the fetch buffer (222), is the first logic value (T), the input port can be changed so that data is input from an adjacent router.
[0139] Referring to Fig. 15, it can be confirmed that the arbiter section has been changed so that data is input from the main input port (①) of each router to the second transmission input port (⑤) since the tail information of the second fleet input to each router is 'T'.
[0140] Meanwhile, the arbiter unit can change the tail information of the input data to the second logic value (F) when the tail maintenance information of the data input from the memory (main input port) is the second logic value (F). In the example illustrated in Fig. 13, since the tail maintenance information of flits F0 to F5 is all 'F', the tail information of flits F0 to F5, particularly the tail information of flits F1, F3, and F5, will be changed to 'F'. On the other hand, since the tail maintenance information of flits F6 and F7 is all 'T', the tail information of flits F6 and F7 is not changed. In particular, it should be noted that the tail information of flit F6 continues to be maintained as 'F', but the tail information of flit F7 is maintained as 'T'.
[0141] And data whose tail information is maintained or changed can be output to an adjacent router. The arbiter unit (233) can control data to be input from the memory when the tail information of the data input from the adjacent router is a first logic value (T). Referring to Fig. 16, it can be confirmed that data has been transmitted to the adjacent router. In particular, router number 3 (232-3) will be examined. Router number 3 (232-3) receives input from the second transmission input port (⑤) flits F6 and F7. Among these, since the tail information of flit F7 is 'T', the input port can be controlled so that data is input from the main input port (①).
[0142] Meanwhile, in a case where no data is input from the adjacent router, such as router number 4 (232-4), the arbiter unit (233) can control data to be input from the main input port (①) when the tail information of the data output to the adjacent router through its first transmission output port (②) is the first logic value (T).
[0143] Looking at FIGS. 17 through 20, we can see that data is transmitted sequentially. In particular, looking at the data output from Router 1 (232-1) in FIG. 20, we can see that it is output in fleet order. Furthermore, we can see that the states of Fleets 8 through 15 in FIG. 20 are quite similar to the initial state mentioned in FIG. 14. Therefore, we can predict that Fleets 8 through 15 will also be transmitted sequentially.
[0144] As previously described with reference to FIGS. 3 to 12 of this specification, the patch unit (200) according to this specification can configure a software topology through a data processing mapping table of a router, and can output data fetched to a network interface by assigning a node ID to the data from a specific router. In addition, as previously described with reference to FIGS. 13 to 20 of this specification, the patch unit (200) according to this specification has a structure capable of guaranteeing sequential output of packet data. Therefore, through the above two functions, the patch unit (200) according to this specification can transpose data. In this specification, since the transpose of data is performed internally by the patch unit (200), this can be called 'inter-transpose'.
[0145] For inter-transposition, the interface controller (221) according to the present specification can assign a destination node ID to the data fetched into each data memory slice. This will be explained in more detail with an example.
[0146] FIG. 21 is a reference diagram of inter-transpose according to one embodiment of the present specification.
[0147] Referring to Fig. 21, it can be confirmed that the software topology of multiple routers (232) is the same as the third embodiment illustrated in Fig. 9 or Fig. 13. However, the embodiment illustrated in Fig. 21 differs from the previous embodiment in the node ID and the 'tail maintenance information' and 'tail information' of each data in Fig. 9 and Fig. 13. In the example illustrated in Fig. 21, it is assumed that one flit is one packet, and all 'tail information' is set to the first logical value (T) indicating the end of the packet. This is for convenience of explanation, and the present invention is not limited by the above embodiment.
[0148] FIG. 22 is an example of a data processing mapping table for inter-transposition according to one embodiment of the present specification.
[0149] Referring to Fig. 22, it can be seen that although it is similar to the data processing mapping table illustrated in Fig. 10, the "output" item has a difference. As previously explained, the fetch network controller (231) can set whether to execute the blocking and reflection according to the topology to be reconstructed. However, the fetch network controller can set the blocking in the data processing mapping table of the routers belonging to the same group in the reconstructed topology to be the same, and set it to output only one node ID. This point will be explained in detail later. Meanwhile, in this specification, the control of data input and output will be described through the third embodiment, but the operation processing device (10) according to this specification is not limited to the above example.
[0150] Figure 23 is a reference diagram of the results of inter-transposition according to one embodiment of the present specification.
[0151] Referring to FIG. 23, the network interface (220) can assign a destination node ID equal to the number of routers (232) belonging to the same group by the fetch network controller (231) to the data fetched to each data memory slice (101). In the example illustrated in FIG. 23, four routers (232-1, 232-2, 232-3, 232-4) constitute one group, and the number of IDs is four (ID#0, ID#1, ID#2, ID#3). In addition, the network interface (220) can sequentially and repeatedly assign a destination node ID to the data fetched to each data memory slice (101) in the order in which they were fetched. For example, if we look at the data fetched in the 1st fetch buffer (222-1), we can see that the destination node IDs "ID#0, ID#1, ID#2, ID#3" are sequentially assigned according to the order in which the data was fetched (F0, F1, F2, F3). Although the drawing does not show additional data fetched thereafter, the IDs may be assigned in the same manner.
[0152] Now, let's assume a situation where the first data is input to each router (232). 'F0 (ID#0,F,T)' is input to the first router (232-1), 'F4 (ID#0,F,T)' is input to the second router (232-2), 'F8 (ID#0,F,T)' is input to the third router (232-3), and 'F12 (ID#0,F,T)' is input to the fourth router (232-4). The above data will be processed according to the sequential delivery guarantee described above. At this time, the first router (232-1) is set to output data with 'ID#0' by the data processing mapping table (see Fig. 22). Accordingly, the above 'F0 (ID#0,F,T), F4 (ID#0,F,T), F8 (ID#0,F,T) and F12 (ID#0,F,T)' can all be output to router 1 (232-1).
[0153] Next, let's assume a situation where the second data is input to each router (232). 'F1 (ID#1,F,T)' is input to the first router (232-1), 'F5 (ID#1,F,T)' is input to the second router (232-2), 'F9 (ID#1,F,T)' is input to the third router (232-3), and 'F13 (ID#1,F,T)' is input to the fourth router (232-4). The above data will be processed according to the sequential delivery guarantee described above. At this time, the second router (232-2) is set to output data with 'ID#1' by the data processing mapping table (see Fig. 22). Accordingly, the above 'F1 (ID#1,F,T), F5 (ID#1,F,T), F9 (ID#1,F,T) and F13 (ID#1,F,T)' can all be output to router 2 (232-2).
[0154] The input and output of the third and fourth data also undergo the same process. Therefore, comparing the output data with the input data confirms that it has been transposed. When one of the two axes constituting the data is stored within a data memory slice and the other is stored as memory slices, the two axes can be transposed, as illustrated in Figures 21 to 23.
[0155] While the embodiments of this specification have been described with reference to the attached drawings, those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering the technical spirit or essential features thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.
[0156] <Explanation of symbols>
[0157] 10: Operational processing unit
[0158] 100: Memory 101: Data Memory Slice
[0159] 200: Fetch Unit 210: Fetch Sequencer
[0160] 220: Network interface 221: Interface controller
[0161] 222: Fetch Buffer 230: Fetch Network
[0162] 231: Fetch Network Controller 232: Router
[0163] 233: Arbiter 240: Feed Module
[0164] 242: Feed Buffer 250: Operation Sequencer Module
[0165] 300: Operational Unit 310: Dot Product Engine
[0166] 400: Commit Unit 410: Commit Sequencer
Claims
1. In a computational processing device including a plurality of fetch units that read data required for computation to perform processing of a neural network from memory and provide it to a computational unit, Each fetch unit, A network interface having a fetch buffer from which data stored in the above memory is fetched; A fetch network having a data processing mapping table describing a method of processing the input data according to the node ID of the data fetched to the network interface, and having a plurality of routers that ensure sequential delivery; and A fetch network controller that reconfigures each data processing mapping table of the plurality of routers to form a software topology according to the operation form; The above memory is, Contains a number of data memory slices equal to the number of said plurality of routers; The above network interface is, An operation processing device characterized by including an interface controller that assigns a destination node ID to data fetched into each of the above data memory slices.
2. In claim 1, The above network interface is an operation processing device characterized in that it assigns a number of destination node IDs equal to the number of routers belonging to the same group by the fetching network controller to the data fetched to each data memory slice.
3. In claim 2, An operation processing device, characterized in that the above network interface sequentially and repeatedly assigns a destination node ID to data fetched to each data memory slice in the order in which they were fetched.
4. In claim 1, Each router, An arithmetic processing device characterized by comprising a main input port into which data is input from the memory, a first transmission output port for transmitting data to an adjacent first router, a first transmission input port for receiving data transmitted from an adjacent first router, a second transmission output port for transmitting data to an adjacent second router, a second transmission input port for receiving data transmitted from an adjacent second router, and a main output port for providing data to the arithmetic unit.
5. In claim 1, An operation processing device characterized in that the above data processing mapping table stores information on whether blocking, reflection, and output of input data are executed.
6. In claim 5, The above fetch network controller is an operation processing device characterized in that it sets whether to execute the blocking and reflection according to the topology to be reconstructed.
7. In claim 6, The above fetch network controller is an operation processing device characterized in that the blocking in the data processing mapping table of the router belonging to the same group in the reconstructed topology is set to be the same and only one node ID is output.
8. In claim 1, An arithmetic processing device, wherein each router further includes an arbiter unit that controls an input port according to tail information of input data and changes the tail information according to tail maintenance information of the input data.
9. In claim 8, The above tail information is either a first logic value (T) meaning the end of the packet or a second logic value (F) meaning not the end of the packet. An arithmetic processing device, characterized in that the tail maintenance information is either a first logic value (T) meaning to maintain the tail information or a second logic value (F) meaning to change the tail information.
10. In claim 9, The above arbiter part, An arithmetic processing device characterized in that, when controlling an input port so that data is input from the above memory, the input port is changed so that data is input from an adjacent router if the tail information of the data input from the above memory is a first logic value (T).
11. In claim 9, The above arbiter part, An arithmetic processing device characterized in that when the tail maintenance information of the data input from the above memory is a second logic value (F), the tail information of the input data is changed to a second logic value (F).
12. In claim 9, The above arbiter part, An arithmetic processing device characterized in that it controls data to be input from the memory when the tail information of data input from the adjacent router is a first logic value (T).
13. In claim 1, The above network interface is an operation processing device that further sets the tail information and the tail maintenance information to the fetched data.
14. In claim 13, The above network interface is an operation processing device characterized in that it sets the tail information of the last data of the packet among the fetched data to the first logic value (T) and sets the tail information of the remaining data to the second logic value (F).
15. In claim 13, The above network interface is an arithmetic processing device characterized in that it sets the tail maintenance information of data having the last node ID among the software topology configurations of the router to a first logic value (T) meaning to maintain the tail information, and sets the tail maintenance information of data having the remaining node IDs to a second logic value (F) meaning to change the tail information.
Citation Information
Patent Citations
Packet network capable of inter-transposing
KR1020250081199A
Scheduling threads in a processor
KR101486025B1
Data storage structure for improving memory cell reliability
KR1020220054763A
Wireless charging system and apparatus for replacement-type battery
KR1020220066855A
Item information offering method of electronic apparatus
KR102441999B1