Compiling device, generation method, program, chip, and system

JP2023064233A5Active Publication Date: 2025-06-24PREFERRED NETWORKS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2021174381
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-06-24
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

Existing deep learning accelerator chips face challenges in efficiently placing tensor elements in multiple memories connected by a tree-structured topology, leading to suboptimal communication costs and operation constraints.

Method used

A compiling device generates machine code that assigns addresses to tensor elements based on the number of divisions and strides in a predetermined hierarchy of the tree structure, allowing for appropriate placement of tensor elements in multiple memories connected by a tree topology, reducing communication costs and enhancing operation efficiency.

Benefits of technology

This approach enables efficient allocation of tensor elements, reducing communication costs and facilitating operations on deep learning accelerator chips with tree-structured topologies, while allowing for intuitive understanding and alignment of element placement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To properly represent an arrangement of each element of a tensor for multiple memories connected by a tree structure topology.SOLUTION: A compiler, for generating a machine code to be executed in a chip including a plurality of distributed memories connected by a tree structure topology, includes at least one memory and at least one processor. The at least one processor is configured to associate each element of a tensor to be processed with an address in the plurality of memories included in the chip, based on a stride and the number of divisions in a predetermined layer of the tree structure with respect to the tensor to be processed.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a compilation device, a generation method, a chip, and an execution method. [Background technology]

[0002] When writing source code, users can specify where in memory each element of a tensor should be located.

[0003] On the other hand, for example, accelerator chips for deep learning may have multiple memory modules (SRAM: Static Random Access Memory) connected by a tree topology, distributed across them, and operate using a SIMD (Single Instruction / Multiple Data) architecture. Therefore, when processing each element of a tensor using such an accelerator chip, it is important to determine which memory module and location each element of the tensor is placed in. [Prior art documents] [Patent Documents]

[0004] [Patent Document 1] Special Publication No. 2020-517006 [Patent Document 2] Japanese Patent Publication No. 2007-242017 [Patent Document 3] Japanese Patent Application Publication No. 06-208501 [Overview of the project] [Problems that the invention aims to solve]

[0005] This disclosure enables the proper representation of the arrangement of each element of a tensor across multiple memories connected by a tree-structure topology. [Means for solving the problem]

[0006] A compilation device according to one aspect of the present disclosure has, for example, the following configuration. That is, A compilation device that generates machine code to be executed on a chip having a plurality of memories connected by a tree-structured topology and distributed, One or more memories, One or more processors, and is provided with The one or more processors Execute associating each element of the tensor to be processed with an address in the plurality of memories of the chip based on the number of divisions and strides in a predetermined hierarchy of the tree structure for the tensor to be processed.

Brief Description of Drawings

[0007] [Figure 1] It is a diagram showing an example of the system configuration of a data processing system and the hardware configuration of each device. [Figure 2] It is a first diagram showing an example of the functional configuration of each device of a data processing system. [Figure 3] It is a diagram showing an example of the hardware configuration of an accelerator chip. [Figure 4] It is a diagram showing a specific example of a plurality of memories connected by a tree-structured topology. [Figure 5] It is a diagram showing a description method of a description regarding layout. [Figure 6] It is a first diagram showing a specific example of a description regarding layout and processing by an allocation unit. [Figure 7] It is a second diagram showing a specific example of a description regarding layout and processing by an allocation unit. [Figure 8] It is a first diagram showing a specific example of processing by a writing unit. [Figure 9] It is a second diagram showing a specific example of processing by a writing unit. [Figure 10] It is a first diagram showing a specific example of processing by an element value reading unit. [Figure 11]The second figure shows a specific example of processing by the element value reading unit. [Figure 12] This flowchart outlines the source code generation process. [Figure 13] This is a flowchart showing the machine code generation process. [Figure 14] This is a flowchart showing the flow of machine code execution. [Figure 15] The second figure shows an example of the functional configuration of each device in a data processing system. [Modes for carrying out the invention]

[0008] Each embodiment will be described below with reference to the attached drawings. In this specification and the drawings, devices having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.

[0009] [First Embodiment] <System configuration of the data processing system and hardware configuration of each device> First, we will describe the overall system configuration of the data processing system having a server device according to the first embodiment, and the hardware configuration of each device constituting the data processing system.

[0010] As shown in Figure 1, the data processing system 100 includes a server device 110 and an external device 160. Also, as shown in Figure 1, the server device 110 includes a compilation device 120 and a data processing device 140.

[0011] The compilation unit 120 may, for example, include a processor 121, main memory 122 (memory), auxiliary memory 123 (memory), a network interface 124, and a device interface 125. The compilation unit 120 may also be implemented as a computer in which these devices are connected via a bus 130.

[0012] The processor 121 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, or ASIC, etc.). Alternatively, the processor 121 may be a semiconductor device including a dedicated processing circuit. Furthermore, the processor 121 is not limited to an electronic circuit using electronic logic elements, but may be implemented using an optical circuit with optical logic elements. The processor 121 may also include computational functions based on quantum computing.

[0013] The processor 121 performs various calculations based on various data and instructions input from the various devices of the internal configuration of the compilation unit 120, and outputs the calculation results and control signals to the respective devices. The processor 121 may also control the various devices of the compilation unit 120 by executing an OS (Operating System) or applications.

[0014] Furthermore, the processor 121 may refer to one or more electronic circuits arranged on a single chip, or to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, each electronic circuit may communicate by wire or wireless.

[0015] The main memory 122 is a storage device that stores instructions executed by the processor 121 and various data, and the various data stored in the main memory 122 is read by the processor 121. The auxiliary storage device 123 is a storage device other than the main memory 122. These storage devices refer to any electronic component capable of storing various data, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The storage device for storing various data in the compilation device 120 may be implemented by the main memory 122 or the auxiliary storage device 123, or by the built-in memory of the processor 121.

[0016] The network interface 124 is an interface for connecting to the communication network 150 wirelessly or via a wired connection. The network interface 124 uses an appropriate interface, such as one conforming to existing communication standards. The communication network 150 may be a WAN (Wide Area Network), LAN (Local Area Network), PAN (Personal Area Network), or a combination thereof. An example of a WAN is the internet; an example of a LAN is IEEE 802.11 or Ethernet; and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication).

[0017] The device interface 125 is an interface such as USB that connects directly to the external device 160.

[0018] The external device 160 is a device connected to the computer. The external device 160 may, for example, be an input device. The input device is, for example, an operating device 161 such as a keyboard, mouse, or touch panel, which provides acquired information to the computer.

[0019] Furthermore, the external device 160 may, for example, be an output device. The output device may be a display device 162 such as an LCD (Liquid Crystal Display), CRT (Cathode Ray Tube), PDP (Plasma Display Panel), or organic EL (Electro Luminescence) panel.

[0020] The data processing device 140 has multiple boards (boards 140_1 to 140_4) as individual devices. Boards 140_1 to 140_4 are equipped with multiple accelerator chips (for example, chips 170_1 to 170_n).

[0021] Furthermore, as shown in Figure 1, each unit of the compilation unit 120 and each unit of the data processing unit 140 are connected via the bus 130. Note that the example in Figure 1 shows the case where the data processing unit 140 has four boards 140_1 to 140_4, but the number of boards that the data processing unit 140 has is arbitrary.

[0022] Chips 170_1 to 170_n are specialized chips specifically designed for the training phase of deep learning, for example. Details of chips 170_1 to 170_n will be described later.

[0023] <Functional configuration of each device in the data processing system> Next, the functional configuration of each device in the data processing system 100 (here, the server device 110 and the display device 162) will be described. Figure 2 is the first diagram showing an example of the functional configuration of each device in the data processing system.

[0024] The compilation device 120 has a generation program for generating source code and a compiler for generating machine code installed, and when these programs are executed, the compilation device 120 performs the following: • Source code description section 211, ·Generation unit 212, Compilation section 213, It functions as such.

[0025] The user of the compilation device 120 starts writing source code by activating the source code writing unit 211. In Figure 2, source code 230 is an example of source code being written, displayed on the display device 162. In this embodiment, source code 230 includes descriptions related to tensors, layouts, indexes, etc. The written source code 230 is then notified to the generation unit 212.

[0026] The generation unit 212 generates a computation graph based on the source code 230. A computation graph is a graphical representation of the computation flow from the input tensor to the generation of the output tensor, or a graphical representation of the computation flow for updating the values ​​of a tensor. For example, if the source code 230 is written in Python® code, the computation graph is generated by executing the source code 230 and converting it to the ONNX representation format. ONNX is an abbreviation for Open Neural Network Exchange.

[0027] Furthermore, the generation unit 212 generates a layout instruction sheet based on the source code 230. The layout instruction sheet is information for executing the process of assigning an address to each element of the tensor, which is generated based on the layout description contained in the source code 230. The "process of assigning an address to each element of the tensor" is an example of the "process of associating each element of the tensor with an address". The "process of associating each element of the tensor with an address" includes at least one of the following: "the process of assigning an address to each element of the tensor" or "the process of assigning each element of the tensor to an address".

[0028] The calculation graph and layout instructions (hereinafter referred to as "calculation graph, etc.") generated in the generation unit 212 are then communicated to the compilation unit 213.

[0029] The compilation unit 213 receives the computation graph and other information notified by the generation unit 212 as input and performs the compilation process to generate machine code. At this time, the compilation unit 213 functions as an allocation unit. Specifically, the compilation unit 213, for example, assigns an address to one of the memories (for example, SRAM) within chips 170_1 to 170_n to each element of the tensor, based on a layout instruction generated based on a description of the layout.

[0030] The generated machine code, along with the data stored in the data storage unit 214, is input to the data processing device 140.

[0031] Boards 140_1 to 140_4 of the data processing device 140 function as execution units 220 that execute machine code generated by the compilation unit 213 and process data stored in the data storage unit 214.

[0032] At this time, the execution unit 220 functions as a writing unit 251. The writing unit 251 writes, for example, the values ​​of each element of a tensor (data stored in the data storage unit 214) to memory addresses within chips 170_1 to 170_n that have been allocated by the allocation unit 241, based on a description of the tensor.

[0033] Furthermore, the execution unit 220 functions as an element value reading unit 252. The element value reading unit 252 reads the value of a specific element of a tensor written to the memory within chips 170_1 to 170_n, for example, based on a description of the index.

[0034] Furthermore, the execution unit 220 functions as an auxiliary writing unit 253. The auxiliary writing unit 253, for example, completes the values ​​of each element of the tensor based on a description of the layout. Specifically, the auxiliary writing unit 253 performs padding to complete the values ​​of missing elements so that the size of the tensor is adjusted according to the memory to which it is written. In addition, the auxiliary writing unit 253 performs broadcast processing to match the shapes when performing operations on the elements of tensors whose array shapes do not match.

[0035] <Hardware configuration of the accelerator chip> Next, we will describe the hardware configuration of the accelerator chips (for example, chips 170_1 to 170_n) mounted on boards 140_1 to 140_4, etc. Figure 3 shows an example of the hardware configuration of an accelerator chip.

[0036] Chip 170_1 (since chips 170_1 to 170_n all have the same hardware configuration, this explanation will focus on chip 170_1) operates, for example, using a SIMD architecture. SIMD stands for Single Instruction / Multiple Data, and refers to a method of applying a single instruction to multiple data simultaneously and processing them in parallel. However, chip 170_1 may also operate using architectures other than SIMD.

[0037] As shown in Figure 3, the chip 170_1 has four third-tier blocks. Each third-tier block has four second-tier blocks. Each second-tier block has multiple first-tier blocks and one second-tier block memory.

[0038] Furthermore, each first-level block has one arithmetic unit and four arithmetic units. The four arithmetic units supply data to the arithmetic unit.

[0039] Thus, in chip 170_1, multiple first-level blocks are distributed across four second-level blocks and four third-level blocks, which are connected by a tree-structure topology. Therefore, the communication cost between memories contained within multiple first-level blocks is not uniform within chip 170_1. For example, communication between nearby memories is low-cost, while communication between memories that require traversing the tree structure is high-cost.

[0040] <Tree structure topology> Next, we will describe a specific example of multiple memories connected by a tree topology. Figure 4 is a diagram showing a specific example of multiple memories connected by a tree topology.

[0041] As shown in the example in Figure 4, the four third-level blocks belong to the Level A hierarchy of the tree structure and are connected to each other. Furthermore, the four second-level blocks contained within each third-level block all belong to the Level B hierarchy of the tree structure and are connected to the corresponding third-level blocks in the Level A hierarchy of the tree structure.

[0042] Furthermore, each of the four first-level blocks contained within the second-level block belonging to the Level B hierarchy of the tree structure belongs to the Level C hierarchy of the tree structure, and each is connected to the corresponding second-level block in the Level B hierarchy of the tree structure.

[0043] Here, for example, The value written to memory 411 included in the first-level block of Level C, indicated by reference numeral 401, • Memory 412 included in the first level block of Level C shown in reference numeral 402, Let's consider the case where we need to move it.

[0044] In this case, chip 170_1 is, • Tracing back through the tree structure hierarchy from Level C to Level B to Level A, • Crossing different blocks within Level A, • Proceed through the tree structure hierarchy from Level A to Level B to Level C. This requires following certain steps, which incurs communication costs. On the other hand, to reduce communication costs, it is effective to write the value to memory near memory 412 instead of writing it to memory 411.

[0045] In other words, in the case of chip 170_1, where multiple memories connected by a tree topology are distributed, it is important to appropriately assign memory addresses to each element of the tensor so that the values ​​of each element of the tensor are written to the memory that takes into account the hierarchy of the tree structure.

[0046] In this embodiment, the data processing system 100, A source code description section 211 that performs a "layout description" using a description method that allows appropriate memory addresses to be allocated to each element of the tensor, The compilation unit 213 assigns addresses to each element of the tensor according to the said description method, The execution unit 220 writes the values ​​of each element of the tensor (data stored in the data storage unit 214) to the assigned address, To provide.

[0047] <How to describe layout-related information> Next, we will explain how to describe the layout. Figure 5 is a diagram showing how to describe the layout.

[0048] As shown in Figure 5, the layout description includes descriptions of vertical and horizontal arrangement within parentheses, separated by a comma.

[0049] Furthermore, as shown in Figure 5, the description of the vertical arrangement includes descriptions of the first level, the second level, ..., and also includes a description of the memory at the lowest level. The description of the first level is, for example, the description of Level A in Figure 4, and the description of the second level is, for example, the description of Level B in Figure 4. The description of the memory at the lowest level is, for example, the description of the memory included in the first level block of Level C in Figure 4.

[0050] Furthermore, as shown in Figure 5, the description of the Nth level (where N is an integer greater than or equal to 1) is in the format "Number of divisions_Level name:Stride", and the description of the memory of the lowest level is in the format "Number of divisions_Memory address:Stride".

[0051] In this context, "stride" refers to information indicating how many block names (addresses in the lowest level) are advanced when one block (or tensor element in the lowest level) is advanced vertically (one element in the lowest level) at each hierarchical level. However, the block name may be an identifier (number, name, etc.) that can identify each block.

[0052] For example, suppose the block names of the four third-level blocks in Level A are "A0" to "A3". Also, suppose the third-level blocks are arranged from left to right in the first row in the order "A0" and "A1", and from left to right in the second row in the order "A2" and "A3". In this case, when the third-level block named "A0" is moved one block vertically, the block name moves two blocks ("A0" → "A2", "A1" → "A3"). Therefore, in this arrangement direction, the stride is "2".

[0053] For example, suppose that in a Level C hierarchy, the memory addresses of 411 allocated to the elements of the first row of a tensor are "0" to "24", the memory addresses of 411 allocated to the elements of the second row are "25" to "49", ... In this case, when the tensor is moved one element vertically, the address moves 25 elements ("0" → "25" → ...). Therefore, for such memory, the stride is "25". Note that the above explanation of stride is just one example, and stride may be expressed in other forms as long as it indicates the change in block names vertically in each hierarchy.

[0054] Similarly, as shown in Figure 5, the description of the horizontal arrangement includes the description of the first level, the description of the second level, ..., and further includes the description of the memory of the lowest level.

[0055] Furthermore, as shown in Figure 5, the description of the Nth level (where N is an integer greater than or equal to 1) is in the format "Number of divisions_Level name:Stride", and the description of the memory of the lowest level is in the format "Number of divisions_Memory address:Stride".

[0056] In this context, "stride" refers to information indicating how many blocks (or addresses in the lowest level) are moved horizontally when a block (or tensor element in the lowest level) is advanced one position in each hierarchy.

[0057] For example, suppose the block names of the four third-level blocks in Level A are "A0" to "A3". Also, suppose the third-level blocks are arranged from left to right in the first row in the order "A0" and "A1", and from left to right in the second row in the order "A2" and "A3". In this case, when the third-level block named "A0" is moved one space horizontally, the block name also moves one space ("A0" becomes "A1", "A2" becomes "A3"). Therefore, in this arrangement direction, the stride is "1".

[0058] For example, suppose that in a Level C hierarchy, the memory addresses of 411 allocated to the elements of the first row of a tensor are "0" to "24", the memory addresses of 411 allocated to the elements of the second row are "25" to "49", ... In this case, when the elements of the tensor are advanced one position vertically, the address advances by one position ("0" → "1" → ...). Therefore, for such memory, the stride is "1". Note that the above explanation of stride is just one example, and stride may be expressed in other forms as long as it indicates the change in block names horizontally in each hierarchy.

[0059] In this way, by separating the description of the vertical arrangement from the description of the horizontal arrangement, and by specifying the number of divisions and the stride for each level, • A highly expressive description method can be achieved, and even when multiple memory locations are connected by a complex tree topology, the arrangement of each element of a tensor across multiple memory locations can be appropriately represented. This allows the chip 170_1 to assign appropriate addresses to each element of the tensor, thereby reducing the communication cost between memory locations. • It enables highly expressive description methods and can accommodate constraints imposed on each operation. Because users can intuitively understand the arrangement of each element of a tensor across multiple memory locations, they can optimize operations considering the arrangement of each tensor element, and arrange the tensor elements considering the characteristics of SIMD. • Because the arrangement of each element can be aligned across tensors, this is advantageous for operation using SIMD architectures. These are some of the advantages.

[0060] <Specific examples of layout descriptions and processing by the assignment section> (1) Specific example 1 Next, we will explain specific examples of layout descriptions. Figure 6 is the first example showing a specific example of a layout description. In the example in Figure 6, for the sake of simplicity, the number of levels is set to "2" (1st level = Level A, 2nd level = bottom level = Level B).

[0061] As shown in Figure 6(b), when all elements of a 100x100 tensor X are allocated to memory contained within the lowest-level Level B block of the chip 600 in Figure 6(a), the layout description would be, for example, ((2_A:2,2_B:2,25_Addr:25),(2_A:1,2_B:1,25_Addr:1)) This is the result.

[0062] Here, regarding the description of the vertical arrangement, 2_A:2 is, In Level A, divide the 100 elements vertically into two groups of 50 elements each. In Level A, moving a block one space vertically results in the block name moving two spaces (e.g., "A0" → "A2" or "A1" → "A3"). It represents.

[0063] Furthermore, regarding the description of the vertical arrangement, 2_B:2 is, In Level B, divide the 50 elements vertically into two groups of 25 elements each. In Level B, moving a block one space vertically results in the block name moving two spaces ("B0" → "B2" or "B1" → "B3"). It represents.

[0064] Furthermore, regarding the description of the vertical arrangement, 25_Addr:25 is, • In the memory contained within a Level B block, divide the 25 elements vertically into 25 sections. In the memory contained within a Level B block, advancing a tensor element by one position vertically advances the address by 25 positions (for example, address "0" → "25", "1" → "26", ...). It represents.

[0065] On the other hand, among the descriptions regarding the horizontal arrangement, 2_A:1 is, In Level A, divide the 100 elements horizontally into two groups of 50 elements each. In Level A, moving a block one space horizontally results in the block name advancing by one space (e.g., "A0" → "A1" or "A2" → "A3"). It represents.

[0066] Furthermore, regarding the description of the horizontal arrangement, 2_B:1 is, In Level B, divide the 50 elements horizontally into two groups of 25 elements each. In Level B, moving a block one space horizontally advances the block name by one space (e.g., "B0" → "B1" or "B2" → "B3"). It represents.

[0067] Furthermore, regarding the description of horizontal arrangement, 25_Addr:1 is, • In the memory contained within a Level B block, divide the 25 horizontal elements into 25 sections. In the memory contained within a Level B block, moving a tensor element one position horizontally advances the address by one position (for example, address "0" → "1", "1" → "2", ...). It represents.

[0068] Thus, based on the above description of the layout, the allocation unit 241 can assign the memory addresses included in the Level B block of the chip 600 to each of the 100 rows x 100 columns.

[0069] (2) Specific example 2 Next, we will explain other specific examples of layout descriptions. Figure 7 is the second figure showing a specific example of a layout description. In the example in Figure 7, for the sake of simplicity, the number of levels is set to "2" (1st level = Level A, 2nd level = bottom level = Level B). However, in the example in Figure 7, the way the blocks are divided is different from the example in Figure 6 (see Figure 7(a)).

[0070] As shown in Figure 7(b), when all elements of a 100x100 tensor X are allocated to memory contained within the lowest-level Level B block of the chip 700 in Figure 7(a), the layout description would be, for example, ((4_A:1,25_Addr:25),(4_B:1,25_Addr:1)) This is the result.

[0071] Here, among the descriptions regarding the vertical arrangement, 4_A:1 is, In Level A, divide the 100 elements vertically into four groups of 25 elements each. In Level A, moving a block one space vertically results in the block name advancing by one space ("A0" → "A1", "A1" → "A2", "A2" → "A3"). It represents.

[0072] Furthermore, regarding the description of the vertical arrangement, 25_Addr:25 is, • In the memory contained within a Level B block, divide the 25 elements vertically into 25 sections. In the memory contained within a Level B block, advancing a tensor element by one position vertically advances the address by 25 positions (for example, address "0" → "25", "1" → "26", ...). It represents.

[0073] On the other hand, among the descriptions regarding horizontal arrangement, 4_B:1 is, In Level B, divide the 100 elements horizontally into four groups of 25 elements each. In Level B, moving a block horizontally by one space corresponds to moving the block name by one space ("B0" → "B1", "B1" → "B2", "B2" → "B3"). It represents.

[0074] Furthermore, regarding the description of horizontal arrangement, 25_Addr:1 is, • In the memory contained within a Level B block, divide the 25 horizontal elements into 25 sections. In the memory contained within a Level B block, moving a tensor element one position horizontally advances the address by one position (for example, address "0" → "1", "1" → "2", ...). It represents.

[0075] Thus, based on the above description of the layout, the allocation unit 241 can assign the memory addresses included in the Level B block of the chip 700 to each of the 100 rows x 100 columns.

[0076] <Specific Example of Processing by Writing Unit> (1) Specific Example 1 Next, a specific example of the process of writing the value of each element of tensor X to the corresponding memory according to the address (Fig. 6) assigned by the assignment unit 241 will be described. Fig. 8 is the first figure showing a specific example of the process by the writing unit.

[0077] In Fig. 8, reference numeral 800 shows a specific example of the value of each element of the 100-row × 100-column tensor X (data stored in the data storage unit 214). Also, in Fig. 8, reference numeral 600' shows how the value of each element of tensor X is written into the memory included in the Level B block of chip 600.

[0078] For example, in the block with block name = "A0", in the memory included in the block with block name = "B0", · To addresses "0" to "24", x 1_1 ~x 1_25 is written, · To addresses "25" to "49", x 2_1 ~x 2_25 is written, ··· · To addresses "600" to "624", x 25_1 ~x 25_25 is written.

[0079] Also, in the block with block name = "A0", in the memory included in the block with block name = "B1", · To addresses "0" to "24", x 1_26 ~x 1_50 is written, · To addresses "25" to "49", x 2_26 ~x 2_50 is written, ··· · To addresses "600" to "624", x 25_26 ~x 25_50 is written.

[0080] Furthermore, within the block named "A0", the memory contained in the block named "B2", • Addresses "0" to "24" contain x 26_1 ~x 26_25 It was written, • Addresses "25" to "49" contain x 27_1 ~x 27_25 It was written, ... • Addresses "600" to "624" contain x 50_1 ~x 50_25 This will be written.

[0081] Furthermore, within the block named "A0", the memory contained in the block named "B3", • Addresses "0" to "24" contain x 26_26 ~x 26_50 It was written, • Addresses "25" to "49" contain x 27_26 ~x 27_50 It was written, ... • Addresses "600" to "624" contain x 50_26 ~x 50_50 This will be written.

[0082] Subsequently, the values ​​of each element of tensor X are written to the memory contained within the Level B block.

[0083] In this way, the writing unit 251 can write each element of the 100x100 grid to the memory contained in the Level B block of the chip 600.

[0084] (2) Specific example 2 Next, we will describe a specific example of the process of writing the values ​​of each element of tensor X to the corresponding memory according to the address (Figure 7) allocated by the allocation unit 241. Figure 9 is a second diagram showing a specific example of the process performed by the writing unit.

[0085] In Figure 9, the symbol 800 represents a specific example of the value of each element of a 100x100 tensor X (data stored in the data storage unit 214). Also in Figure 9, the symbol 700' shows how the values ​​of each element of the tensor X are written to the memory included in the Level B block of the chip 700.

[0086] For example, within block A0, the memory contained in block B0, • Addresses "0" to "24" contain x 1_1 ~x 1_25 It was written, • Addresses "25" to "49" contain x 2_1 ~x 2_25 It was written, ... • Addresses "600" to "624" contain x 25_1 ~x 25_25 This will be written.

[0087] Furthermore, within the block named "A0", the memory contained in the block named "B1", • Addresses "0" to "24" contain x 1_26 ~x 1_50 It was written, • Addresses "25" to "49" contain x 2_26 ~x 2_50 It was written, ... • Addresses "600" to "624" contain x 25_26 ~x 25_50 This will be written.

[0088] Furthermore, within the block named "A0", the memory contained in the block named "B2", • Addresses "0" to "24" contain x 1_51 ~x 1_75 It was written, • Addresses "25" to "49" contain x 2_51 ~x 2_75 It was written, ... • Addresses "600" to "624" contain x 25_51 ~x 25_75 This will be written.

[0089] Furthermore, within the block named "A0", the memory contained in the block named "B3", • Addresses "0" to "24" contain x 1_76 ~x 1_100 It was written, • Addresses "25" to "49" contain x 2_76 ~x 2_100 It was written, ... • Addresses "600" to "624" contain x 25_76 ~x 25_100 This will be written.

[0090] Subsequently, the values ​​of each element of tensor X are written to the memory contained within the Level B block.

[0091] In this way, the writing unit 251 can write each element of the 100x100 grid to the memory contained in the Level B block of the chip 700.

[0092] <Specific example of processing by the element value reading unit> Next, a specific example of processing by the element value reading unit 252 will be described. As mentioned above, the element value reading unit 252 reads the value of a specific element of the tensor written to memory based on the index description contained in the source code 230.

[0093] (1) Specific example 1 Figure 10 is the first diagram showing a specific example of processing by the element value reading unit. The example in Figure 10 shows how the values ​​of each element of the tensor X, indicated by code 800 in Figure 8, are written to the chip 600 (see code 600') under the "Layout Description" in Figure 6(b), and how the value of index (91,36) is read.

[0094] As shown in Figure 10, the element value reading unit 252 identifies the vertical block of Level A based on the quotient obtained by dividing the value for identifying the vertical address ("91") by the number of vertical elements per block of Level A ("50").

[0095] In the example in Figure 10, since the quotient value is "1", the element value reading unit 252 identifies the vertical block of Level A as the first block (block name "A2" or "A3").

[0096] Next, the element value reading unit 252 identifies the vertical block of Level B based on the quotient obtained by dividing the remainder value ("41") by the number of vertical elements per block of Level B ("25").

[0097] In the example in Figure 10, since the quotient value is "1", the element value reading unit 252 identifies the vertical block of Level B as the first block (block name "B2" or "B3").

[0098] Next, the element value reading unit 252 determines from the remainder value ("16") that the vertical position of the tensor is the 16th row.

[0099] Similarly, the element value reading unit 252 identifies the horizontal block of Level A based on the quotient obtained by dividing the value for identifying the horizontal address ("36") by the number of horizontal elements per block of Level A ("50").

[0100] In the example in Figure 10, since the quotient value is "0", the element value reading unit 252 identifies the horizontal block of Level A as the 0th block (block name "A0" or "A2").

[0101] Next, the element value reading unit 252 identifies the horizontal block of Level B based on the quotient obtained by dividing the remainder value ("36") by the number of horizontal elements per block of Level B ("25").

[0102] In the example in Figure 10, since the quotient value is "1", the element value reading unit 252 identifies the horizontal block of Level B as the first block (block name "B1" or "B3").

[0103] Next, the element value reading unit 252 identifies the horizontal position of the tensor as the 11th column based on the remainder value ("11").

[0104] As a result, the element value reading unit 252, • The block at Level A has the block name "A2", • The block at Level B has the block name "B3", • The memory address is the 411th address in row 16 × 25 + column 11 (see symbol 1000). To identify that it is so.

[0105] As a result, the element value reading unit 252 can read the value written to the address identified based on the description of the index.

[0106] Thus, the index (91,36) is decomposed into ((1,1,16),(0,1,11)), and each of them, As a Level A block, 1 × stride (="2") + 0 × stride (="1") = 2, As a Level B block, 1 × stride (="2") + 1 × stride (="1") = 3, · As a memory address, 16 × stride (="25") + 11 × stride (="1") = 411, By performing this calculation, we can identify the Level A block as "A2", the Level B block as "B3", and the memory address as "address number 411".

[0107] As described above, the ((1,1,16),(0,1,11)) obtained by the element value reading unit 252 decomposing the index (91,36) is referred to in this embodiment, for example, as the "decomposed index". Also, as described above, the block name "A2", block name "B3", and memory address "address 411" identified by the element value reading unit 252 from the index (91,36) are referred to in this embodiment, for example, as the "hierarchical index".

[0108] Expressions such as "decomposed indexes" ((1,1,16),(0,1,11)) and "hierarchical indexes" ("A2", "B3", "address 411") may be used in the machine code generation process by the compilation unit 213. For example, they may be used as a method for identifying each element of a tensor when generating machine code that performs layout changes on the same tensor.

[0109] Furthermore, when the values ​​of each element of tensor X, indicated by the symbol 800 in Figure 8, are written to the chip 600, it is possible to reduce communication costs when performing operations on tensor X in units such as 3x3 or 5x5 matrices. This is because the number of times different blocks are crossed at Level A when performing operations in units such as 3x3 or 5x5 matrices can be reduced. Operations performed in units such as 3x3 or 5x5 matrices include, for example, convolution and pooling operations.

[0110] (2) Specific example 2 Figure 11 is the second figure showing a specific example of processing by the element value reading unit. The example in Figure 11 shows how the values ​​of each element of the tensor X, indicated by the symbol 900 in Figure 9, are written to the chip 700 (see symbol 700') and the value of index (91,36) is read.

[0111] As shown in Figure 11, the element value reading unit 252 identifies the vertical block of Level A based on the quotient obtained by dividing the value for identifying the vertical address ("91") by the number of vertical elements per block of Level A ("25").

[0112] In the example in Figure 11, since the quotient value is "3", the element value reading unit 252 identifies that the third vertical block of Level A is the third block (block name "A3").

[0113] Next, the element value reading unit 252 determines from the remainder value ("16") that the vertical position of the tensor is the 16th row.

[0114] Similarly, the element value reading unit 252 identifies the horizontal block of Level B based on the quotient obtained by dividing the value for identifying the horizontal address ("36") by the number of horizontal elements per block of Level B ("25").

[0115] In the example in Figure 11, since the quotient value is "1", the element value reading unit 252 identifies the horizontal block of Level B as the first block (block name "B1").

[0116] Next, the element value reading unit 252 identifies the horizontal position of the tensor as the 11th column based on the remainder value ("11").

[0117] As a result, the element value reading unit 252, • The block at Level A has the block name "A3", • The block at Level B has the block name "B1", • The memory address is the 411th address in row 16 × 25 + column 11 (see code 1100). To identify that it is so.

[0118] As a result, the element value reading unit 252 can read the value written to the address identified based on the description of the index.

[0119] In this way, the index (91,36) is decomposed into ((3,16),(1,11)), and each of them, As a Level A block, 3 × stride (="1") = 3, As a Level B block, 1 × stride (="1") = 1, · As a memory address, 16 × stride (="25") + 11 × stride (="1") = 411, By performing this calculation, we can identify the Level A block as "A3", the Level B block as "B1", and the memory address as "address number 411".

[0120] Furthermore, when the values ​​of each element of tensor X, indicated by the symbol 900 in Figure 9, are written to chip 700, communication costs can be reduced, for example, when calculating row-by-row statistics for tensor X. This is because there is no need to cross different blocks in Level A when calculating row-by-row statistics for tensor X.

[0121] <Data processing flow by the data processing system> Next, the data processing flow by the data processing system 100 will be explained. Here, the explanation will be divided into three parts: the source code generation process by the source code description unit 211 and the generation unit 212, the machine code generation process by the compilation unit 213, and the machine code execution process by the execution unit 220.

[0122] (1) Source code generation process First, the flow of the source code generation process by the source code description unit 211 and the generation unit 212 will be explained. Figure 12 is a flowchart showing the flow of the source code generation process. When the user starts the source code description unit 211, the source code generation process shown in Figure 12 begins.

[0123] In step S1201, the user begins writing source code. The source code writing unit 211 then accepts the source code written by the user.

[0124] In step S1202, the user determines whether or not they have made a description related to tensors. If they determine that they have made a description related to tensors (i.e., the answer in step S1202 is YES), the process proceeds to step S1203. As a result, the source code description unit 211 accepts the user's description related to tensors.

[0125] In step S1203, the user describes the layout and proceeds to step S1204. As a result, the source code description unit 211 accepts the user's description of the layout.

[0126] On the other hand, if it is determined in step S1202 that no description of tensors has been made (i.e., the answer in step S1202 is NO), the process proceeds directly to step S1204.

[0127] In step S1204, the user decides whether or not to finish writing the source code. If the user decides not to finish writing the source code in step S1204 (i.e., the answer is NO in step S1204), the user returns to step S1202 and continues writing the source code.

[0128] On the other hand, if it is determined in step S1204 that the source code writing is finished (if the answer is YES in step S1204), the process proceeds to step S1205.

[0129] In step S1205, the user activates the generation unit 212 and instructs it to generate the calculation graph, etc. The generation unit 212 then retrieves the source code from the source code description unit 211 and generates the calculation graph, etc. The generation unit 212 also notifies the compilation unit 213 of the generated calculation graph, etc.

[0130] (2) Machine code generation process Next, the flow of the machine code generation process by the compilation unit 213 will be explained. Figure 13 is a flowchart showing the flow of the machine code generation process. When the user starts the compilation unit 213 of the compilation device 120, the compilation unit 213 starts the machine code generation process shown in Figure 13.

[0131] In step S1301, the compilation unit 213 starts the compilation process based on the calculation graph, etc.

[0132] In step S1302, the compilation unit 213 determines whether or not there is a description regarding the layout. If it is determined in step S1302 that there is a description regarding the layout (i.e., the answer is YES in step S1302), the process proceeds to step S1303.

[0133] In step S1303, the compilation unit 213 assigns a memory address to each element of the tensor based on the layout description, and proceeds to step S1304.

[0134] On the other hand, if it is determined in step S1302 that there is no description regarding the layout (i.e., the answer is NO in step S1302), the process proceeds directly to step S1304.

[0135] In step S1304, it is determined whether the compilation process for the calculation graph, etc., has finished. If it is determined in step S1304 that the compilation process has not finished (i.e., the answer is NO in step S1304), the process returns to step S1302 and continues.

[0136] On the other hand, if it is determined in step S1304 that the compilation process for the computation graph, etc., has finished (if the answer in step S1304 is YES), the machine code generation process is terminated.

[0137] (3) Machine code execution process Next, the flow of machine code execution processing by the execution unit 220 will be explained. Figure 14 is a flowchart showing the flow of machine code execution processing. When the user specifies the data to be processed stored in the data storage unit 214 and inputs an execution instruction to the execution unit 220 of the server device 110, the execution unit 220 starts the machine code execution processing shown in Figure 14.

[0138] In step S1401, the execution unit 220 starts the calculation of the machine code.

[0139] In step S1402, the execution unit 220 writes the values ​​of each element of the tensor (the data to be processed stored in the data storage unit 214) to the allocated memory address.

[0140] In step S1403, the execution unit 220 sequentially executes various processes included in the machine code 1410. For example, the execution unit 220 performs padding processing according to the code indicating padding processing and updates the allocated memory with the values ​​of each element of the processed tensor. The execution unit 220 also performs broadcast processing according to the code indicating broadcast processing and updates the allocated memory with the values ​​of each element of the processed tensor.

[0141] Once all the various processes included in the machine code 1410 have been executed, or once a predetermined termination condition has been met, the execution unit 220 terminates the machine code execution process.

[0142] <Summary> As is clear from the above description, the compilation device 120 according to the first embodiment is • Generates machine code to be executed on an accelerator chip having multiple memory locations connected by a tree-structure topology and distributed across them. Based on the number of divisions and stride (vertical or horizontal) for each level of the tensor to be processed, addresses within multiple memories on the accelerator chip are assigned to each element of the tensor to be processed.

[0143] As a result, according to the first embodiment, the arrangement of each element of a tensor across multiple memories connected by a tree topology can be appropriately represented.

[0144] [Second Embodiment] In the first embodiment described above, the compilation device 120 was described as being located within the server device 110, but the compilation device 120 may be configured separately from the server device 110. Also, in the first embodiment described above, the compilation unit 213 was described as being implemented in the compilation device 120, but the compilation unit 213 may be implemented in, for example, a terminal device (not shown). Alternatively, the compilation unit 213 may be implemented in another external device other than a terminal (for example, another server device).

[0145] Furthermore, in the first embodiment described above, the source code description unit 211, generation unit 212, and compilation unit 213 were implemented in the compilation device 120. However, the source code description unit 211 may be implemented in a terminal device connected via a network to the server device 110 on which the compilation device 120 is located. Alternatively, the source code description unit 211 and the generation unit 212 may be implemented in a terminal device connected via a communication network 150 to the server device 110 on which the compilation device 120 is located.

[0146] Figure 15 is a second diagram showing an example of the functional configuration of each device in a data processing system. In the example in Figure 15, the source code description unit 211 and the generation unit 212 are implemented in the terminal device 1510, and the source code 230 is displayed on the display device 1520 connected to the terminal device 1510. In the example in Figure 15, the calculation graph and the like generated by the terminal device 1510 are sent to the compilation device 120.

[0147] Furthermore, although the computation graph was described in the first embodiment above as being generated by executing the source code 230 and converting it to the ONNX representation format, the method for generating the computation graph is not limited to this, and the computation graph may be generated by other methods.

[0148] Furthermore, in the first embodiment described above, the generation unit 212 generates a layout instruction based on a layout description entered by the user, and the compilation unit 213 assigns addresses to each element of the tensor according to the layout instruction. However, the method of assigning addresses is not limited to this, and for example, the compilation unit 213 may select a layout and assign addresses to each element of the tensor according to the selected layout.

[0149] Furthermore, in the first embodiment described above, for example, the chip 170_1 was described as having four third-level blocks in Level A and four second-level blocks in Level B (Figure 3). However, the number of blocks (memory) in each level and the number of levels (depth) are not limited to this and are arbitrary.

[0150] Furthermore, in the first embodiment described above, • Level A hierarchy: 3rd tier block, • Level B hierarchy: 2nd tier block, • Level C hierarchy: 1st level block, However, the definitions of each level are not limited to this; for example, • Level A hierarchy: Tips, • Level B hierarchy: 3rd tier block, • Level C hierarchy: 2nd level block, • Level D hierarchy: 1st level block, You could also do that, • Level A hierarchy: Chips and 3rd tier blocks, • Level B hierarchy: 2nd tier block, • Level C hierarchy: 1st level block, That is also acceptable.

[0151] Furthermore, when referring to "Level A hierarchy: chips and third-tier blocks," for example, suppose that one board has four chips mounted on it, and each chip has four third-tier blocks. In this case, the Level A hierarchy can be described as if there were 16 third-tier blocks, in terms of the layout.

[0152] Furthermore, the hierarchy to which the memory belongs is not limited to the lowest level, but may change to other levels. In addition, the first and second embodiments described above may be applied by defining hierarchies such as a structure that bundles the highest level of memory (e.g., a chip), a structure that bundles the chips (e.g., a node), and a structure that bundles the nodes.

[0153] [Other embodiments] Where the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used in this specification (including the claims), it includes any of a, b, c, ab, ac, bc, or abc. It also includes multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it includes adding other elements other than the enumerated elements (a, b, and c), such as abcd which has d.

[0154] Furthermore, where expressions such as "data as input / based on / according to / in accordance with data" (including similar expressions) are used in this specification (including the claims), unless otherwise specified, this includes cases where the various data themselves are used as input, or where the various data have been processed in some way (e.g., data with added noise, normalized data, intermediate representations of the various data, etc.) are used as input. Also, where it is stated that some result is obtained "based on / according to / in accordance with data", this includes cases where the result is obtained based solely on the data in question, as well as cases where the result is also influenced by other data, causes, conditions, and / or states other than the data in question. Also, where it is stated that "data is output", unless otherwise specified, this includes cases where the various data themselves are used as output, or where the various data have been processed in some way (e.g., data with added noise, normalized data, intermediate representations of the various data, etc.) are used as output.

[0155] Furthermore, where the terms “connected” and “coupled” are used in this specification (including the claims), they are intended to be non-restrictive terms that include any of the following: direct connection / coupling, indirect connection / coupling, electrical connection / coupling, communicative connection / coupling, operational connection / coupling, and physical connection / coupling. The terms should be interpreted appropriately in the context in which they are used, but any form of connection / coupling that is not intentionally or naturally excluded should be interpreted non-restrictively as being included in the terms.

[0156] Furthermore, in this specification (including the claims), when the expression "A configured to B" is used, it may include that the physical structure of element A has a configuration capable of performing operation B, and that the permanent or temporary setting / configuration of element A is configured to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and that it is configured to actually perform operation B by the setting of a permanent or temporary program (instruction). Also, if element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.

[0157] Furthermore, where terms meaning "comprising" or "having" are used in this specification (including the claims), they are intended to be open-ended terms, including cases where the object of such term contains or possesses something other than the object indicated by the object of the term. If the object of such terms meaning "comprising" or "having" is an expression that does not specify a quantity or suggests a singular number (an expression with the article a or an), such expression should be interpreted as not being limited to a specific number.

[0158] Furthermore, even if expressions such as "one or more" or "at least one" are used in some places within this specification (including the claims), and expressions that do not specify a quantity or suggest a singular number (expressions using the articles a or an) are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or suggest a singular number (expressions using the articles a or an) should not necessarily be interpreted as not being limited to a specific number.

[0159] Furthermore, if this specification describes that a particular configuration of a certain embodiment yields a specific advantage / result, it should be understood that, unless otherwise stated, the same advantage / result can also be obtained from one or more other embodiments having that configuration. However, it should be understood that the presence or absence of such advantage / result generally depends on various causes, conditions, and / or states, and that the configuration does not necessarily guarantee that the advantage / result will be obtained. The advantage / result is merely obtained by the configuration described in the embodiment when various causes, conditions, and / or states are met, and the advantage / result is not necessarily obtained in the invention claimed to define that configuration or a similar configuration.

[0160] In this specification (including the claims), when terms such as "optimize" are used, they include finding a global optimal value, finding an approximate value of the global optimal value, finding a local optimal value, and finding an approximate value of the local optimal value, and should be interpreted appropriately depending on the context in which the terms are used. Furthermore, this includes finding these approximate values ​​of optimal values ​​probabilistically or heuristically.

[0161] Furthermore, in this specification (including the claims), when multiple hardware components perform a predetermined process, each component may cooperate to perform the predetermined process, or some components may perform all of the predetermined process. Alternatively, some components may perform part of the predetermined process, while other components perform the remainder. In this specification (including the claims), when expressions such as "one or more hardware components perform a first process, and the one or more hardware components perform a second process" are used, the hardware component performing the first process and the hardware component performing the second process may be the same or different. In other words, it is sufficient that the hardware component performing the first process and the hardware component performing the second process are included in the one or more hardware components. Note that hardware may include electronic circuits or devices containing electronic circuits.

[0162] Furthermore, in this specification (including the claims), when multiple memory devices store data, each of the multiple memory devices may store only a portion of the data or the entire data.

[0163] While embodiments of this disclosure have been described in detail above, this disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, and partial deletions are possible, provided that they do not depart from the conceptual idea and spirit of the present invention derived from the claims and their equivalents. For example, where numerical values ​​or mathematical formulas are used in the description in all of the embodiments described above, they are provided as examples only and are not limited thereto. Also, the order of operations in the embodiments is provided as examples only and is not limited thereto.

Claims

1. A compilation device for generating machine code to be executed on a chip having at least a first hierarchical layer and a second hierarchical layer, comprising: the second hierarchy is higher than the first hierarchy, the first hierarchy has a plurality of first blocks, The compiling device includes: At least one memory; At least one processor; The at least one processor Obtaining a tensor to be processed on the chip; executes a process of associating each element of the tensor with any one of the first blocks included in the chip, based on at least a division number in the first hierarchy of the chip; generating the machine code to be executed on the chip based on the mapping process; the first layer used in the association process corresponds to a hardware configuration of the chip; Compilation device.

2. The at least one processor performs the matching process based on at least the number of divisions in the first hierarchical layer and the stride in the first hierarchical layer. The compiling device according to claim 1 .

3. Each of the plurality of first blocks includes at least one memory among a plurality of memories included in the chip; As the associating process, the at least one processor executes a process of associating each element of the tensor with an address of the plurality of memories included in the chip, based on at least the number of divisions in the first hierarchy. The compiling device according to claim 1 or 2.

4. The number of divisions in the first hierarchy includes at least the number of vertical divisions and the number of horizontal divisions in the first hierarchy, The compiling device according to any one of claims 1 to 3.

5. The stride in the first hierarchical layer includes at least a vertical stride and a horizontal stride in the first hierarchical layer. The compiling device according to claim 2 .

6. The second hierarchical layer has a plurality of second blocks, each of the plurality of second blocks having the plurality of first blocks; The association process includes: another process performed by the at least one processor to associate each element of the tensor with any one of the second blocks included in the chip, based on at least a division number in the second hierarchy of the chip; the second hierarchy used in the other process of associating corresponds to a hardware configuration of the chip; The compiling device according to any one of claims 1 to 5.

7. The at least one processor performs the other processing of associating based on at least the number of divisions in the second hierarchical layer and the stride in the second hierarchical layer. The compiling device according to claim 6.

8. The number of divisions in the first hierarchy and the number of divisions in the second hierarchy are different from each other. The compiling device according to claim 6 or 7.

9. The at least one processor further obtains a computation graph to be processed on the chip, and the tensor is a tensor used in the computation graph. The compiling device according to any one of claims 1 to 8.

10. The at least one processor further generates the computation graph based on a source code. The compiling device according to claim 9.

11. The number of divisions in the first hierarchy is described in the source code. The compiling device according to claim 10.

12. A method for generating a machine code using a compiling device according to any one of claims 1 to 11. Generation method.

13. A method for generating a code for a computer-readable medium comprising: causing at least one processor of a compiling device to execute the generating method according to claim 12; program.

14. A chip having at least a first hierarchical layer and a second hierarchical layer, the chip executing machine code generated using a compiling device according to any one of claims 1 to 11. Tips.

15. A compilation device having at least one memory and at least one processor; a chip having at least a first tier and a second tier, the second hierarchy is higher than the first hierarchy, and the first hierarchy has a plurality of first blocks; The at least one processor Obtaining a tensor to be processed on the chip; executes a process of associating each element of the tensor with any one of the first blocks included in the chip, based on at least a division number in the first hierarchy of the chip; generating machine code to be executed on the chip based on the mapping process; The chip comprises: By executing the machine code generated by the compiling device performing the corresponding process, at least one of a process of writing values ​​of each element of the tensor to the first block corresponding to each element of the tensor, and a process of reading values ​​of each element of the tensor from the first block corresponding to each element of the tensor, the first layer used in the association process corresponds to a hardware configuration of the chip; system.

16. When executing a process of writing values ​​of each element of the tensor to the first block corresponding to each element of the tensor, the chip executes a padding process to adjust the size according to a memory to which the tensor is written. The system of claim 15.

17. The chip further performs broadcast processing when performing an operation between tensors whose array shapes do not match.

17. A system according to claim 15 or claim 16.

18. The at least one processor performs the matching process based at least on the number of divisions in the first hierarchical layer and the stride in the first hierarchical layer. A system according to any one of claims 15 to 17.

19. The chip further comprises a plurality of memories, each of the plurality of first blocks including at least one memory of the plurality of memories; As the associating process, the at least one processor executes a process of associating each element of the tensor with an address of the plurality of memories included in the chip, based on at least the number of divisions in the first hierarchy. A system according to any one of claims 15 to 18.

20. The plurality of memories included in the chip are connected in a tree structure.

20. The system of claim 19.

21. The second tier of the chip has a plurality of second blocks, each of the second blocks having the plurality of first blocks; The association process includes: another process performed by the at least one processor to associate each element of the tensor with any one of the second blocks included in the chip, based on at least a division number in the second hierarchy of the chip; the second hierarchy used in the other process of associating corresponds to a hardware configuration of the chip; A system according to any one of claims 15 to 20.

22. Each of the plurality of first blocks includes at least one arithmetic unit.

22. A system according to any one of claims 15 to 21.

23. The chip operates according to a SIMD architecture.

23. A system according to any one of claims 15 to 22.

24. A system according to any one of claims 15 to 23, Tips.