Information processing device, information processing system, scheduling method, and scheduling program

By optimizing the scheduling of forward and backward calculations on multiple workers by subdividing backward calculations, the method addresses inefficiencies in training time, enhancing the efficiency of model training processes.

JP2025097587APending Publication Date: 2025-07-01PREFERRED NETWORKS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023213844
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing scheduling methods for training models on multiple workers in pipeline parallelization result in inefficiencies due to workers waiting for calculations to complete, leading to increased training time.

Method used

An information processing apparatus that identifies the execution order and timing of forward and backward calculations for each worker, subdividing backward calculations into data and weight calculations to optimize scheduling and minimize idle time.

Benefits of technology

The proposed method reduces the time when workers are idle and shortens the overall training time by efficiently assigning calculations to available time slots, thereby improving the training process efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025097587000001_ABST
    Figure 2025097587000001_ABST
Patent Text Reader

Abstract

To generate a schedule for making a plurality of workers efficiently execute processing.SOLUTION: In an information processing device for specifying information including an execution sequence when each worker executes a forward calculation of each batch to be used for training of a model, and scheduling execution timing of the forward calculation and backward calculation of each batch to be executed by each worker during training the model on the basis of the specified information so as to satisfy a prescribed constraint condition, at least one processor schedules execution timing of at least one worker so as to execute second calculation included in the backward calculation of the i-th batch after first calculation included in the backward calculation of the i-th batch and the first calculation included in the backward calculation of a batch after the i-th batch, and the second calculation includes calculation to be executed by using a result of the first calculation.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing system, a scheduling method, and a scheduling program.

Background Art

[0002] Various scheduling methods have been proposed for generating a schedule for efficiently executing training processing for a model to be trained in pipeline parallelization on a plurality of workers (for example, a plurality of server devices). By generating a schedule by this method and causing a plurality of workers to execute the training processing, the training time can be shortened.

[0003] On the other hand, when generating a schedule, it is necessary to determine the execution timing of each calculation by each worker so as to satisfy a predetermined constraint condition. For this reason, the generated schedule includes, for example, a time during which the second worker does not execute processing, such as a need to wait for the calculation by the second worker until the calculation by the first worker is completed.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The present disclosure generates a schedule for efficiently executing processing on a plurality of workers.

Means for Solving the Problems

[0006] An information processing apparatus according to an aspect of the present disclosure has, for example, the following configuration. That is, at least one memory, and at least one processor, wherein the at least one processor identifies information including an execution order when each worker executes forward calculation of each batch used for training of a model, and schedules execution timings of the forward calculation and backward calculation of each batch to be executed by each worker during training of the model based on the identified information so as to satisfy a predetermined constraint condition, the information processing apparatus being wherein the at least one processor schedules execution timings of at least one worker so that a second calculation included in the backward calculation of the i-th batch is executed after a first calculation included in the backward calculation of the i-th batch and the first calculation included in the backward calculation of a batch after the i-th batch, wherein the second calculation includes a calculation executed using a result of the first calculation.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8A

Figure 8B

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 12

Figure 13A

Figure 13B

Embodiments for Carrying Out the Invention

[0008] Hereinafter, each embodiment will be described with reference to the accompanying drawings. In the present specification and drawings, for components having substantially the same functional configuration, the same reference numerals are given and duplicate descriptions are omitted.

[0009] [First Embodiment] <System Configuration of Information Processing System> First, the system configuration of the information processing system according to the first embodiment will be described. FIG. 1 is a diagram showing an example of the system configuration of the information processing system. As shown in FIG. 1, the information processing system 100 according to the first embodiment includes a plurality of server devices (server device group 110) and an information processing device 120.

[0010] The server device group 110 executes a training process on a model to be trained (for example, a neural network). The training process by the server device group 110 is executed based on a pipeline-parallelized schedule generated by the information processing device 120.

[0011] The information processing device 120 generates a schedule for efficiently executing the training process on the model to be trained in a pipeline-parallelized manner and causing a plurality of workers to execute it. In this embodiment, a worker refers to a plurality of servers included in the server device group 110. That is, a plurality of servers are included in one worker.

[0012] However, the definition of a worker is not limited to this, and a worker may refer to one or a plurality of servers included in the server device group 110. That is, one worker = one or a plurality of servers may be used, or one worker = one or a plurality of information processing devices may be used. Using a more general expression, one worker may be one device or a plurality of device groups specified as the destination for schedule allocation.

[0013] Alternatively, a worker may refer to a plurality of accelerators included in one server. That is, one worker may include a plurality of accelerators. Alternatively, a worker may refer to one accelerator included in one server. That is, one worker = one accelerator may be possible. Here, an accelerator is taken as an example, but the accelerator may be read as a GPU (Graphics Processing Unit). Alternatively, the accelerator may be read as a processor. Using a more general expression, one worker may be one component or a plurality of component groups specified as a schedule assignment destination.

[0014] During the training process, the processes executed by each worker for each micro-batch of training data include forward calculation and backward calculation. Therefore, in the information processing apparatus 120, the timing at which each worker executes the forward calculation and backward calculation of each micro-batch is scheduled.

[0015] However, in the information processing system 100 according to the present embodiment, when causing each worker to execute backward calculation, it is executed by being divided into backward data calculation and backward weight calculation. Note that the backward data calculation refers to, for example, the part that calculates the gradient of activation among the backward calculations, and the backward weight calculation refers to, for example, the part that calculates the gradient of parameters among the backward calculations.

[0016] Therefore, in the information processing apparatus 120, the backward calculation of each micro-batch is divided into backward data calculation and backward weight calculation, and the timing of executing each calculation is scheduled respectively.

[0017] Specifically, the information processing apparatus 120 first receives the input of scheduling information. In the present embodiment, the scheduling information received by the information processing apparatus 120 includes, for example, · Configuration information indicating the configuration of the model to be trained, · The number of mini - batches of training data used in the training process, · The number of workers used in the training process, · The execution order of each mini - batch when executing the training process, · The memory capacity of each worker used in the training process, · The network bandwidth between each worker used in the training process, etc. are included.

[0018] Subsequently, the information processing device 120 generates a number of forward - calculation identifiers and backward - calculation identifiers corresponding to the number of mini - batches included in the above scheduling information. Further, the information processing device 120 divides the generated backward - calculation identifier into a backward - data - calculation identifier and a backward - weight - calculation identifier.

[0019] Subsequently, the information processing device 120 arranges the generated forward - calculation identifiers, backward - data - calculation identifiers, and backward - weight - calculation identifiers at positions indicating the execution timing of each worker based on the above scheduling information. Thereby, the information processing device 120 can schedule the execution timing of the forward calculation, backward data calculation, and backward weight calculation of each mini - batch. Note that the information processing device 120 schedules so as to satisfy the pre - stored constraint conditions (constraint conditions regarding the execution order of forward calculation, backward data calculation, and backward weight calculation).

[0020] The information processing device 120 transmits the generated schedule to the server device group 110. Thereby, the server device group 110 can execute the training process based on the schedule generated by the information processing device 120.

[0021] Note that in the server device group 110, as an example of the training process executed by each worker, for example, when the model to be trained is a neural network (NN: Neural Network), · Worker name = Worker0: The first layer of NN, · Worker name = Worker1: The second layer of NN, ··· Examples include cases where each worker executes training processing for the corresponding layer. Or, · Worker name = Worker0: From the first layer to the Nth layer of NN, · Worker name = Worker1: From the (N + 1)th layer to the 2Nth layer of NN, ··· Examples include cases where each worker executes training processing for the corresponding multiple layers. That is, cases where NN is divided as evenly as possible, and in order from the layer closer to the input, each worker is in charge of the training processing.

[0022] However, when the number of layers of NN is not divisible by the number of workers, there may be cases where the number of layers a part of the workers are in charge of when executing training processing is less than the number of layers the other workers are in charge of when executing training processing. Or, when special calculations are included in the layers around the input and the layers around the output, there may be cases where the computational load is unbalanced among the workers.

[0023] <Hardware Configuration of the Information Processing Device> Next, the hardware configuration of the information processing device 120 will be described. FIG. 2 is a diagram showing an example of the hardware configuration of the information processing device. The information processing device 120 includes, as components, a processor 201, a main storage device 202 (memory), an auxiliary storage device 203 (memory), a network interface 204, and a device interface 205. The information processing device 120 may be realized as a computer in which these components are connected via a bus 206. In the example of FIG. 2, the information processing device 120 is shown as having one of each component, but the information processing device 120 may have multiple of the same components.

[0024] The various operations of the information processing apparatus 120 may be executed in parallel using one or more processors. Further, the various operations may be distributed to a plurality of arithmetic cores within the processor 201 and executed in parallel. Also, part or all of the processing, means, etc. of the present disclosure may be executed by an external device 230 (at least one of a processor and a storage device) provided on a cloud that can communicate with the information processing apparatus 120 via the network interface 204.

[0025] The processor 201 may be an electronic circuit (processing circuit, Processing circuit, Processing circuitry, CPU, GPU, FPGA, or ASIC, etc.). Further, the processor 201 may be a semiconductor device including a dedicated processing circuit, etc. Note that the processor 201 is not limited to an electronic circuit using electronic logic elements, and may be realized by an optical circuit using optical logic elements. Also, the processor 201 may include an arithmetic function based on quantum computing.

[0026] The processor 201 performs various operations based on various data and instructions input from each device, etc. of the internal configuration of the information processing apparatus 120, and outputs the operation results and control signals to each device, etc. The processor 201 controls each component included in the information processing apparatus 120 by executing an OS (Operating System), an application, etc.

[0027] Also, the processor 201 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or devices. When using a plurality of electronic circuits, each electronic circuit may communicate wired or wirelessly.

[0028] The main memory device 202 is a storage device that stores instructions executed by the processor 201 and various data, etc. Various data stored in the main memory device 202 are read by the processor 201. The auxiliary storage device 203 is a storage device other than the main memory device 202. Note that these storage devices mean any electronic components capable of storing various data (for example, constraint conditions stored in the constraint condition storage unit 305 described later), and may be semiconductor memories. The semiconductor memory may be either a volatile memory or a non-volatile memory. The storage device for storing various data in the information processing device 120 may be realized by the main memory device 202 or the auxiliary storage device 203, or may be realized by a built-in memory built into the processor 201.

[0029] Also, a plurality of processors 201 may be connected (coupled) to one main memory device 202, or a single processor 201 may be connected. Alternatively, a plurality of main memory devices 202 may be connected (coupled) to one processor 201. When the information processing device 120 is composed of at least one main memory device 202 and a plurality of processors 201 connected (coupled) to this at least one main memory device 202, a configuration in which at least one of the plurality of processors 201 is connected (coupled) to at least one main memory device 202 may be included.

[0030] The network interface 204 is an interface for connecting to the communication network 220 wirelessly or by wire.

[0031] The device interface 205 is an interface such as USB that directly connects to the external device 240.

[0032] The external device 240 may be, for example, an input device. In the present embodiment, the input device is an electronic device such as a camera, a microphone, various sensors, a keyboard, a mouse, or a touch panel, and gives the acquired information to the information processing device 120.

[0033] Further, the external device 240 may be, for example, an output device. In this embodiment, the output device may be, for example, a display device such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), or an organic EL (Electro Luminescence) panel, or may be a speaker or the like that outputs sound or the like.

[0034] Further, the external device 240 may be a storage device (memory). For example, the external device 240 may be a network storage or the like, or the external device 240 may be a storage such as an HDD.

[0035] Further, the external device 240 may be a device having some functions of the components of the information processing device 120. That is, the information processing device 120 may transmit and receive processing results to and from the external device 240.

[0036] Here, the hardware configuration of the information processing device 120 has been described, and the hardware configuration of each of the plurality of server devices included in the server device group 110 has not been mentioned. However, at least one server device included in the server device group 110 may have the same hardware configuration as the information processing device 120.

[0037] <Functional Configuration of Information Processing Device> Next, the functional configuration of the information processing device 120 will be described. FIG. 3 is a diagram showing an example of the functional configuration of the information processing device. A scheduling program is installed in the information processing device 120, and when the program is executed, the information processing device 120 functions as a specifying unit 301, a dividing unit 302, a scheduling unit 303, and a transmitting unit 304.

[0038] The specific part 301 receives the input of information for scheduling. Since the details of the information for scheduling received by the specific part 301 have been described with reference to FIG. 1, the description is omitted here. The specific part 301 notifies the scheduling part 303 of the received input information for scheduling. Further, the specific part 301 generates a number of forward calculation identifiers and backward calculation identifiers corresponding to the "number of micro-batches" included in the received input information for scheduling. Also, the specific part 301 notifies the generated forward calculation identifiers and backward calculation identifiers to the splitting part 302.

[0039] The splitting part 302 splits the backward calculation identifiers among the number of forward calculation identifiers and backward calculation identifiers corresponding to the "number of micro-batches" notified from the specific part 301 into backward data calculation identifiers and backward weight calculation identifiers. The splitting part 302 notifies the number of forward calculation identifiers corresponding to the number of micro-batches, and the number of backward data calculation identifiers and backward weight calculation identifiers corresponding to the number of micro-batches to the scheduling part 303.

[0040] The scheduling part 303 acquires the information for scheduling notified from the specific part 301, and the forward calculation identifiers, backward data calculation identifiers, and backward weight calculation identifiers notified from the splitting part 302. Also, the scheduling part 303 is based on the information for scheduling and the constraint conditions read from the constraint condition storage part 305, · The forward calculation identifiers of each micro-batch, · The backward data calculation identifiers of each micro-batch, · The backward weight calculation identifiers of each micro-batch, By arranging them at positions indicating the execution timing of each worker, the execution timing of the forward calculation, backward data calculation, and backward weight calculation of each micro-batch is scheduled.

[0041] The transmission unit 304 transmits the schedule generated by the scheduling unit 303 to the server device group 110.

[0042] <An example of the constraint conditions> Next, the details of the constraint conditions stored in the constraint condition storage unit 305 will be described. FIG. 4 is a diagram showing an example of the constraint conditions in the first embodiment. As shown in FIG. 4, the information processing apparatus 120 according to the first embodiment schedules so as to satisfy the constraint conditions (1) to (4). (1) The forward calculation of each micro-batch is executed in a specific execution order among the workers. (2) Each worker executes the backward data calculation of each micro-batch after the forward calculation of each micro-batch. (3) The backward data calculation of each micro-batch is executed in an order opposite to the above specific execution order among the workers. (4) Each worker executes the backward weight calculation of each micro-batch after the backward calculation of the micro-batch.

[0043] The scheduling unit 303, for example, arranges each calculation identifier notified from the division unit 302 at a position indicating the execution timing of each worker so as to satisfy the above constraint conditions, and, for example, searches for an arrangement that minimizes the training time. Note that the scheduling unit 303 may search for an arrangement that minimizes the training time by solving an optimization problem.

[0044] <The schedule generated by the information processing apparatus of the comparative example> Next, before explaining the schedule generated by the information processing apparatus 120 according to the present embodiment, the schedule generated by the information processing apparatus of the comparative example will be explained. The information processing apparatus of the comparative example is an apparatus that performs scheduling without dividing backward calculation into backward data calculation and backward weight calculation. Here, for simplicity of explanation, among the information for scheduling, only the number of micro-batches, the number of workers, and the execution order of the micro-batches will be used for scheduling, and the case where the following information is input will be explained respectively. · Number of micro-batches: 4, · Number of workers: 4 (worker names = Worker0, Worker1, Worker2, Worker3), · Execution order of micro-batches: in the order of Worker0 → Worker1 → Worker2 → Worker3.

[0045] FIG. 5 is a diagram showing a specific example of the schedule generated by the information processing apparatus of the comparative example. In FIG. 5, the horizontal axis represents time, and the vertical axis represents the worker names of each worker. The calculation identifiers are arranged in the area where each time intersects with the worker name of each worker.

[0046] Among the calculation identifiers arranged in the schedule of FIG. 5, · Fwd0 to Fwd3 represent the forward calculation identifiers of micro-batch 0 to the forward calculation identifiers of micro-batch 3, · Bwd0 to Bwd3 represent the backward calculation identifiers of micro-batch 0 to the backward calculation identifiers of micro-batch 3, respectively. Note that the length of each calculation identifier in the time axis direction represents the time required for each calculation. Also, the time when no calculation identifier is arranged represents the time when the worker is not executing the process. The micro-batch i (where i is any of 0 to 3) represents the micro-batch in which the forward calculation is executed at the i-th position.

[0047] Assume that the information for scheduling is input and among the constraint conditions shown in FIG. 4, (1), (2), and (3) are applied (however, for (2) and (3), assume that the backward data calculation is read as backward calculation and applied). In this case, the information processing apparatus of the comparative example generates a schedule as shown in FIG. 5, for example.

[0048] Note that in the example of FIG. 5, the information processing apparatus of the comparative example schedules such that the backward calculation for all micro-batches 0 to 3 is executed after the forward calculation for all micro-batches 0 to 3 so as to satisfy the constraint condition (2).

[0049] According to the example of FIG. 5, when the training process is executed based on the schedule generated by the information processing apparatus of the comparative example, the training time = Ta.

[0050] FIG. 6 is a diagram showing another specific example of the schedule generated by the information processing apparatus of the comparative example. In FIG. 6, the horizontal axis and the vertical axis are the same as those in FIG. 5, and the calculation identifier is also the same as that in FIG. 5. Further, the input information for scheduling is also the same as that in FIG. 5.

[0051] The difference from FIG. 5 is that in the case of FIG. 6, · When scheduling to satisfy the constraint condition (2), the backward calculation for each micro-batch in the worker with worker name = Worker3 is scheduled to be executed immediately after the forward calculation for each micro-batch. Specifically, in the worker with worker name = Worker3, · The backward calculation of micro-batch 0 is executed immediately after the forward calculation of micro-batch 0, · The backward calculation of micro-batch 1 is executed after the forward calculation of micro-batch 1, · The backward calculation of micro-batch 2 is executed after the forward calculation of micro-batch 2, · The backward calculation of micro-batch 3 is executed after the forward calculation of micro-batch 3. in terms of scheduling in this way.

[0052] According to the example of FIG. 6, when the training process is executed based on the schedule generated by the information processing apparatus of the comparative example, the training time = Tb (Ta = Tb).

[0053] Thus, even when the same scheduling information is input and the same constraint conditions are applied, the generated schedule is not necessarily determined uniquely, and there may be cases where a plurality of schedules are generated. In such cases, finally, the schedule with the minimum training time is selected. However, in the examples of FIGS. 5 and 6, the training time is the same for any schedule, and any schedule includes a lot of time when each worker is not executing the process.

[0054] On the other hand, in the information processing apparatus 120 according to the present embodiment, the backward calculation of each micro-batch is subdivided by dividing it into backward data calculation and backward weight calculation. Then, the information processing apparatus 120 according to the present embodiment schedules the subdivided calculations so that they are assigned to the idle time when each worker is not executing the process. Thereby, according to the information processing apparatus 120 according to the present embodiment, it is possible to reduce the time when the worker is not executing the process, and the training time can be shortened. Hereinafter, the details of the information processing apparatus 120 according to the present embodiment will be described.

[0055] <Division processing by the division unit> First, the division processing by the division unit 302 of the information processing apparatus 120 according to the present embodiment will be described. FIG. 7 is a first diagram showing a specific example of the division processing by the division unit.

[0056] As shown in FIG. 7, the division unit 302 · notifies the scheduling unit 303 of the forward calculation identifiers (Fwd0 to Fwd3) of micro-batches 0 to 3 notified from the specifying unit 301. ·Divide each backward calculation identifier (Bwd0 to Bwd3) of micro-batches 0 to 3 notified from the specific unit 301 into each backward data calculation identifier (BD0 to BD3) and each backward weight calculation identifier (BW0 to BW3). Also, notify each divided backward data calculation identifier (BD0 to BD3) and each backward weight calculation identifier (BW0 to BW3) to the scheduling unit 303.

[0057] Note that in the example of FIG. 7, the case where one backward calculation identifier is divided into one backward data calculation identifier and one backward weight calculation identifier is shown, but the number of divisions by the division unit 302 is not limited to this.

[0058] For example, the division unit 302 may divide one backward calculation identifier into one backward data calculation identifier and a plurality of backward weight calculation identifiers.

[0059] Also, in the example of FIG. 7, the case where the backward calculation identifier is divided into the backward data calculation identifier and the backward weight calculation identifier is shown, but the division method by the division unit 302 is not limited to this.

[0060] For example, the division unit 302 may divide it into a first calculation identifier indicating a calculation to be executed earlier than other calculations and a second calculation identifier indicating a calculation to be executed later than the first calculation among a plurality of calculations included in the backward calculation.

[0061] Specifically, when the model to be trained is a Transformer, it is conceivable to classify the backward weight calculation of the normalization process (such as Layer Normalization or RMS Normalization) into the backward data calculation. This is because the memory usage can be reduced by executing the backward weight calculation of the normalization process first.

[0062] <Scheduling Process by Scheduling Unit> Next, the scheduling process by the scheduling unit 303 of the information processing apparatus 120 according to the present embodiment will be described. FIGS. 8A and 8B are first and second flowcharts showing the flow of the scheduling process by the scheduling unit of the information processing apparatus according to the first embodiment.

[0063] In step S801 of FIG. 8A, the scheduling unit 303 schedules the timing at which each worker executes the forward calculation of micro-batch 0.

[0064] According to the constraint condition (1), the forward calculation of the micro-batch is defined to be executed in a specific execution order. Also, according to the scheduling information, the specific execution order of the forward calculation of the micro-batch is "Worker0→Worker1→Worker2→Worker3". Therefore, the scheduling unit 303 arranges the forward calculation identifier of micro-batch 0 as shown by reference numeral 811.

[0065] In step S802, the scheduling unit 303 schedules the timing at which all calculations of all micro-batches are executed in the worker (worker name = Worker3) that performs processing close to the output of the model to be trained.

[0066] Specifically, the scheduling unit 303 arranges the feedback calculation identifiers, backward data calculation identifiers, and backward weight calculation identifiers of micro-batches 0 to 3 so as to satisfy the constraint conditions (2) and (4) (see reference numeral 812).

[0067] In the case of the example of reference numeral 812, the scheduling unit 303 arranges the backward data calculation identifiers (BD0 to BD3) to the immediate right of the corresponding forward calculation identifiers (Fwd0 to Fwd3). That is, the scheduling unit 303 schedules the backward data calculation to be executed immediately after the corresponding forward calculation.

[0068] Also, in the case of the example of reference numeral 812, the scheduling unit 303 arranges the backward weight calculation identifiers (BW0 to BW3) to the right of all the backward data calculation identifiers (BD0 to BD3). That is, the scheduling unit 303 schedules the backward weight calculation to be executed after all the backward data calculations are executed.

[0069] Scheduling in this way in step S802 is to reduce the time during which the processing is not executed by the workers with worker names = Worker0 to Worker2 in the subsequent scheduling.

[0070] In step S803, the scheduling unit 303 schedules the timing at which each of the workers with worker names = Worker2 and Worker1 executes the backward data calculation for micro-batches 0 to 3.

[0071] According to the constraint condition (3), it is stipulated that the backward data calculation of the micro-batch is executed in the order opposite to the specific execution order. Also, according to the information for scheduling, the order opposite to the specific execution order of the forward calculation of the micro-batch is "Worker3 → Worker2 → Worker1 → Worker0". Therefore, the scheduling unit 303 arranges the backward data calculation identifiers (BD0 to BD3) for micro-batches 0 to 3 as shown by reference numeral 813.

[0072] In step S804 of FIG. 8B, the scheduling unit 303 schedules the timing at which the forward calculations for micro-batches 1 to 3 are executed so as to be allocated to the idle time of each of the workers with worker names = Worker0 to Worker2. Specifically, the scheduling unit 303 schedules in the same manner as the forward calculation of micro-batch 0, and arranges the forward calculation identifiers for micro-batches 1 to 3 as shown by reference numeral 814.

[0073] In step S805, the scheduling unit 303 schedules the timing to execute the backward weight calculation of micro-batches 0 to 3 so that it is assigned to the idle time of each worker with worker names = Worker1 and Worker2.

[0074] According to constraint condition (4), it is stipulated that the backward weight calculation of the micro-batch should be executed after the backward data calculation. Therefore, the scheduling unit 303 can also arrange the backward weight calculation identifiers (BW0 to BW3) on the right side of the backward data calculation identifier (BD3). That is, the scheduling unit 303 can also schedule to execute the backward weight calculation after all the backward data calculations are executed.

[0075] However, the scheduling unit 303 schedules so that the backward weight calculation of the i-th micro-batch is executed during the idle time between the backward data calculation of the i-th micro-batch and the backward data calculation of the i + 1-th micro-batch. This is to reduce the time when the workers with worker names = Worker1 and Worker2 are not performing processing (see reference numeral 815).

[0076] In step S806, the scheduling unit 303 schedules the timing for the worker with worker name = Worker0 to execute the backward data calculation and backward weight calculation of micro-batches 0 to 3.

[0077] According to constraint condition (3), it is stipulated that the backward data calculation of micro-batch 0 by the worker with worker name = Worker0 should be executed after the backward data calculation of micro-batch 0 by the worker with worker name = Worker1.

[0078] Similarly, according to constraint condition (3), the backward data calculation of micro-batch 1 by the worker with worker name = Worker0 is stipulated to be executed after the backward data calculation of micro-batch 1 by the worker with worker name = Worker1.

[0079] Similarly, according to constraint condition (3), the backward data calculation of micro-batch 2 by the worker with worker name = Worker0 is stipulated to be executed after the backward data calculation of micro-batch 2 by the worker with worker name = Worker1.

[0080] Similarly, according to constraint condition (3), the backward data calculation of micro-batch 3 by the worker with worker name = Worker0 is stipulated to be executed after the backward data calculation of micro-batch 3 by the worker with worker name = Worker1.

[0081] Therefore, the scheduling unit 303 arranges the backward data calculation identifiers of micro-batches 0 to 3 as shown by reference numeral 816.

[0082] Furthermore, according to constraint condition (4), it is stipulated that the backward weight calculation of the micro-batch be executed after the backward data calculation of the micro-batch. Therefore, the scheduling unit 303 arranges the backward weight calculation identifiers of micro-batches 0 to 3 as shown by reference numeral 816.

[0083] The reason for arranging the backward weight calculation identifiers of micro-batches 0 to 3 as shown by reference numeral 816 is the same as the reason explained in step S805 above.

[0084] <Schedule generated by the information processing apparatus according to the present embodiment> Next, the schedule generated by the scheduling unit 303 of the information processing apparatus 120 according to the present embodiment will be described. FIG. 9 is a diagram showing an example of a schedule generated by the scheduling unit of the information processing apparatus according to the first embodiment.

[0085] According to the example of FIG. 9, when the training process is executed based on the schedule generated by the scheduling unit of the information processing apparatus according to the first embodiment, the training time = Tc. As is clear from comparing FIGS. 5 and 6 with FIG. 9, in the schedule shown in FIG. 9, the time when each worker is not executing the process is less than that in the schedules shown in FIGS. 5 and 6. Therefore, when the training process is executed based on the schedule shown in FIG. 9, the training time can be shortened compared to the case where the training process is executed based on the schedules shown in FIGS. 5 and 6 (Ta, Tb > Tc).

[0086] <Summary> As is clear from the above description, the information processing apparatus 120 according to the first embodiment · As information for scheduling, identify the number of workers, the execution order when each worker executes the forward calculation of each micro-batch, and the number of micro-batches. · Based on the identified information, schedule the execution timings of the forward calculation and the backward calculation of each micro-batch executed by each worker so as to satisfy a predetermined constraint condition. · When scheduling the execution timing of the backward calculation, divide the backward calculation into a backward data calculation and a backward weight calculation, and schedule the execution timings of each calculation executed by each worker respectively.

[0087] In this way, the information processing apparatus 120 according to the first embodiment subdivides the backward calculation of each micro-batch so that the subdivided backward calculation can be assigned to the idle time when each worker is not executing the process during scheduling.

[0088] Accordingly, according to the information processing apparatus 120 according to the first embodiment, it is possible to generate a schedule for efficiently executing processing by a plurality of workers.

[0089] As a result, according to the information processing apparatus 120 according to the first embodiment, in the scheduling when the training process is pipelined in parallel, it is possible to reduce the time when the worker is not executing the process, and the training time can be shortened.

[0090] [Second Embodiment] In the first embodiment described above, the case of performing scheduling so as to satisfy the constraint conditions (1) to (4) has been described. However, the constraint conditions are not limited to (1) to (4), and another constraint condition (5) may be added. In the second embodiment, a case will be described in which, even when the constraint condition (5) is added, by subdividing the backward calculation of each micro-batch, it is possible to generate a schedule that satisfies the constraint conditions (1) to (5). The description will be centered on the differences from the first embodiment described above.

[0091] <An Example of Constraint Conditions> First, the constraint conditions used in the information processing apparatus 120 according to the second embodiment will be described. FIG. 10 is a diagram showing an example of the constraint conditions in the second embodiment. The difference from the constraint conditions in the first embodiment shown in FIG. 4 is that a new (5) is added to the constraint conditions shown in FIG. 10.

[0092] As shown in FIG. 10, in the constraint condition (5), it is stipulated that each worker executes the forward calculation and backward data calculation of each micro-batch so that the memory usage amount at each time does not exceed the memory capacity.

[0093] Therefore, when scheduling in the scheduling unit 303 of the information processing apparatus 120 according to the second embodiment, it is determined whether the memory usage at each time exceeds the memory capacity. When it exceeds the memory capacity, the scheduling unit 303 changes the execution timing of the backward weight calculation. Thus, according to the scheduling unit 303 of the information processing apparatus 120 according to the second embodiment, a schedule that satisfies the constraint conditions (1) to (5) can be generated.

[0094] <Scheduling process by the scheduling unit> Next, the scheduling process by the scheduling unit 303 of the information processing apparatus 120 according to the second embodiment will be described. FIGS. 11A and 11B are the first and second flowcharts showing the flow of the scheduling process by the scheduling unit of the information processing apparatus according to the second embodiment.

[0095] The difference from the flowchart described with reference to FIGS. 8A and 8B is that it includes steps S1101, S1102, and S1103. Therefore, here, the processing of steps S1101, S1102, and S1103 will be mainly described.

[0096] In step S1101 of FIG. 11A, the scheduling unit 303 determines whether the constraint condition (5) is satisfied for each worker (worker name = Worker3) for which the arrangement of each calculation identifier of each microbatch is completed.

[0097] Specifically, the scheduling unit 303 determines whether the memory usage at each time exceeds the memory capacity. In the case of the arrangement indicated by reference numeral 812, it is shown that the memory usage exceeds the memory capacity by executing the forward calculation of microbatch 2 (refer to the forward calculation identifier (Fwd2)) (arrow). Here, if the amount of data to be stored in the memory for the backward data calculation of each microbatch is D and the amount of data to be stored in the memory for the backward weight calculation of each microbatch is E, · Each time the forward calculation of each micro-batch is completed, the memory usage increases by only D, · Each time the backward data calculation of each micro-batch is completed, the memory usage increases by E - D, · Each time the backward weight calculation of each micro-batch is completed, the memory usage decreases by only E.

[0098] Therefore, at the timing when the forward calculation of micro-batch 2 (see forward calculation identifier (Fwd2)) is completed, the worker with worker name = Worker3, · The forward calculation of each of micro-batches 0 to 2 (3 forward calculations) and, · The backward data calculation of each of micro-batches 0 and 1 (2 backward data calculations) and, have been completed. Therefore, the memory usage by the worker with worker name = Worker3 is, Memory usage = 3×D + 2×(E - D) = D + 2E becomes.

[0099] Therefore, the scheduling unit 303 changes the arrangement of the backward weight calculation identifier of micro-batch 0 to before the forward calculation identifier of micro-batch 2. By executing the backward weight calculation of micro-batch 0 (see backward weight calculation identifier (BW0)), the memory usage decreases by only E. Therefore, the memory usage by the worker with worker name = Worker3 at the timing when the forward calculation of micro-batch 2 (see forward calculation identifier (Fwd2)) is completed is, Memory usage = D + 2E - E = D + E becomes.

[0100] In this way, the scheduling unit 303 changes the execution timing of the backward weight calculation for micro-batch 0 (see the backward weight calculation identifier (BW0)) by taking advantage of the fact that the backward calculation is subdivided. That is, the backward weight calculation (see the backward weight calculation identifier (BW0)) is scheduled to be executed before the timing when the memory usage exceeds the memory capacity. As a result, it is possible to avoid a situation where the memory usage exceeds the memory capacity at the timing when the forward calculation for micro-batch 2 (see the forward calculation identifier (Fwd2)) is completed, and a schedule that satisfies the constraint condition (5) can be generated.

[0101] However, afterwards, according to the arrangement of the code 812, · the backward data calculation for micro-batch 2 (see the backward data calculation identifier (BD2)), · the forward calculation (see the forward calculation identifier (Fwd3)) and the backward data calculation (see the backward data calculation identifier (BD3)) for micro-batch 3, are executed. For this reason, at the timing when the backward data calculation for micro-batch 3 (see the backward data calculation identifier (BD3)) is completed, again, the memory usage exceeds the memory capacity. Note that at the timing when the backward data calculation for micro-batch 3 (see the backward data calculation identifier (BD3)) is completed, the worker with the worker name = Worker3 further · the backward data calculation for micro-batches 2 and 3 (two backward calculations), and · the forward calculation for micro-batch 3 (one forward calculation), are completed. Therefore, the memory usage by the worker with the worker name = Worker3 is Memory usage = D + E + 2×(E - D) + D = 3×E becomes.

[0102] Therefore, as shown by the code 1111, the scheduling unit 303 changes the arrangement of the backward weight calculation identifier of micro-batch 1 to before the forward calculation identifier of micro-batch 3. By executing the backward weight calculation of micro-batch 1 (see the backward weight calculation identifier (BW1)), the memory usage is reduced by only E. Therefore, at the timing when the backward data calculation of micro-batch 3 (see the backward data calculation identifier (BD3)) is completed, the memory usage by the worker with the worker name = Worker3 is Memory usage = 3×E - E = 2×E becomes.

[0103] In this way, the scheduling unit 303 changes the execution timing of the backward weight calculation of micro-batch 1 (see the backward weight calculation identifier (BW1)) by taking advantage of the fact that the backward calculation is subdivided. That is, the backward weight calculation (see the backward weight calculation identifier (BW1)) is scheduled to be executed before the timing when the memory usage exceeds the memory capacity. Thereby, it is possible to avoid a situation where the memory usage exceeds the memory capacity at the timing when the backward data calculation of micro-batch 3 (see the backward data calculation identifier (BD3)) is completed, and a schedule that satisfies the constraint condition (5) can be generated.

[0104] In FIG. 11A, the code 1111 shows how the scheduling unit 303 generates a schedule so as to satisfy the constraint condition (5) for the worker with the worker name = Worker3 in step S1101.

[0105] Subsequently, in step S803, the scheduling unit 303 schedules the timing at which each worker with the worker name = Worker2 and Worker1 executes the backward data calculation of micro-batches 0 to 3.

[0106] Specifically, the scheduling unit 303 arranges the backward calculation identifiers of micro-batches 0 to 3 so as to satisfy the constraint condition (3).

[0107] In the process of step S1101, · Backward weight calculation of micro-batch 0 (refer to the backward weight calculation identifier (BW0)), and · Backward weight calculation of micro-batch 1 (refer to the backward weight calculation identifier (BW1)), the timing of execution is changed. Along with this, the timing of the worker with worker name = Worker3 to execute backward data calculation (refer to the backward data calculation identifiers (BD2, BD3)) is also changed.

[0108] Therefore, in the scheduling unit 303, the backward data calculation identifiers of micro-batch 2 and the backward data calculation identifiers of micro-batch 3 are arranged at positions different from the reference numeral 813 in FIG. 8A.

[0109] As a result, in the scheduling unit 303, in step S803, a schedule indicated by reference numeral 1112 is generated as a schedule different from the schedule indicated by reference numeral 803 in FIG. 8A.

[0110] In steps S804 and S805 of FIG. 11B, the same processes as steps S804 and S805 of FIG. 8B are executed. However, the schedule (reference numeral 1112) generated in step S803 is different from the schedule (reference numeral 813) generated in step S803 of FIG. 8A. Therefore, the schedules (reference numerals 1113, 1114) generated in steps S804 and S805 of FIG. 11B are also different from the schedules (reference numerals 814, 815) generated in steps S804 and S805 of FIG. 8B.

[0111] In step S1102, the scheduling unit 303 determines whether the constraint condition (5) is satisfied for the workers with worker names = Worker1 and 2 for which the arrangement of the backward data calculation identifier and the backward weight calculation identifier of each micro-batch is completed.

[0112] Specifically, the scheduling unit 303 determines whether the memory usage at each time exceeds the memory capacity. In the example of FIG. 11B, it is assumed that in the case of the arrangement indicated by reference numeral 1114, it is determined that the memory usage does not exceed the memory capacity at any time. In this case, since it is not necessary to change the arrangement of the backward weight calculation identifier, the scheduling unit 303 proceeds to step S806 without changing the schedule generated in step 805.

[0113] In step S806 of FIG. 11B, the same processing as in step S806 of FIG. 8B is executed. However, in the case of FIG. 11B, the schedule (reference numeral 1114) generated in step S805 is different from the schedule (reference numeral 815) generated in step S805 of FIG. 8B. For this reason, the schedule generated in step S806 of FIG. 11B is also different from the schedule (reference numeral 816) generated in step S806 of FIG. 8B (see reference numeral 1115).

[0114] In step S1103, the scheduling unit 303 determines whether the constraint condition (5) is satisfied for the worker with worker name = Worker0 for which the arrangement of the backward data calculation identifier and the backward weight calculation identifier of each micro-batch is completed.

[0115] Specifically, the scheduling unit 303 determines whether the memory usage at each time exceeds the memory capacity. In the example of FIG. 11B, in the case of the arrangement indicated by reference numeral 1115, it is assumed that at any time, it is determined that the memory usage does not exceed the memory capacity. In this case, since there is no need to change the arrangement of the backward weight calculation identifier, the scheduling unit 303 outputs the schedule generated in step 806 without changing it, and ends the scheduling process.

[0116] <Summary> As is clear from the above description, the information processing apparatus 120 according to the second embodiment · When scheduling by the scheduling unit, along with the addition of constraint condition (5), for each worker, it is determined whether the memory usage at each time exceeds the memory capacity. · When the memory usage exceeds the memory capacity, the memory usage is reduced by changing the timing for executing the backward weight calculation.

[0117] In this way, the information processing apparatus 120 according to the second embodiment utilizes the fact that the backward calculation of each microbatch is subdivided, and changes the execution timing of the backward weight calculation according to constraint condition (5).

[0118] Thereby, according to the information processing apparatus 120 according to the second embodiment, a situation where the memory usage at each time exceeds the memory capacity can be avoided, and a schedule that satisfies the constraint conditions can be generated.

[0119] [Third Embodiment] In the second embodiment described above, the case where constraint condition (5) is added to constraint conditions (1) to (4) was explained. In contrast, in the third embodiment, the case where constraint condition (6) is added to constraint conditions (1) to (4) will be explained. Similar to the second embodiment, also in the third embodiment, by subdividing the backward calculation of each microbatch, the case where it becomes possible to generate a schedule that satisfies constraint conditions (1) to (4) and (6) will be explained. However, the explanation will focus on the differences from the first or second embodiment described above.

[0120] <An example of constraint conditions> First, the constraint conditions used in the information processing apparatus 120 according to the third embodiment will be explained. FIG. 12 is a diagram showing an example of the constraint conditions in the third embodiment. The difference from the constraint conditions in the first embodiment shown in FIG. 4 is that in the constraint conditions shown in FIG. 12, (6) is newly added.

[0121] As shown in FIG. 12, in constraint condition (6), it is stipulated that the processing involving communication within the worker is executed after all the backward data calculations of the microbatches are executed.

[0122] Therefore, in the scheduling unit 303 of the information processing apparatus 120 according to the third embodiment, when scheduling, the timing for executing the processing involving communication within the worker is determined. Thus, according to the scheduling unit 303 of the information processing apparatus 120 according to the third embodiment, a schedule that satisfies constraint conditions (1) to (4) and (6) can be generated.

[0123] <Scheduling process by the scheduling unit> Next, the scheduling process by the scheduling unit 303 of the information processing apparatus 120 according to the third embodiment will be explained. FIGS. 13A and 13B are the first and second flowcharts showing the flow of the scheduling process by the scheduling unit of the information processing apparatus according to the third embodiment.

[0124] The differences from the flowchart described with reference to FIGS. 8A and 8B are steps S1301, S1302, and S1303. Therefore, here, the processing in steps S1301, S1302, and S1303 will be mainly described.

[0125] In step S1301 of FIG. 13A, the scheduling unit 303 identifies the timing that satisfies the constraint condition (6) for the worker (worker name = Worker3) in which the arrangement of each calculation identifier of each micro-batch is completed.

[0126] Specifically, the scheduling unit 303 identifies the timing at which the backward data calculation (see backward data calculation identifier (BD3)) of micro-batch 3 by the worker with the worker name = Worker3 is completed. Thereby, the scheduling unit 303 can schedule to execute the process involving communication within the worker at the identified timing.

[0127] As a result, it becomes possible to execute the process involving communication within the worker in parallel with the backward weight calculation (see backward weight calculation identifiers (BW0 to BW3)) of micro-batches 0 to 3 by the worker with the worker name = Worker3. That is, the scheduling unit 303 can generate a schedule that satisfies the constraint condition (6). In reference numeral 813 in FIG. 13A, the arrow indicated by "parallel calculation start" indicates the start timing at which the process involving communication within the worker with the worker name = Worker3 can be executed in parallel with the backward weight calculation. Note that the process involving communication within the worker includes, for example, · A process of calculating the sum of gradients of model parameters. Specifically, A process of calculating the sum of gradients of model parameters by transmitting and receiving (and adding) the results of all backward data calculations within the worker, or A process of calculating the sum of gradients of model parameters on the receiving side by collecting the results of all backward calculations within the worker at one location and then transmitting them, · A process of updating model parameters, · A process of distributing updated model parameters, · A process of distributing the summed gradients, includes at least any one of the above processes.

[0128] In step S1302 of FIG. 13B, the scheduling unit 303 identifies the timing that satisfies the constraint condition (6) for the workers with worker names = Worker2 and Worker1, for which the arrangement of each calculation identifier of each micro-batch has been completed.

[0129] Specifically, the scheduling unit 303 identifies the timing when the backward data calculation of micro-batch 3 (refer to the backward data calculation identifier (BD3)) by the worker with worker name = Worker2 is completed. Thereby, the scheduling unit 303 can schedule to execute the process involving communication within the worker at the identified timing.

[0130] As a result, it becomes possible to execute the process involving communication within the worker in parallel with the backward weight calculation of micro-batches 2 to 3 (refer to the backward weight calculation identifiers (BW2, BW3)) by the worker with worker name = Worker2. That is, the scheduling unit 303 can generate a schedule that satisfies the constraint condition (6). In reference numeral 815 in FIG. 13B, the arrows indicated by "parallel calculation start" respectively indicate the start timing at which the process involving communication within the worker with worker name = Worker2 or Worker1 can be executed in parallel with the backward weight calculation.

[0131] Similarly, the scheduling unit 303 identifies the timing at which the backward data calculation of micro-batch 3 (see backward data calculation identifier (BD3)) by the worker with worker name = Worker1 is completed. Thereby, the scheduling unit 303 can schedule to execute a process involving communication within the worker. In reference numeral 816 in FIG. 13B, the arrow indicated by "parallel calculation start" indicates the start timing at which it becomes possible to execute a process involving communication within the worker with worker name = Worker0 in parallel with the backward weight calculation.

[0132] As a result, it becomes possible to execute a process involving communication within the worker in parallel with the backward weight calculation of micro-batches 2 to 3 (see backward weight calculation identifier (BW3)) by the worker with worker name = Worker1. That is, the scheduling unit 303 can generate a schedule that satisfies the constraint condition (6).

[0133] In step S1303, the scheduling unit 303 identifies the timing that satisfies the constraint condition (6) for the worker with worker name = Worker0 in which the arrangement of each calculation identifier of each micro-batch is completed.

[0134] Specifically, the scheduling unit 303 identifies the timing at which the backward data calculation of micro-batch 3 (see backward data calculation identifier (BD3)) by the worker with worker name = Worker0 is completed. Thereby, the scheduling unit 303 can schedule to execute a process involving communication within the worker.

[0135] As a result, it becomes possible to execute a process involving communication within the worker in parallel with the backward weight calculation of micro-batch 3 (see backward weight calculation identifier (BW3)) by the worker with worker name = Worker0. That is, the scheduling unit 303 can generate a schedule that satisfies the constraint condition (6).

[0136] <Summary> As is clear from the above description, the information processing apparatus 120 according to the third embodiment · Along with the addition of the constraint condition (6), when scheduling by the scheduling unit, the timing for executing the process involving communication within the worker is specified.

[0137] In this way, the information processing apparatus 120 according to the third embodiment utilizes the fact that the backward calculation of each micro-batch is subdivided to specify, for each worker, the timing when the backward data calculation of all micro-batches is completed.

[0138] Thereby, according to the information processing apparatus 120 according to the third embodiment, it becomes possible to execute the process involving communication within the worker in parallel with the backward weight calculation, and a schedule that satisfies the constraint conditions can be generated.

[0139] [Other Embodiments] In each of the above embodiments, as the constraint condition (4), it is stipulated that each worker executes the backward weight calculation of each micro-batch after the backward data calculation of each micro-batch.

[0140] However, the constraint condition (4) is a constraint condition on the premise that the backward weight calculation of each micro-batch uses the result of the backward data calculation of each micro-batch. That is, when the backward weight calculation is executed without using the result of the backward data calculation, it is not necessary to satisfy the constraint condition (4). In this case, the scheduling unit 303 may schedule the backward weight calculation to be executed before the backward data calculation.

[0141] Also, in each of the above embodiments, the method of using the network bandwidth between workers included in the scheduling information is not mentioned, but the scheduling unit 303 may perform scheduling including the communication time based on the network bandwidth.

[0142] In each of the above embodiments, it has been described that the number of workers is input in advance as information for scheduling. However, the number of workers may also be the subject of optimization processing by the scheduling unit 303.

[0143] In each of the above embodiments, a micro-batch has been described as an example of the processing unit (batch) of the training data executed by each worker during the training process. However, the batch of training data is not limited to a micro-batch and may be a mini-batch. Also, one batch may be one of a plurality of divisions of the training data, or may include one or more data included in the training data.

[0144] In each of the above embodiments, when causing each worker to execute backward calculation, it has been described that the backward calculation is divided into backward data calculation and backward weight calculation and executed. However, in the scheduling of some workers (for example, worker name = Worker0), the backward data calculation and the backward weight calculation may be scheduled integrally without division.

[0145] In each of the above embodiments, the case where the information processing apparatus 120 applies a scheduling method to the training process has been described. However, the information processing apparatus 120 may apply a scheduling method to processes other than the training process. That is, the information processing apparatus 120 may apply a scheduling method to data other than the training data.

[0146] In each of the above embodiments, the case where the information processing apparatus 120 is provided separately from the server apparatus group 110 has been described. However, the information processing apparatus 120 may be integrated with the server apparatus group 110.

[0147] Specifically, all functions of the information processing apparatus 120 may be realized in some of the servers in the server apparatus group 110. That is, the information processing system 100 may include N server apparatus groups 110 and one information processing apparatus 120, or may include (N - 1) server apparatus groups 110 and one server apparatus. Alternatively, the information processing apparatus 120 itself may be a worker or a part of a worker.

[0148] Also, in each of the above embodiments, the information processing system 100 has been described assuming that there is one information processing apparatus 120, but the information processing apparatus 120 may be constituted by a plurality of apparatuses.

[0149] In this specification (including the claims), when an expression such as "at least one (one) of a, b, and c" or "at least one (one) of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, a - b, a - c, b - c, or a - b - c. Also, it may include multiple instances of any element, such as a - a, a - b - b, a - a - b - b - c - c, etc. Furthermore, it also includes adding other elements other than the enumerated elements (a, b, and c), such as having d like a - b - c - d.

[0150] In addition, in this specification (including the claims), when expressions such as "using data as an input / based on data / in accordance with data / in response to data" (including similar expressions) are used, unless otherwise specified, it includes cases where various data themselves are used as an input, and cases where something obtained by performing some processing on various data (for example, data with noise added, normalized data, intermediate representations of various data, etc.) is used as an input. Also, when it is described that some result is obtained "based on / in accordance with / in response to data", it includes cases where the result is obtained based only on the said data, and may also include cases where the result is obtained under the influence of other data, factors, conditions, and / or states, etc. other than the said data. Further, when it is described that "data is output", unless otherwise specified, it includes cases where various data themselves are used as an output, and cases where something obtained by performing some processing on various data (for example, data with noise added, normalized data, intermediate representations of various data, etc.) is output.

[0151] In addition, in this specification (including the claims), when the terms "connected" and "coupled" are used, they are intended as non - limiting terms that include any of direct connection / coupling, indirect connection / coupling, electrical connection / coupling, communicative connection / coupling, operative connection / coupling, physical connection / coupling, etc. The said terms should be appropriately interpreted according to the context in which they are used, but connection / coupling forms that are not intentionally or naturally excluded should be interpreted non - limitatively as being included in the said terms.

[0152] Also, in this specification (including the claims), when the expression "A is configured to B" is used, the physical structure of element A has a configuration capable of performing operation B, and it may include that a permanent or temporary setting / configuration of element A is set to actually perform operation B. For example, when element A is a general-purpose processor, the processor has a hardware configuration capable of performing operation B, and it may be set to actually perform operation B by a permanent or temporary program (instruction) setting. Also, when element A is a dedicated processor or a dedicated arithmetic circuit, etc., regardless of whether control instructions and data are actually attached, the circuit structure of the processor may be implemented to actually perform operation B.

[0153] Also, in this specification (including the claims), when terms meaning inclusion or possession (e.g., "comprising / including" and "having", etc.) are used, they are intended as open-ended terms including cases where they contain or possess things other than the object indicated by the object of the term. When the object of these terms meaning inclusion or possession does not specify a quantity or is an expression suggesting a singular number (an expression with "a" or "an" as an article), the expression should be interpreted as not being limited to a specific number.

[0154] Also, in this specification (including the claims), even if an expression such as "one or more" or "at least one" is used in one place and an expression that does not specify a quantity or implies a singular number (an expression with "a" or "an" as an article) is used in another place, the latter expression is not intended to mean "one." Generally, an expression that does not specify a quantity or implies a singular number (an expression with "a" or "an" as an article) should be interpreted as not necessarily being limited to a specific number.

[0155] Also, in this specification, if it is described that a specific effect (advantage / result) is obtained for a specific configuration of a certain embodiment, unless there are other reasons, it should be understood that the same effect can also be obtained for one or more other embodiments having the same configuration. However, the presence or absence of the effect generally depends on various factors, conditions, and / or states, etc., and it should be understood that the effect is not necessarily obtained by the configuration. The effect is only obtained by the configuration described in the embodiment when various factors, conditions, and / or states, etc., are satisfied, and the effect is not necessarily obtained in the invention according to the claim that defines the configuration or a similar configuration.

[0156] Also, in this specification (including the claims), when a plurality of hardware performs a predetermined process, each piece of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Also, some of the hardware may perform a part of the predetermined process, and another piece of hardware may perform the remainder of the predetermined process. In this specification (including the claims), when an expression such as "one or more pieces of hardware perform a first process and the one or more pieces of hardware perform a second process" is used, the hardware that performs the first process and the hardware that performs the second process may be the same or different. That is, it is sufficient that the hardware that performs the first process and the hardware that performs the second process are included in the one or more pieces of hardware. Note that the hardware may include an electronic circuit or a device including an electronic circuit, etc.

[0157] Also, in this specification (including the claims), when a plurality of storage devices (memories) store data, each of the plurality of storage devices (memories) may store only a part of the data or may store all of the data.

[0158] As described above in detail for the embodiments of the present disclosure, the present disclosure is not limited to the individual embodiments described above. Various additions, changes, replacements, and partial deletions are possible without departing from the conceptual ideas and spirit of the present invention derived from the content defined in the claims and their equivalents. For example, in all of the embodiments described above, the numerical values used in the description are shown as examples and are not limited thereto. Also, the order of each operation in the embodiments is shown as an example and is not limited thereto.

[0159] In addition, in the disclosed technology, forms such as the following supplementary notes can be considered. (Supplementary Note 1) At least one memory, and at least one processor, wherein the at least one processor identifies information including the execution order when each worker executes the forward calculation of each batch used for training the model, and schedules the execution timing of the forward calculation and backward calculation of each batch executed by each worker during the training of the model based on the identified information so as to satisfy a predetermined constraint condition, an information processing apparatus, wherein the at least one processor schedules the execution timing of at least one worker so that the second calculation included in the backward calculation of the i-th batch is executed after the first calculation included in the backward calculation of the i-th batch and the first calculation included in the backward calculation of the batch after the i-th batch, The second calculation includes a calculation executed using the result of the first calculation. An information processing apparatus. (Appendix 2) The at least one worker is a worker that executes the forward calculation of each batch last in the execution order. The information processing apparatus according to Appendix 1. (Appendix 3) The at least one processor schedules the execution timing so that the second calculation of each batch is executed after the first calculation of all batches is completed in the at least one worker. The information processing apparatus according to Appendix 2. (Appendix 4) The at least one worker includes a first worker and a second worker. In the execution order, the first worker is a worker that executes the forward calculation of each batch before the second worker, The at least one processor schedules the execution timing so that the second calculation of the i-th batch by the first worker is executed before the second calculation of the i-th batch by the second worker. The information processing apparatus according to Appendix 1. (Appendix 5) The at least one processor divides the plurality of calculations included in the backward calculation into at least one first calculation and one or more second calculations. The information processing apparatus according to Appendix 1. (Appendix 6) The predetermined constraint conditions include at least each worker executes the forward calculation of each batch in the execution order, each worker executes the first calculation of each batch after the forward calculation of each batch, each worker executes the first calculation of each batch in an order opposite to the execution order. The information processing apparatus according to Supplementary Note 5, including (Supplementary Note 7) The predetermined constraint condition further Each of the workers executes the second calculation of each batch after the first calculation of each batch The information processing apparatus according to Supplementary Note 6, including (Supplementary Note 8) The at least one processor Identifies the memory capacity of each worker, Calculates the memory usage of each worker at each time during the training of the model, Schedules based on the identified memory capacity and the calculated memory usage, The information processing apparatus according to any one of Supplementary Notes 5 to 7. (Supplementary Note 9) The at least one processor Schedules so that the second calculation is executed before the timing when the calculated memory usage exceeds the identified memory capacity, The information processing apparatus according to Supplementary Note 8. (Supplementary Note 10) The at least one processor Identifies the network bandwidth when transmitting and receiving the result of the first calculation, When scheduling to satisfy the predetermined constraint condition, schedules including the communication time based on the identified network bandwidth, The information processing apparatus according to any one of Supplementary Notes 5 to 9. (Supplementary Note 11) The at least one processor The second calculation that is executed after the first calculations for a plurality of batches by each worker is executed by each worker, Processing for calculating the total sum of the gradients of the model parameters, Processing for updating the model parameters, Processing for distributing the updated model parameters, Processing for distributing the summed gradients, Scheduling at least any one of the processes to be executed in parallel, An information processing apparatus according to any one of Appendices 5 to 10. (Appendix 12) The first calculation is a backward data calculation, The second calculation is a backward weight calculation. An information processing apparatus according to any one of Appendices 5 to 11. (Appendix 13) Equipped with one or more workers, Each worker executes the forward calculation, the first calculation, and the second calculation of each batch at the execution timing scheduled by the information processing apparatus according to any one of Appendices 1 to 12 to train the model. An information processing system. (Appendix 14) At least one processor Identifies information including the execution order when each worker executes the forward calculation of each batch used for training the model and the number of batches, Based on the identified information, scheduling the execution timing of the forward calculation and the backward calculation of each batch executed by each worker during the training of the model so as to satisfy a predetermined constraint condition, which is a scheduling method, The at least one processor Schedules the execution timing of at least one worker so that the second calculation included in the backward calculation of the i-th batch is executed after the first calculation included in the backward calculation of the i-th batch and the first calculation included in the backward calculation of the batches after the i-th batch, The second calculation includes a calculation executed using the result of the first calculation. A scheduling method. (Appendix 15) To at least one processor Identifies information including the execution order when each worker executes the forward calculation of each batch used for training the model and the number of batches, Execute a process of scheduling the execution timing of the forward calculation and backward calculation of each batch, which is executed by each worker during the training of the model, based on the specified information so as to satisfy the specified constraint conditions. Further cause the at least one processor to Schedule the execution timing of at least one worker so that the second calculation included in the backward calculation of the i-th batch is executed after the first calculation included in the backward calculation of the i-th batch and the first calculation included in the backward calculation of the batches after the i-th batch. Execute a process. The second calculation includes a calculation executed using the result of the first calculation. Scheduling program.

Claims

1. At least one memory and, At least one processor, and comprises, The at least one processor, Identifies information including the execution order when each worker executes the forward calculation of each batch used for training the model, An information processing apparatus that schedules the execution timings of the forward calculation and backward calculation of each batch, which are executed by each worker during training of the model, based on the identified information so as to satisfy a predetermined constraint condition, The at least one processor, Schedules the execution timing of at least one worker so that the second calculation included in the backward calculation of the i-th batch is executed after the first calculation included in the backward calculation of the i-th batch and the first calculation included in the backward calculation of the batch after the i-th batch, The second calculation includes a calculation executed using the result of the first calculation, Information processing apparatus.

2. The at least one worker is a worker that executes the forward calculation of each batch last in the execution order, The information processing apparatus according to claim 1.

3. The at least one processor, Schedules the execution timing so that the second calculation of each batch is executed after the first calculation of all batches has been completed in the at least one worker, The information processing apparatus according to claim 2.

4. The at least one worker includes a first worker and a second worker, In the execution order, the first worker is a worker that executes the forward calculation of each batch before the second worker, The at least one processor, Schedules the execution timing so that the second calculation of the i-th batch by the first worker is executed before the second calculation of the i-th batch by the second worker, The information processing apparatus according to claim 1.

5. The at least one processor, Divides a plurality of calculations included in the backward calculation into at least one first calculation and one or more second calculations, The information processing apparatus according to claim 1.

6. The predetermined constraint condition is at least, Each worker executes the forward calculation of each batch in the execution order, Each worker executes the first calculation for each batch after the forward calculation of each batch. Each worker executes the first calculation for each batch in an order opposite to the execution order. The information processing apparatus according to claim 5, comprising:

7. The predetermined constraint condition further includes: Each worker executes the second calculation for each batch after the first calculation for each batch. The information processing apparatus according to claim 6, comprising:

8. The at least one processor: Identifies the memory capacity of each worker; Calculates the memory usage of each worker at each time during the training of the model; Performs scheduling based on the identified memory capacity and the calculated memory usage. The information processing apparatus according to claim 5.

9. The at least one processor: Performs scheduling so that the second calculation is executed before the timing when the calculated memory usage exceeds the identified memory capacity. The information processing apparatus according to claim 8.

10. The at least one processor: Identifies the network bandwidth when transmitting and receiving the result of the first calculation; When performing scheduling to satisfy the predetermined constraint condition, performs scheduling including the communication time based on the identified network bandwidth. The information processing apparatus according to claim 5.

11. The at least one processor: Performs scheduling so that the second calculation, which is executed after the first calculations for a plurality of batches by each worker, is executed by each worker; A process of calculating the total sum of the gradients of the model parameters; A process of updating the model parameters; A process of distributing the updated model parameters; A process of distributing the summed gradients; Performs scheduling to be executed in parallel with at least any one of the processes. The information processing apparatus according to claim 5.

12. The first calculation is backward data calculation, and the second calculation is backward weight calculation. The information processing apparatus according to claim 5.

13. Comprising one or more workers, wherein each worker trains a model by executing the forward calculation, the first calculation, and the second calculation for each batch at the execution timing scheduled by the information processing apparatus according to any one of claims 1 to 12. An information processing system.

14. At least one processor: ​ Identify information including the execution order when each worker executes the forward calculation for each batch used in training the model and the number of batches, A scheduling method for scheduling the execution timing of the forward calculation and backward calculation of each batch to be executed by each worker during training of the model based on the identified information so as to satisfy a predetermined constraint condition, The at least one processor, Schedule the execution timing of at least one worker so that the second calculation included in the backward calculation of the i-th batch is executed after the first calculation included in the backward calculation of the i-th batch and the first calculation included in the backward calculation of the batches after the i-th batch, The second calculation includes a calculation executed using the result of the first calculation, Scheduling method.

15. To at least one processor, Identify information including the execution order when each worker executes the forward calculation for each batch used in training the model and the number of batches, Execute a process of scheduling the execution timing of the forward calculation and backward calculation of each batch to be executed by each worker during training of the model based on the identified information so as to satisfy a predetermined constraint condition, To the at least one processor, further, Schedule the execution timing of at least one worker so that the second calculation included in the backward calculation of the i-th batch is executed after the first calculation included in the backward calculation of the i-th batch and the first calculation included in the backward calculation of the batches after the i-th batch, Execute the process, The second calculation includes a calculation executed using the result of the first calculation, Scheduling program.