Electronic device and control method of electronic device
The electronic device optimizes resource allocation and communication costs across multiple devices to enhance the efficiency and performance of distributed learning for large-scale AI models, addressing the challenges of uneven resource distribution and high communication costs.
Patent Information
- Application Number
- PCT/KR2024/017091
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-11-01
- Publication Date
- 2025-06-05
AI Technical Summary
Current methods for distributed learning of large-scale artificial intelligence models face challenges in achieving high performance with minimal resources and costs, particularly when resources across devices are unevenly distributed, leading to increased communication costs and deteriorated learning performance.
An electronic device equipped with a processor that obtains information about resources across multiple devices, identifies optimal resource combinations for various parallelization methods (data, pipeline, tensor parallelization), and adjusts resource allocation to minimize communication costs and maximize learning efficiency.
The proposed solution enables efficient distributed learning by optimizing resource allocation and communication costs, thereby improving learning performance and reducing resource utilization inefficiencies.
Smart Images

Figure KR2024017091_05062025_PF_FP_ABST
Abstract
Description
Electronic devices and methods of controlling electronic devices
[0001] The present disclosure relates to an electronic device and a control method thereof, and more particularly, to an electronic device that performs distributed learning using multiple parallelism methods and a control method thereof.
[0002] In the current field of artificial intelligence, training large-scale AI models requires significant resources and costs. This situation necessitates a training method for AI models that can achieve high performance with minimal resources and costs.
[0003] In particular, to train large-scale artificial intelligence models, methods for performing distributed learning across multiple devices using multiple parallelization methods (e.g., data parallelization, pipeline parallelization, tensor parallelization, etc.) have recently been provided. Conventionally, when performing distributed learning using multiple parallelization methods, the resources (e.g., GPUs) contained within the multiple devices were evenly divided to perform distributed learning. However, when performing distributed learning by evenly dividing resources in an environment where the types or numbers of resources differ across the multiple devices, learning performance may deteriorate. Specifically, when the types or numbers of resources differ across the multiple devices and the resources are evenly divided, the resources contained within the multiple devices can be allocated as a single group for performing tensor parallelization. In this case, tensor parallelization requires frequent communication between resources, whereas when the resources contained within the multiple devices are allocated as a single group, communication costs increase, which may degrade learning performance.
[0004] An electronic device for distributed learning of a neural network model through multiple parallelization methods, comprising: a communication interface; a memory storing a neural network model including multiple layers and at least one instruction; And a processor for controlling the electronic device; wherein the processor, by executing the at least one instruction, obtains information about resources included in a plurality of devices for performing distributed learning of the neural network model through a plurality of parallelism methods, wherein the plurality of parallelism methods include data parallelism, pipeline parallelism, and tensor parallelism, and identifies a plurality of resource combinations for performing the plurality of parallelism methods based on the information about the resources included in the plurality of devices, obtains first throughput information about the resources included in the plurality of devices for each of the plurality of layers, obtains stage information for partitioning the layers included in the neural network model for each of the plurality of resource combinations based on the first throughput information about the resources included in the plurality of devices, performs distributed learning through the plurality of parallelism methods for each of the plurality of resource combinations based on the stage information, obtains second throughput information about the plurality of resource combinations, and obtains information about an optimal distributed learning strategy based on the second throughput information about the plurality of resource combinations.
[0005] The processor can group resources included in each of the plurality of devices based on the number of resources included in the plurality of devices, and identify a plurality of resource combinations for performing distributed learning through the plurality of parallelization methods based on the grouped resources.
[0006] The processor may obtain a common divisor of the number of resources included in the plurality of devices, group the resources included in each of the plurality of devices into resources for performing the tensor parallelization based on the obtained common divisor, obtain information on the number of divisible resource groups based on the obtained common divisor and the number of the devices, obtain information on the number of resources for performing the data parallelization and the number of resources for performing the pipeline parallelization based on the divisor of the number of divisible resource groups, and identify a combination of a plurality of resources for performing the plurality of parallelization methods based on the information on the resource groups for performing the tensor parallelization, the number of resources for performing the data parallelization, and the number of resources for performing the pipeline parallelization.
[0007] The processor may divide the plurality of devices into a plurality of virtual devices based on the number of resources included in the plurality of devices so that the plurality of devices may have the same number of resources, and may identify a plurality of resource combinations for performing the plurality of parallelization methods based on the resources included in the plurality of virtual devices.
[0008] The processor may obtain the greatest common divisor of resources included in the plurality of devices, divide virtual devices into each of the plurality of devices based on the greatest common divisor, and obtain the plurality of virtual devices, obtain the greatest common divisor of resources included in the plurality of virtual devices, obtain information on the number of resources for performing pipeline parallelization based on the common divisor, obtain information on the number of resources for performing data parallelization and the number of resources for performing tensor parallelization based on resources remaining from among the resources included in the plurality of virtual devices excluding the resources for performing the pipeline parallelization, and identify a plurality of resource combinations for performing the plurality of parallelization methods based on the information on the number of resources for performing the tensor parallelization, the number of resources for performing the data parallelization, and the number of resources for performing the pipeline parallelization.
[0009] The processor performs a profile for the plurality of layers using resources included in the plurality of devices to obtain first processing rate information of the resources included in the plurality of devices for each of the plurality of layers, and the first processing rate information may include at least one of execution time and memory usage information.
[0010] The processor, based on first processing rate information of resources included in the plurality of devices, performs learning on a plurality of stages obtained by partitioning the layers included in the neural network model using the resources included in the plurality of devices, and can obtain stage information by partitioning the layers so that the deviation in execution time for performing the plurality of stages is the smallest.
[0011] The processor obtains stage information for partitioning the layers included in the neural network model by batch size based on first processing rate information of resources included in the plurality of devices, and performs distributed learning for each of the plurality of resource combinations through the plurality of parallelization methods based on the stage information obtained for each batch size to obtain second processing rate information for the plurality of resource combinations, and obtains information on an optimal distributed learning strategy that enables the fastest distributed learning based on the second processing rate information for the plurality of resource combinations.
[0012] The above optimal distributed learning strategy may include information about optimal resource combination, optimal stage information, and optimal batch size.
[0013] According to one embodiment of the present disclosure, a control method of an electronic device for distributed learning of a neural network model through a plurality of parallelization methods comprises: obtaining information about resources included in a plurality of devices for distributed learning of the neural network model through a plurality of parallelization methods, wherein the plurality of parallelization methods include data parallelism, pipeline parallelism, and tensor parallelism; identifying a plurality of resource combinations for performing the plurality of parallelization methods based on the information about the resources included in the plurality of devices; obtaining first throughput information about resources included in the plurality of devices for each of the plurality of layers; obtaining stage information for partitioning the layers included in the neural network model for each of the plurality of resource combinations based on the first throughput information about resources included in the plurality of devices; performing distributed learning through the plurality of parallelization methods for each of the plurality of resource combinations based on the stage information to obtain second throughput information about the plurality of resource combinations; and a step of obtaining information on an optimal distributed learning strategy based on the second processing rate information for the plurality of resource combinations.
[0014] The identifying step may group resources included in each of the plurality of devices based on the number of resources included in the plurality of devices, and identify a plurality of resource combinations for performing distributed learning through the plurality of parallelization methods based on the grouped resources.
[0015] The identifying step may include: obtaining a common divisor of the number of resources included in the plurality of devices; grouping the resources included in each of the plurality of devices into resources for performing the tensor parallelization based on the obtained common divisor; obtaining information on the number of divisible resource groups based on the obtained common divisor and the number of devices; obtaining information on the number of resources for performing the data parallelization and the number of resources for performing the pipeline parallelization based on the divisor of the number of divisible resource groups; and identifying a plurality of resource combinations for performing the plurality of parallelization methods based on the information on the resource groups for performing the tensor parallelization, the number of resources for performing the data parallelization, and the number of resources for performing the pipeline parallelization.
[0016] The identifying step may divide the plurality of devices into a plurality of virtual devices based on the number of resources included in the plurality of devices so that the plurality of devices can have the same number of resources, and identify a plurality of resource combinations for performing the plurality of parallelization methods based on the resources included in the plurality of virtual devices.
[0017] The identifying step may include: obtaining the greatest common divisor of resources included in the plurality of devices; dividing virtual devices into each of the plurality of devices based on the greatest common divisor to obtain the plurality of virtual devices; obtaining the greatest common divisor of resources included in the plurality of virtual devices; obtaining information on the number of resources for performing pipeline parallelization based on the common divisor; obtaining information on the number of resources for performing data parallelization and the number of resources for performing tensor parallelization based on resources remaining from among the resources included in the plurality of virtual devices excluding the resources for performing the pipeline parallelization; and identifying a plurality of resource combinations for performing the plurality of parallelization methods based on the information on the number of resources for performing the tensor parallelization, the number of resources for performing the data parallelization, and the number of resources for performing the pipeline parallelization.
[0018] The step of obtaining the first processing rate information includes performing a profile for the plurality of layers using resources included in the plurality of devices to obtain first processing rate information of the resources included in the plurality of devices for each of the plurality of layers, and the first processing rate information may include at least one of execution time and memory usage information.
[0019] The step of obtaining the stage information may be performed by partitioning the layers included in the neural network model based on first processing rate information of the resources included in the plurality of devices, and learning is performed on the plurality of stages obtained by using the resources included in the plurality of devices, thereby obtaining stage information by partitioning the layers so that the deviation in execution time for performing the plurality of stages is the smallest.
[0020] The step of obtaining the stage information may obtain stage information by partitioning the layers included in the neural network model by batch size based on first processing rate information of resources included in the plurality of devices, and the step of obtaining information on the second processing rate may obtain second processing rate information for the plurality of resource combinations by performing distributed learning through the plurality of parallelization methods for each of the plurality of resource combinations based on the stage information obtained by the batch size, and the step of obtaining information on the optimal distributed learning strategy may obtain information on the optimal distributed learning strategy that enables the fastest distributed learning based on the second processing rate information for the plurality of resource combinations.
[0021] The above optimal distributed learning strategy may include information about optimal resource combination, optimal stage information, and optimal batch size.
[0022] FIG. 1 is a block diagram illustrating a configuration of an electronic device according to one embodiment of the present disclosure;
[0023] FIG. 2 is a flowchart illustrating a method for controlling an electronic device for distributed learning of a neural network model through multiple parallelization methods according to one embodiment of the present disclosure;
[0024] FIGS. 3A to 3D are diagrams illustrating a method for obtaining a resource combination by grouping resources included in a plurality of devices according to one embodiment of the present disclosure.
[0025] FIGS. 4A to 4D are diagrams for explaining a method of obtaining a resource combination by identifying multiple virtual devices through resources included in multiple devices according to one embodiment of the present disclosure.
[0026] FIG. 5 is a drawing illustrating a method for obtaining optimal stage information for performing pipeline parallelization according to one embodiment of the present disclosure;
[0027] FIG. 6 is a diagram illustrating an optimal distributed learning strategy for a neural network model according to one embodiment of the present disclosure.
[0028] The present embodiments may be modified and have various embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope to specific embodiments, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0029] In describing the present disclosure, if it is determined that a specific description of a related known function or configuration may unnecessarily obscure the gist of the present disclosure, a detailed description thereof will be omitted.
[0030] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concepts of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to further faithfully and completely convey the technical concepts of the present disclosure to those skilled in the art.
[0031] The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the scope of the rights. Singular expressions include plural expressions unless the context clearly dictates otherwise.
[0032] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0033] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0034] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0035] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that said component may be directly coupled to said other component, or may be coupled via another component (e.g., a third component).
[0036] On the other hand, when it is said that a component (e.g., a first component) is "directly connected" or "directly connected" to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between said component and said other component.
[0037] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0038] Instead, in some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.
[0039] In the embodiments, a 'module' or 'part' performs at least one function or operation, and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, a plurality of 'modules' or 'parts' may be integrated into at least one module and implemented as at least one processor, except for a 'module' or 'part' that needs to be implemented as a specific hardware.
[0040] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0041] According to the present disclosure, a "neural network model" is a model implemented based on the neural network of the human brain, and may refer to an overall model in which artificial neurons that form a network by combining synapses change the strength of the synapses through learning to have problem-solving capabilities. At this time, the neural network model includes multiple layers (or strata), and for example, may include an input layer, an output layer, and a number of hidden layers between them. At this time, the neural network model may include an artificial neural network (ANN) model, a deep neural network (DNN) model, etc.
[0042] According to the present disclosure, "distributed learning" refers to a method of dividing and learning a neural network model or data across multiple devices (wherein the multiple devices may be referred to as "multiple nodes" or "multiple machines"). This method can broadly utilize data parallelism or model parallelism. Here, model parallelism may include pipeline parallelism and tensor parallelism.
[0043] According to the present disclosure, "parallelism" refers to a method of performing distributed learning in parallel by dividing a neural network model across multiple devices (or multiple nodes) or multiple resources. The parallelization method may include data parallelism, pipeline parallelism, and tensor parallelism.
[0044] According to the present disclosure, "data parallelism" is a technique for parallel processing of training data in a neural network model. In other words, data parallelism is a technique for training a neural network model by allocating the entire neural network model to each of multiple resources, dividing the entire training data into batches of a certain size, and allocating them to each resource.
[0045] According to the present disclosure, "pipeline parallelism" is a technique for parallel processing a neural network model by dividing it into layers. In other words, pipeline parallelism is a technique for training a neural network model by dividing multiple resources included in a neural network model into multiple stages and assigning each of the multiple stages to a resource.
[0046] According to the present disclosure, "tensor parallelism" is a technique for dividing and parallelizing layers included in a neural network model. In other words, tensor parallelism is a technique for training a neural network model by dividing multiple layers included in a neural network model and assigning them to multiple resources. Unlike pipeline parallelism, which divides each layer individually, tensor parallelism is a method of assigning a single layer to multiple resources, and thus requires a sync operation that combines the outputs of intermediate or final layers.
[0047] According to the present disclosure, a “resource” includes a computing resource for performing various functions such as specific calculations or tasks, and may include hardware such as a processor or core included in a processor, such as a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), etc.
[0048] Hereinafter, with reference to the attached drawings, embodiments according to the present disclosure will be described in detail so that a person having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement the present disclosure.
[0049] FIG. 1 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure. As illustrated in FIG. 1, the electronic device (100) may include a communication interface (110), a memory (120), and a processor (130). Meanwhile, the configuration of the electronic device (100) illustrated in FIG. 1 is merely an example, and it is understood that some additional configurations may be added depending on the type of the electronic device (100).
[0050] An electronic device (100) according to one embodiment of the present disclosure may be implemented as a server, but this is merely one embodiment, and may be implemented as a user terminal such as a smart phone, tablet PC, laptop PC, etc., and may be implemented as various devices such as a smart TV, home appliance, IoT device, etc. In addition, when the electronic device (100) is implemented as a server, it goes without saying that it may be implemented as two or more servers.
[0051] The communication interface (110) includes a circuit and can communicate with an external device (server or user terminal). Specifically, the processor (130) can receive various data or information from an external device connected via the communication interface (110) and can also transmit various data or information to the external device.
[0052] The communication interface (110) may include at least one of a WiFi module, a Bluetooth module, a wireless communication module, an NFC module, and a UWB module (Ultra Wide Band). At this time, the wireless communication module may perform communication according to various communication standards such as IEEE, Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), 5G (5th Generation), etc.
[0053] In particular, the communication interface (110) can receive information about resources from multiple devices for distributed learning of neural network models through multiple parallelization methods. The information about the resources may include information such as the type of the resource, the product name or number of the resource, and the number of resources.
[0054] The memory (120) may store an operating system (OS) for controlling the overall operation of the components of the electronic device (100) and instructions or data related to the components of the electronic device (100). In particular, the memory (120) may include various modules for performing distributed learning using multiple parallelization methods. In particular, when a function for performing distributed learning using multiple parallelization methods is executed, the electronic device (100) may load data for various modules for performing distributed learning using multiple parallelization methods stored in a non-volatile memory into a volatile memory. Here, loading means an operation of loading and storing data stored in a non-volatile memory into a volatile memory so that the processor (130) can access it.
[0055] Additionally, the memory (120) can store information about a neural network model including multiple layers. In this case, the neural network model may be a language model capable of processing large amounts of data, such as speech recognition or natural language understanding. However, this is merely an example, and may be a neural network model capable of performing various operations, such as object recognition, translation, document summarization, and image generation.
[0056] Meanwhile, the memory (120) may be implemented as a non-volatile memory (e.g., hard disk, SSD (Solid state drive), flash memory), volatile memory (may also include memory within the processor (120)), etc.
[0057] The processor (130) can control the electronic device (100) according to at least one instruction stored in the memory (120).
[0058] In particular, the processor (130) may include a plurality of processors. Specifically, the plurality of processors may include one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), a Neural Processing Unit (NPU), a hardware accelerator, or a machine learning accelerator. The one or more processors may control one or any combination of other components of the electronic device and perform operations related to communication or data processing. The one or more processors may execute one or more programs or instructions stored in a memory. For example, the plurality of processors may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in a memory.
[0059] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one processor or by a plurality of processors. That is, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-dedicated processor). For example, the operation of identifying an optimal distributed learning strategy may be performed by the first processor (e.g., a CPU), and the operation of distributed learning a neural network model may be performed by the second processor (e.g., a GPU or NPU).
[0060] At least one processor (120) may be implemented as one or more multi-core processors including multiple cores (e.g., homogeneous multi-cores or heterogeneous multi-cores). When at least one processor (120) is implemented as a multi-core processor, each of the multiple cores included in the multi-core processor may include internal processor memory, such as cache memory or on-chip memory, and a common cache shared by the multiple cores may be included in the multi-core processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multi-core processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.
[0061] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among a plurality of cores included in a multi-core processor, or may be performed by a plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.
[0062] In embodiments of the present disclosure, a processor may mean a system on a chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but embodiments of the present disclosure are not limited thereto.
[0063] Meanwhile, the processor (130) obtains information about resources included in multiple devices for distributed learning of a neural network model through multiple parallelism methods by executing at least one instruction. Here, the multiple parallelism methods include data parallelism, pipeline parallelism, and tensor parallelism. In addition, the multiple devices may include the electronic device (100), but this is only one embodiment, and the multiple devices may be implemented as multiple external devices that do not include the electronic device (100).
[0064] The processor (130) identifies a plurality of resource combinations for performing a plurality of parallelization methods based on information about resources included in a plurality of devices. At this time, the resource combination may mean a combination of resources for performing data parallelization, pipeline parallelization, and tensor parallelization. For example, if a total of 8 resources are included in a plurality of devices, the resource combinations are (1,1,8), (1,2,4),(1,4,2),(1,8,1),(2,1,4),(2,2,2,),(2,4,1),(4,2,1),(4,1,2),(8,1,1). At this time, in (x,y,z), x may mean the number (or size) of resources for performing pipeline parallelization, y may mean the number (or size) of resources for performing data parallelization, and z may mean the number (or size) of resources for performing tensor parallelization. Here, performing parallelization may mean performing distributed learning according to a parallelization method.
[0065] The processor (130) obtains first processing rate information of resources included in multiple devices for each of multiple layers. At this time, the first processing rate information is information related to the speed at which one layer included in a neural network model is processed when performing training of a neural network model using the resources, and may include at least one of information regarding the execution time for processing one layer, the amount of data stored in the memory (120) through an operation for processing one layer, etc.
[0066] The processor (130) obtains information about a stage into which layers included in a neural network model are partitioned for each combination of resources based on first processing rate information of resources included in multiple devices. At this time, the stage may refer to a group of layers into which layers included in a neural network model are divided in order to perform pipeline parallelization. For example, in a situation where there are 10 layers included in a neural network model, if layers are divided into 5 each in order to perform pipeline parallelization, the 1st to 5th layers may be referred to as the 1st stage, and the 6th to 10th layers may be referred to as the 2nd stage.
[0067] Based on the processor (130) stage information, distributed learning is performed for each of a plurality of resource combinations through a plurality of parallelization methods to obtain second processing rate information for the plurality of resource combinations. The second processing rate information is information related to the learning speed of the neural network model when performing learning of the neural network model through the resource combination, and may include at least one of information on the execution time for learning the neural network model per iteration, the amount of data stored in the memory (120) through the operation of learning the neural network model, etc.
[0068] The processor (130) obtains information on an optimal distributed learning strategy based on second processing rate information for multiple resource combinations. The optimal distributed learning strategy may include information on an optimal resource combination, optimal stage information, and optimal batch size that can achieve the best (or fastest) learning performance when performing distributed learning on a neural network model.
[0069] In one embodiment, the processor (130) may group the resources included in each of the plurality of devices based on the number of resources included in the plurality of devices, and identify a plurality of resource combinations for performing distributed learning through a plurality of parallelization methods based on the grouped resources. Specifically, the processor (130) may obtain a common divisor of the number of resources included in the plurality of devices. The processor (130) may group the resources included in each of the plurality of devices into resources for performing tensor parallelization based on the obtained common divisor. The processor (130) may obtain information on the number of resource groups that can be split based on the obtained common divisor and the number of devices. The processor (130) may obtain information on the number of resources for performing data parallelization and the number of resources for performing pipeline parallelization based on the divisor of the number of resource groups that can be split. The processor (130) may identify a plurality of resource combinations for performing a plurality of parallelization methods based on the information on the resource groups for performing tensor parallelization, the number of resources for performing data parallelization, and the number of resources for performing pipeline parallelization.
[0070] In one embodiment, the processor (130) may divide the plurality of devices into a plurality of virtual devices based on the number of resources included in the plurality of devices so that the plurality of devices can have the same number of resources, and identify a plurality of resource combinations for performing a plurality of parallelization methods based on the resources included in the plurality of virtual devices. Specifically, the processor (130) may obtain the greatest common divisor of the resources included in the plurality of devices. The processor (130) may obtain a plurality of virtual devices by dividing the virtual devices into the plurality of devices based on the greatest common divisor. The processor (130) may obtain a common divisor of the resources included in the plurality of virtual devices. The processor (130) may obtain information about the number of resources for performing pipeline parallelization based on the common divisor. The processor (130) may obtain information about the number of resources for performing data parallelization and the number of resources for performing tensor parallelization based on the remaining resources excluding the resources for performing pipeline parallelization among the resources included in the plurality of virtual devices. The processor (130) can identify a plurality of resource combinations for performing a plurality of parallelization methods based on information about the number of resources for performing tensor parallelization, the number of resources for performing data parallelization, and the number of resources for performing pipeline parallelization.
[0071] In one embodiment, the processor (130) may perform a profile on multiple layers using resources included in multiple devices to obtain first throughput information on the resources included in each of the multiple layers. Here, "profile" refers to an operation of measuring the performance of a layer using specific resources. The first throughput information may include at least one of execution time and memory usage information.
[0072] In one embodiment, the processor (130) partitions layers included in a neural network model based on first processing rate information of resources included in multiple devices, and performs learning for multiple stages obtained by using resources included in multiple devices, and as a result, acquires stage information for partitioning layers so that the deviation in execution time for performing multiple stages is the smallest.
[0073] In one embodiment, the processor (130) may obtain stage information for partitioning the layers included in the neural network model by batch size based on first processing rate information of resources included in multiple devices. In this case, the batch size refers to the number of data samples input to the neural network model at one time to learn the neural network model. The processor (130) may perform distributed learning for each of the plurality of resource combinations through multiple parallelization methods based on the stage information obtained by batch size, thereby obtaining second processing rate information for the plurality of resource combinations. The processor (130) may obtain information on an optimal distributed learning strategy that enables the fastest distributed learning based on the second processing rate information for the plurality of resource combinations.
[0074]
[0075] Hereinafter, the present disclosure will be described in more detail with reference to FIGS. 2 to 6.
[0076] FIG. 2 is a flowchart illustrating a method for controlling an electronic device for distributed learning of a neural network model through multiple parallelization methods according to one embodiment of the present disclosure.
[0077] First, the electronic device (100) acquires information about resources included in multiple devices for distributed learning of a neural network model through multiple parallelization methods (S210). Here, the multiple parallelization methods may include data parallelization, pipeline parallelization, and tensor parallelization.
[0078] In one embodiment, the electronic device (100) may receive information about resources included in the plurality of devices from each of the plurality of devices. The information about the resources may include the types of resources included in the plurality of devices (e.g., GPU, NPU, CPU, etc.), product information about the resources (e.g., product name, product number, etc.), the number of resources, etc. In one embodiment, the electronic device (100) may obtain information about resources included in the plurality of devices through user input.
[0079] The electronic device (100) identifies multiple resource combinations for performing multiple parallelization methods based on information about resources included in multiple devices (S220). Specifically, in the past, regardless of the resources included in multiple devices, the same number of resources were allocated and distributed evenly. For example, as illustrated in FIG. 3A, when a first device (310) includes four resources (310-1 to 310-4) and a second device (320) includes eight resources (320-1 to 320-8) and an even partitioning method is used, if a resource combination with three or six resources for performing tensor parallelization occurs among the multiple resource combinations, resources from other devices may be included in a group for performing a single tensor parallelization. As described above, since tensor parallelization requires frequent communication between resources, if resources included in different devices are included in a group for performing a single tensor parallelization, a problem of performance degradation may occur due to frequent communication.
[0080] Accordingly, the present invention can create a resource combination in the following manner so as not to include resources included in different devices within a group for one tensor parallelization.
[0081] In one embodiment, the electronic device (100) can group resources contained in each of the multiple devices based on the number of resources contained in the multiple devices, and identify multiple resource combinations for performing distributed learning using multiple parallelization methods based on the grouped resources. The present disclosure will be described in more detail below with reference to FIGS. 3A to 3D .
[0082] Specifically, the electronic device (100) can obtain a common divisor of the number of resources included in the plurality of devices. For example, as illustrated in FIG. 3A, if the first device (310) includes four resources (310-1 to 310-4) and the second device (320) includes eight resources (320-1 to 320-8), the electronic device (100) can obtain 4, which is a common divisor of the number of resources included in the plurality of devices (310, 320). Then, the electronic device (100) can group the resources included in each of the plurality of devices into resources for performing tensor parallelization based on the obtained common divisor. For example, as illustrated in FIG. 3B, the electronic device (100) can group the resources included in the plurality of devices (310, 320) so that one device has a number of groups corresponding to the common divisor 4. That is, the electronic device (100) can group the 1-1 to 1-4 resources (310-1 to 310-4) included in the first device (310) into one group, and can group the 2-1 to 2-8 resources (320-1 to 320-8) included in the second device (320) into four groups by grouping them in pairs. When grouping in this manner, the electronic device (100) can group the resources included in the first device (310) into four groups (310-1 to 310-4), as illustrated in FIG. 3C, and can group the resources included in the second device (310) into four groups (330-1 to 330-4). By this, in an environment where two devices have 12 resources, the electronic device (100) can obtain information about two devices including four resource groups, thereby obtaining a search space that enables equal division as shown in FIG. 3d.
[0083] As another example, as illustrated in FIG. 3A, if the first device (310) includes four resources (310-1 to 310-4) and the second device (320) includes eight resources (320-1 to 320-8), the electronic device (100) can obtain 1 or 2, which is a common divisor of the number of resources included in the plurality of devices (310, 320). The electronic device (100) can group the resources included in the plurality of (310, 320) so that one device has a number of groups corresponding to the common divisor 2. That is, the electronic device (100) can group the 1-1 to 1-4 resources (310-1 to 310-4) included in the first device (310) into two groups by grouping them in pairs, and can group the 2-1 to 2-8 resources (320-1 to 320-8) included in the second device (320) into two groups by grouping them in pairs. As another example, the electronic device (100) can group the resources included in a plurality of (310, 320) so that one device has a number of groups corresponding to the common divisor 1. That is, the electronic device (100) can group the 1-1 to 1-4 resources (310-1 to 310-4) included in the first device (310) into four groups and group them into one group, and can group the 2-1 to 2-8 resources (320-1 to 320-8) included in the second device (320) into eight groups and group them into one group.
[0084] In addition, the electronic device (100) can obtain information on the number of divisible resource groups based on the obtained common divisor and the number of devices. At this time, the number of divisible resource groups can be obtained by multiplying the common divisor and the number of devices. For example, if the common divisor is 4, the number of divisible resource groups can be 8, which is the product of the common divisor and the number of devices. As another example, if the common divisor is 2, the number of divisible resource groups can be 4, which is the product of the common divisor and the number of devices. As another example, if the common divisor is 1, the number of divisible resource groups can be 2, which is the product of the common divisor and the number of devices.
[0085] The electronic device (100) can obtain information about the number of resources for performing data parallelization and the number of resources for performing pipeline parallelization based on a divisor of the number of divisible resource groups. For example, when the number of divisible resource groups is 8, since the divisor of the number of divisible resource groups is [1,2,4,8], the electronic device (100) can obtain (1,8), (2,4), (4,2), (8,1) as a combination of the number of resources for performing data parallelization and the number of resources for performing pipeline parallelization. As another example, when the number of divisible resource groups is 4, since the divisor of the number of divisible resource groups is [1,2,4], the electronic device (100) can obtain (1,4), (2,2), (4,1) as a combination of the number of resources for performing data parallelization and the number of resources for performing pipeline parallelization. As another example, if the number of divisible resource groups is 2, since the divisor of the number of divisible resource groups is [1,2], the electronic device (100) can obtain (1,2) and (2,1) as a combination of the number of resources for performing data parallelism and the number of resources for performing pipeline parallelism.
[0086] The electronic device (100) can identify multiple resource combinations for performing multiple parallelization methods based on information about resource groups for performing tensor parallelization, the number of resources for performing data parallelization, and the number of resources for performing pipeline parallelization. In one embodiment, the electronic device (100) can obtain resource combinations as shown in Table 1 below.
[0087] - Number of resources of the first device: 4 - Number of resources of the second device: 8 - Common divisor of resources included in the first device and the second device: [1,2,4] - Resource combination (PP_size, DP_size, TP_group) (1, 8, TP_group(4)) At this time, TP_group(4) is 1,2 (2, 4, TP_group(4)) At this time, TP_group(4) is 1,2 (4, 2, TP_group(4)) At this time, TP_group(4) is 1,2 (8, 1, TP_group(4)) At this time, TP_group(4) is 1,2 (1, 4, TP_group(2)) At this time, TP_group(4) is 2,4 (2, 2, TP_group(2)) At this time, TP_group(2) is 2,4 (4, 1, TP_group(2)) At this time, TP_group(2) is 2,4(2, 1, TP_group(1)) At this time, TP_group(1) is 4,8(1, 2, TP_group(1)) At this time, TP_group(1) is 4,8
[0088] In the above-described Table 1, PP_size is the number of resources for performing pipeline parallelization, DP_size is the number of resources for performing data parallelization, and TP_group is a resource group for performing tensor parallelization. TP_group(4) means that 1 and 2 resources of the first device (100) and 2 resources of the second device (200) are used, respectively, TP_group(2) means that 2 and 4 resources of the first device (100) and 2 resources of the second device (200) are used, respectively, and TP_group(1) may mean that 4 and 8 resources of the first device (100) and 2 resources of the second device (200) are used, respectively. In the same manner as the above-described table, when the number of resources included in multiple devices is equalized by grouping resources included in multiple devices, it is possible to prevent resources included in different devices from being included in order to perform tensor parallelization.
[0089] In one embodiment, the electronic device (100) may divide the plurality of devices into a plurality of virtual devices based on the number of resources included in the plurality of devices so that the plurality of devices may have the same number of resources, and may identify a plurality of resource combinations for performing a plurality of parallelization methods based on the resources included in the plurality of virtual devices. The present disclosure will be described in more detail below with reference to FIGS. 4A to 4D . Meanwhile, for the present embodiment to be applied, the number of resources included in a device having the smallest number of resources among the plurality of devices may be a divisor of the number of resources of the other devices.
[0090] The electronic device (100) can obtain the greatest common divisor of resources included in multiple devices. For example, as illustrated in FIG. 4A, if a first device (410) includes four resources (410-1 to 410-4) and a second device (420) includes eight resources (420-1 to 420-8), the electronic device (100) can obtain 4, which is the greatest common divisor of the number of resources included in the multiple devices (310, 320).
[0091] And, the electronic device (100) can obtain multiple virtual devices by dividing the virtual devices for each of the multiple devices based on the greatest common divisor. For example, as illustrated in FIG. 4a, in a situation where the first device (410) includes four resources (410-1 to 410-4) and the second device (420) includes eight resources (420-1 to 420-8), since the greatest common divisor is 4, the electronic device (100) can divide the resources of the multiple devices (410, 420) so that they have a number of resources corresponding to the greatest common divisor, as illustrated in FIG. 4b. That is, as illustrated in FIG. 4b, the electronic device (100) can obtain one virtual device for the first device (410) so that it has four resources, and can divide the second device (420) into two virtual devices so that it has four resources. Accordingly, the electronic device (100) can obtain a first virtual device (410) having four resources (410-1 to 410-4), a second virtual device (430) having four resources (420-1, 420-2, 420-5, 420-6), and a third virtual device (440) having four resources (420-3, 420-4, 420-7, 420-8), as illustrated in FIG. 4C. Meanwhile, the embodiment of dividing the second device (420) into a plurality of two virtual devices is not limited to FIG. 4C, and can be divided into two virtual devices in various combinations. For example, the second virtual device (430) may include four resources (420-1 to 420-4), and the third virtual device (440) may include four resources (420-5 to 420-8). Accordingly, in an environment where two devices have 12 resources, the electronic device (100) can obtain information about three virtual devices, each of which includes four resources, thereby obtaining a search space that enables an even division as illustrated in FIG. 4d.
[0092] In addition, the electronic device (100) can obtain a common divisor of the resources included in the plurality of virtual devices. For example, since there are 12 resources included in the plurality of virtual devices, the electronic device (100) can obtain [1,2,3,4,6,12] as a common divisor of the resources included in the plurality of virtual devices.
[0093] In addition, the electronic device (100) can obtain information on the number of resources for performing pipeline parallelism based on the common divisor. That is, the electronic device (100) can set 1, 2, 3, 4, 6, and 12 as resources for performing pipeline parallelism.
[0094] In addition, the electronic device (100) can obtain information about the number of resources for performing data parallelism and the number of resources for performing tensor parallelism based on the remaining resources, excluding the resources for performing pipeline parallelism, among the resources included in the multiple virtual devices. In this case, cases where the number of resources for performing tensor parallelism exceeds the number of resources included in the virtual device (i.e., 4) can be excluded.
[0095] For example, if there is 1 resource for performing pipeline parallelization, the number of remaining resources is 12, and since the number of resources for performing tensor parallelization cannot exceed 4, the electronic device (100) can obtain the number of resources for performing data parallelization and the number of resources for performing tensor parallelization as (3,4), (4,3), (6,2), (12,1). In addition, if there are 2 resources for performing pipeline parallelization, the number of remaining resources is 6, and since the number of resources for performing tensor parallelization cannot exceed 4, the electronic device (100) can obtain the number of resources for performing data parallelization and the number of resources for performing tensor parallelization as (2,3), (3,2), (6,1). If there are 3 resources for performing pipeline parallelism, the number of remaining resources is 4, and since the number of resources for performing tensor parallelism cannot exceed 4, the electronic device (100) can obtain the number of resources for performing data parallelism and the number of resources for performing tensor parallelism as (1,4), (2,2), (4,1). If there are 4 resources for performing pipeline parallelism, the number of remaining resources is 3, and thus the electronic device (100) can obtain the number of resources for performing data parallelism and the number of resources for performing tensor parallelism as (1,3), (3,1). If there are 6 resources for performing pipeline parallelism, the number of remaining resources is 2, and thus the electronic device (100) can obtain the number of resources for performing data parallelism and the number of resources for performing tensor parallelism as (2,1), (2,1). If there are 12 resources for performing pipeline parallelization, the number of resources remaining is 1, so the electronic device (100) can obtain the number of resources for performing data parallelization and the number of resources for performing tensor parallelization as (1,1).
[0096] In addition, the electronic device (100) can identify multiple resource combinations for performing multiple parallelization methods based on information about the number of resources for performing tensor parallelization, the number of resources for performing data parallelization, and the number of resources for performing pipeline parallelization. In one embodiment, the electronic device (100) can obtain resource combinations as shown in Table 2 below.
[0097] - Number of resources in the first device: 4 - Number of resources in the second device: 8 - Greatest common divisor of resources in the first and second devices: 4 - Resource combination (PP_size, DP_size, TP_size) (1, 3, 4) (1, 4, 3) (1, 6, 2) (1, 12, 1) (2, 2, 3) (2, 3, 2) (2, 6, 1) (3, 1, 4) (3, 2, 2) (3, 4, 1) (4, 3, 1) (4, 1, 3) (6, 2, 1) (6, 1, 2) (12, 1, 1)
[0098] In the above-described Table 2, PP_size may refer to the number of resources for performing pipeline parallelism, DP_size may refer to the number of resources for performing data parallelism, and TP_size may refer to the number of resources for performing tensor parallelism. In the same manner as the above-described table, when multiple devices are divided into multiple virtual devices with the same number of resources and the number of resources included in the multiple virtual devices is equalized, it is possible to prevent resources included in different devices from being included in order to perform tensor parallelism.
[0099] That is, the electronic device (100) can obtain information on at least one of the nine resource combinations included in Table 1 and the 15 resource combinations included in Table 2 through step S220.
[0100] Referring back to FIG. 2, the electronic device (100) acquires first processing rate information for resources contained in multiple devices for each of the multiple layers (S230). Specifically, the electronic device (100) can perform a profile for each of the multiple layers using the resources contained in the multiple devices, thereby acquiring first processing rate information for resources contained in the multiple devices for each of the multiple layers. For example, as illustrated in FIG. 3A, when a neural network model having a total of 12 layers is distributedly trained using a first type of resource (e.g., a GPU having a first performance) included in a first device (310) and a second type of resource (a GPU having a second performance, wherein the second performance may be superior to the first performance) included in a second device (320), the electronic device (100) can obtain information on the execution time and memory usage information for each layer processed after inputting data to each of the 12 layers using the first type of resource, and can obtain information on the execution time and memory usage information for each layer processed after inputting data to each of the 12 layers using the second type of resource. Thereby, the electronic device (100) can obtain first processing rate information of resources included in a plurality of devices for each of the plurality of layers.
[0101] The electronic device (100) obtains stage information for dividing layers included in the neural network model for each combination of the plurality of resources based on the first processing rate information of the resources included in the plurality of devices (S240). At this time, the electronic device (100) obtains stage information for dividing layers included in the neural network model based on the first processing rate information of the resources included in the plurality of devices, and performs learning using the resources included in the plurality of devices for the obtained plurality of stages, thereby dividing layers so that the deviation in execution time for performing the plurality of stages is the smallest. For example, in a situation where a neural network model including 12 layers is distributedly learned, as illustrated on the left side of FIG. 5, the electronic device (100) may obtain a first stage (510-1) and a second stage (510-2) by dividing the layers into groups of 6 each. The electronic device (100) can obtain a first execution time required as a result of performing learning for the first stage (510-1) using a first resource among the plurality of resources based on the first processing rate information, and a second execution time required as a result of performing learning for the second stage (510-2) using a second resource among the plurality of resources. In addition, the electronic device (100) can identify a first deviation between the first execution time and the second execution time. In addition, as illustrated on the right side of FIG. 5 for the same neural network model, the electronic device (100) can obtain a third stage (510-3) and a fourth stage (510-4) by dividing the layers into 5 and 7, respectively. The electronic device (100) can obtain a third execution time required as a result of performing learning for the third stage (510-3) using a first resource among a plurality of resources based on the first processing rate information, and a fourth execution time required as a result of performing learning for the fourth stage (510-4) using a second resource among a plurality of resources.And, the electronic device (100) can identify the second deviation of the third execution time and the fourth execution time. In this way, the electronic device (100) can obtain information on a plurality of stage combinations by dividing a plurality of layers included in the neural network model in a plurality of ways, and obtain deviations in the execution times required as a result of learning the stage combinations. And, the electronic device (100) can obtain stage information by dividing the layers so as to correspond to the stage combination with the smallest deviation in execution time among the plurality of stage combinations. For example, as illustrated on the right side of FIG. 5, if the deviation in execution time is the smallest when divided into a combination of the third stage (510-3) and the fourth stage (510-4), the electronic device (100) can obtain stage information corresponding to the first resource combination having two resources for performing pipeline parallelization among the plurality of resource combinations, among the combinations of the third stage (510-3) and the fourth stage (510-4).
[0102] The electronic device (100) can repeatedly perform step S240 on the plurality of resource combinations acquired in step S220 to obtain stage information for each of the plurality of resource combinations. For example, the electronic device (100) can obtain stage information when the number of resources for performing pipeline parallelization is 2 and stage information when the number of resources for performing pipeline parallelization is 4. In this case, the electronic device (100) may not obtain stage information because there is no need to divide the layers when the number of resources for performing pipeline parallelization is 1.
[0103] Meanwhile, when performing step S240, the electronic device (100) can obtain stage information that divides the layers included in the neural network model by batch size. Specifically, the first processing rate information may include time required and memory usage information for each batch size. Accordingly, the electronic device (100) can use the first processing rate information to obtain stage information that divides the layers included in the neural network model by batch size for multiple resource combinations.
[0104] The electronic device (100) can acquire second processing rate information for each of the plurality of resource combinations by performing distributed learning through multiple parallelization methods based on stage information (S250). That is, when distributed learning is performed based on the plurality of resource combinations acquired in step S220 and the stage information acquired in step S240, the electronic device (100) can acquire second processing rate information for each of the plurality of resource combinations.
[0105] At this time, the electronic device (100) can perform distributed learning for each of a plurality of resource combinations through a plurality of parallelization methods based on the stage information acquired for each batch size, thereby obtaining second processing rate information for each of the plurality of resource combinations. That is, the electronic device (100) can obtain second processing rate information based on stage information for each of a plurality of batch sizes for each of the plurality of resource combinations. For example, the electronic device (100) can obtain second processing rate information based on stage information acquired for each of the plurality of resource combinations when the first batch size is used, and can obtain second processing rate information based on stage information acquired for each of the plurality of resource combinations when the second batch size is used.
[0106] The electronic device (100) can obtain information on an optimal distributed learning strategy based on the second processing rate information for multiple resource combinations. That is, the electronic device (100) can obtain information on an optimal distributed learning strategy that enables the fastest distributed learning based on the second processing rate information for each of the multiple resource combinations obtained in step S250. At this time, the optimal distributed learning strategy may include information on an optimal resource combination, optimal stage information, and optimal batch size. For example, when performing distributed learning on a neural network model using multiple devices including a total of 16 resources, the electronic device (100) can identify a resource combination including four resources for performing pipeline parallelization, two resources for performing data parallelization, and four resources for performing tensor parallelization among the multiple resource combinations, as illustrated in FIG. 6 . Specifically, as illustrated in FIG. 6 , the electronic device (100) can identify resources for training the first neural network model (610) and the second neural network model (620) as resources for performing data parallelization. In addition, the electronic device (100) can identify resources for learning four stages (610-1 to 610-4) included in the first neural network model (610) and four stages (610-1 to 610-4) included in the second neural network model (620) as resources for performing pipeline parallelization. In addition, the electronic device (100) can identify resources for learning four tensors (611-1 to 611-4) included in one stage of the first neural network model (610) and four tensors (621-1 to 621-4) included in one stage of the second neural network model (620) as resources for performing tensor parallelization.
[0107] Additionally, the electronic device (100) can configure the batch size to process the maximum number of samples per hour for the entire pipeline, taking into account the predicted execution time for each pipeline. This allows the electronic device (100) to obtain information on an optimal distributed learning strategy, including information on the optimal batch size input to each neural network model.
[0108]
[0109] The artificial intelligence-related functions according to the present disclosure are operated through the processor and memory of the electronic device (100).
[0110] The processor may be composed of one or more processors. In this case, the one or more processors may include at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an NPU (Neural Processing Unit), but is not limited to the examples of the processors described above.
[0111] CPUs are general-purpose processors capable of performing not only general calculations but also artificial intelligence calculations. Their multi-layered cache structure allows for the efficient execution of complex programs. CPUs are advantageous for serial processing, enabling organic linking of previous and subsequent calculation results through sequential calculations. General-purpose processors are not limited to the examples described above, except where specifically identified as CPUs.
[0112] A GPU is a processor designed for large-scale computations, such as floating-point operations used in graphics processing. It integrates a large number of cores to perform large-scale computations in parallel. In particular, GPUs may be advantageous over CPUs in parallel processing methods, such as convolution operations. Furthermore, GPUs can be used as coprocessors to supplement the functions of CPUs. Processors for large-scale computations are not limited to the examples described above, except in cases where they are specifically referred to as GPUs.
[0113] An NPU is a processor specialized in artificial intelligence computation using artificial neural networks, and each layer of the artificial neural network can be implemented in hardware (e.g., silicon). Since NPUs are designed specifically according to the company's specifications, they have less freedom than CPUs or GPUs, but can efficiently process the AI computations requested by the company. Meanwhile, as a processor specialized in AI computation, an NPU can be implemented in various forms, such as a Tensor Processing Unit (TPU), an Intelligence Processing Unit (IPU), or a Vision Processing Unit (VPU). Except as specifically designated as an NPU, an AI processor is not limited to the examples described above.
[0114] Additionally, one or more processors may be implemented as a System on Chip (SoC). In this case, in addition to one or more processors, the SoC may further include memory and a network interface, such as a bus, for data communication between the processor and the memory.
[0115] When a System on Chip (SoC) included in an electronic device includes multiple processors, the electronic device may use some of the multiple processors to perform operations related to artificial intelligence (e.g., operations related to learning or inference of an artificial intelligence model). For example, the electronic device may use at least one of a GPU, NPU, VPU, TPU, or hardware accelerator specialized in artificial intelligence operations, such as convolution operations or matrix multiplication operations, among the multiple processors to perform operations related to artificial intelligence. However, this is merely an example, and it is of course possible to process operations related to artificial intelligence using a CPU or other general-purpose processor.
[0116] Additionally, electronic devices can perform computations related to artificial intelligence (AI) functions by utilizing multiple cores (e.g., dual cores, quad cores, etc.) contained within a single processor. In particular, electronic devices can perform AI operations, such as convolution operations and matrix multiplication operations, in parallel by utilizing the multiple cores contained within the processor.
[0117] One or more processors are controlled to process input data according to predefined operating rules or artificial intelligence models stored in memory. The predefined operating rules or artificial intelligence models are characterized by being created through learning.
[0118] Here, "created through learning" means that a predefined set of behavioral rules or an AI model with desired characteristics is created by applying a learning algorithm to a large number of learning data. This learning may be performed on the device itself, where the AI according to the present disclosure is implemented, or through a separate server / system.
[0119] An artificial intelligence model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs its operation through the operation result of the previous layer and at least one defined operation. Examples of neural networks include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, and a transformer. The neural networks in the present disclosure are not limited to the above-described examples unless otherwise specified.
[0120] A learning algorithm is a method for training a target device (e.g., a robot) using a large amount of learning data, enabling the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Unless otherwise specified, the learning algorithms in this disclosure are not limited to the aforementioned examples.
[0121] Meanwhile, the methods according to various embodiments of the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0122] The methods according to various embodiments of the present disclosure may be implemented as software including commands stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call commands stored in the storage medium and operate according to the called commands, and may include an electronic device (e.g., a TV) according to the disclosed embodiments.
[0123] Meanwhile, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0124] When the above instruction is executed by the processor, the processor may perform the function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter.
[0125] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. In an electronic device for distributed learning of a neural network model through multiple parallel methods, communication interface; A neural network model including multiple layers and a memory storing at least one instruction; and A processor for controlling the electronic device; The above processor, by executing at least one instruction, Obtaining information about resources included in multiple devices for distributed learning of the neural network model through multiple parallelism methods, wherein the multiple parallelism methods include data parallelism, pipeline parallelism, and tensor parallelism. Identifying a plurality of resource combinations for performing the plurality of parallelization methods based on information about the resources included in the plurality of devices; Obtain first processing rate information of resources included in the plurality of devices for each of the plurality of layers, Based on the first processing rate information of the resources included in the plurality of devices, stage information for partitioning the layers included in the neural network model is obtained for each combination of the plurality of resources. Based on the stage information, distributed learning is performed for each of the plurality of resource combinations through the plurality of parallelization methods to obtain second processing rate information for the plurality of resource combinations. An electronic device that obtains information about an optimal distributed learning strategy based on the second processing rate information for the plurality of resource combinations.
2. In paragraph 1, The above processor, Grouping the resources included in each of the plurality of devices based on the number of resources included in the plurality of devices, An electronic device for identifying a plurality of resource combinations for performing distributed learning through the plurality of parallelization methods based on the grouped resources.
3. In paragraph 2, The above processor, Obtain the common divisor of the number of resources included in the above multiple devices, Based on the common divisor obtained above, the resources included in each of the plurality of devices are grouped into resources for performing the tensor parallelization, Obtain information about the number of divisible resource groups based on the obtained common divisor and the number of devices, Obtain information about the number of resources for performing the data parallelization and the number of resources for performing the pipeline parallelization based on a divisor of the number of the divisible resource groups, An electronic device that identifies a plurality of resource combinations for performing the plurality of parallelization methods based on information about a resource group for performing the tensor parallelization, a number of resources for performing the data parallelization, and a number of resources for performing the pipeline parallelization.
4. In paragraph 1, The above processor, Dividing the plurality of devices into a plurality of virtual devices so that the plurality of devices can have the same number of resources based on the number of resources included in the plurality of devices, An electronic device that identifies a plurality of resource combinations for performing the plurality of parallelization methods based on resources included in the plurality of virtual devices.
5. In paragraph 4, The above processor, Obtain the greatest common divisor of the resources contained in the above multiple devices, Divide the virtual devices into the plurality of devices based on the greatest common divisor and obtain the plurality of virtual devices, Obtain the common divisor of the resources included in the above multiple virtual devices, Obtain information about the number of resources for performing pipeline parallelization based on the above common divisor, Obtain information about the number of resources for performing data parallelism and the number of resources for performing tensor parallelism based on the remaining resources excluding the resources for performing the pipeline parallelism among the resources included in the above plurality of virtual devices, An electronic device that identifies a plurality of resource combinations for performing the plurality of parallelization methods based on information about the number of resources for performing the tensor parallelization, the number of resources for performing the data parallelization, and the number of resources for performing the pipeline parallelization.
6. In paragraph 1, The above processor, By performing a profile for the plurality of layers using the resources included in the plurality of devices, first processing rate information of the resources included in the plurality of devices is obtained for each of the plurality of layers. The above first processing rate information is, An electronic device comprising at least one of execution time and memory usage information.
7. In paragraph 6, The above processor, An electronic device that obtains stage information by partitioning the layers included in the neural network model based on first processing rate information of resources included in the plurality of devices, and performing learning using the resources included in the plurality of devices for the plurality of stages so that the deviation in execution time for performing the plurality of stages is the smallest.
8. In paragraph 1, The above processor, Based on the first processing rate information of the resources included in the above multiple devices, stage information is obtained by partitioning the layers included in the above neural network model by batch size. Based on the stage information obtained for each of the above batch sizes, distributed learning is performed for each of the above multiple resource combinations through the above multiple parallelization methods to obtain second processing rate information for the above multiple resource combinations. An electronic device that obtains information on an optimal distributed learning strategy that enables the fastest distributed learning based on second processing rate information for the above combinations of multiple resources.
9. In paragraph 1, The above optimal distributed learning strategy is, An electronic device comprising information about optimal resource combinations, optimal stage information, and optimal batch sizes.
10. A method for controlling an electronic device for distributed learning of a neural network model through multiple parallelization methods, A step of obtaining information about resources included in a plurality of devices for distributed learning of the neural network model through a plurality of parallelism methods, wherein the plurality of parallelism methods include data parallelism, pipeline parallelism, and tensor parallelism; A step of identifying a plurality of resource combinations for performing the plurality of parallelization methods based on information about resources included in the plurality of devices; A step of obtaining first processing rate information of resources included in the plurality of devices for each of the plurality of layers; A step of obtaining stage information for partitioning the layers included in the neural network model for each combination of the plurality of resources based on first processing rate information of the resources included in the plurality of devices; A step of performing distributed learning for each of the plurality of resource combinations through the plurality of parallelization methods based on the stage information to obtain second processing rate information for the plurality of resource combinations; and A control method comprising: a step of obtaining information on an optimal distributed learning strategy based on the second processing rate information for the plurality of resource combinations.
11. In paragraph 10, The above identifying step is, A control method for grouping resources included in each of the plurality of devices based on the number of resources included in the plurality of devices, and identifying a plurality of resource combinations for performing distributed learning through the plurality of parallelization methods based on the grouped resources.
12. In paragraph 11, The above identifying step is, A step of obtaining a common divisor of the number of resources included in the above plurality of devices; A step of grouping resources included in each of the plurality of devices into resources for performing the tensor parallelization based on the obtained common divisor; A step of obtaining information on the number of divisible resource groups based on the obtained common divisor and the number of devices; A step of obtaining information about the number of resources for performing the data parallelization and the number of resources for performing the pipeline parallelization based on a divisor of the number of the divisible resource groups; and A control method comprising: a step of identifying a plurality of resource combinations for performing the plurality of parallelization methods based on information about a resource group for performing the tensor parallelization, a number of resources for performing the data parallelization, and a number of resources for performing the pipeline parallelization.
13. In paragraph 10, The above identifying step is, A control method for dividing the plurality of devices into a plurality of virtual devices based on the number of resources included in the plurality of devices so that the plurality of devices can have the same number of resources, and identifying a plurality of resource combinations for performing the plurality of parallelization methods based on the resources included in the plurality of virtual devices.
14. In paragraph 13, The above identifying step is, A step of obtaining the greatest common divisor of resources included in the above plurality of devices; A step of dividing a virtual device into each of the plurality of devices based on the greatest common divisor and obtaining the plurality of virtual devices; A step of obtaining a common divisor of resources included in the plurality of virtual devices; A step of obtaining information on the number of resources for performing pipeline parallelization based on the above common divisor; A step of obtaining information about the number of resources for performing data parallelism and the number of resources for performing tensor parallelism based on the remaining resources excluding the resources for performing the pipeline parallelism among the resources included in the plurality of virtual devices; and A control method comprising: a step of identifying a plurality of resource combinations for performing the plurality of parallelization methods based on information about the number of resources for performing the tensor parallelization, the number of resources for performing the data parallelization, and the number of resources for performing the pipeline parallelization.
15. In paragraph 10, The step of obtaining the above first processing rate information is: By performing a profile for the plurality of layers using the resources included in the plurality of devices, first processing rate information of the resources included in the plurality of devices is obtained for each of the plurality of layers. The above first processing rate information is, A control method including at least one of execution time and memory usage information.
Citation Information
Patent Citations
Methods, systems, and devices for distributed training task scheduling for intelligent computing
CN115248728B
reflective member driving device
KR1020220136831A
System for evaluating SLA
KR1020230134358A