Method and apparatus for controlling inference task execution by split inference of artificial neural network
Through the split inference and update strategy methods of artificial neural network, the problem of split execution of inference tasks between multiple devices is solved, and efficient and highly adaptable task execution is achieved.
Patent Information
- Application Number
- CN202380079683.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-09
- Filing Date
- 2023-08-31
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively split the inference tasks of the artificial intelligence model between multiple devices, and it is difficult for users to directly correct inappropriate task execution strategies.
The execution of inference tasks is controlled through split inference by artificial neural networks, and the update strategy method is used to determine the device used for split inference. The method includes selecting a policy from multiple task execution strategies based on the requirements and correction index of the inference task, and updating the correction index and strategy based on the failure results of split inference.
The reasoning task of efficiently splitting the execution of artificial intelligence models between multiple devices is realized, which improves the adaptability and efficiency of task execution and reduces the complexity of users when correcting strategies.
Smart Images

Figure CN120226022A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method and apparatus for controlling the execution of an inference task through split inference of an artificial neural network. More specifically, the present disclosure relates to a method of an update strategy for determining an apparatus for performing an inference task through split inference. Background Art
[0002] With the advancement of artificial intelligence technology, the functions of artificial intelligence models are being used in various apparatuses. However, artificial intelligence models require a large amount of computation. In a case where it is difficult to perform all operations for artificial intelligence on a single apparatus due to differences in hardware performance and security issues, inference can be performed by distributing the inference across multiple apparatuses. Performing the inference task of an artificial intelligence model by distributing the inference task of the artificial intelligence model across multiple apparatuses may be referred to as split inference.
[0003] In order to perform split inference using multiple apparatuses, a method for determining which apparatus to use for split inference is required. An electronic apparatus can determine an apparatus for split inference by using a task execution strategy including a method for determining an apparatus. Depending on the current state of the apparatus and the type of task, the task execution strategy may not be appropriate. However, it is generally difficult for a user to directly correct an inappropriate task execution strategy. Therefore, it may be necessary to modify the task execution strategy to adapt to the user's execution environment. Summary of the Invention
[0004] According to an embodiment of the present disclosure, there is provided a method for controlling the execution of an inference task through split inference of an artificial neural network. The method may include determining one strategy from a plurality of task execution strategies based on one of requirements of the inference task and a correction index, wherein the correction index is determined based on a failure rate of each task execution strategy. The method may include determining one or more apparatuses for performing split inference of the artificial neural network based on the strategy. The method may include updating the correction index corresponding to the strategy based on a result of whether the split inference performed by the one or more apparatuses fails. The method may include updating the strategy by using an execution record of split inference obtained from the one or more apparatuses, wherein the execution record includes information on a cause of failure of the split inference. Each task execution strategy among the plurality of task execution strategies may include at least one of a priority of considering one or more apparatus conditions for selecting an apparatus for performing split inference and the number of apparatuses for split inference.
[0005] According to an embodiment of the present disclosure, there is provided a computer-readable recording medium having recorded thereon a program for performing the method on a computer.
[0006] According to an embodiment of the present disclosure, there is provided an apparatus for controlling the execution of an inference task through split inference of an artificial neural network. The apparatus may include a memory and at least one processor, the memory including one or more instructions. The at least one processor may be configured to determine a policy from a plurality of task execution policies based on the requirements of the inference task and a correction index indicating the failure rate of each task execution policy. The at least one processor may be configured to determine one or more apparatuses for performing split inference of the artificial neural network based on the policy. The at least one processor may be configured to update the correction index corresponding to the policy based on the result of whether the split inference performed by the one or more apparatuses fails. The at least one processor may be configured to update the policy by using the execution record of the split inference obtained from the one or more apparatuses. The execution record may include information about the cause of failure of the split inference. Each task execution policy among the plurality of task execution policies may include at least one of a priority for considering one or more apparatus conditions for selecting an apparatus for performing split inference and the number of apparatuses for split inference. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 FIG. is a diagram illustrating a method of controlling the execution of split inference by an electronic device according to an embodiment of the present disclosure by using one or more apparatuses.
[0008] Figure 2 FIG. is a flowchart of a method of controlling split inference according to an embodiment of the present disclosure.
[0009] Figure 3 FIG. is a block diagram illustrating a process of an electronic device controlling split inference according to an embodiment of the present disclosure.
[0010] Figure 4 FIG. is a diagram for describing a task execution policy according to an embodiment of the present disclosure.
[0011] Figure 5 FIG. is a block diagram illustrating a process of an electronic device determining a task execution policy according to an embodiment of the present disclosure.
[0012] Figure 6 FIG. is a block diagram illustrating a process of an electronic device determining an apparatus according to an embodiment of the present disclosure.
[0013] Figure 7 FIG. is a block diagram illustrating a process of an electronic device determining an inference ratio according to an embodiment of the present disclosure.
[0014] Figure 8 FIG. is a block diagram illustrating a process of an electronic device inferring apparatus state information according to an embodiment of the present disclosure.
[0015] Figure 9It is a diagram for describing a method by which an electronic device controls split inference according to an inference ratio according to an embodiment of the present disclosure.
[0016] Figure 10 It is a diagram showing a method of updating a correction index performed by an electronic device according to an embodiment of the present disclosure.
[0017] Figure 11 It is a diagram showing a method of updating a policy performed by an electronic device according to an embodiment of the present disclosure.
[0018] Figure 12a It is a diagram showing a method of training an artificial intelligence model for determining an inference task policy performed by an electronic device according to an embodiment of the present disclosure.
[0019] Figure 12b It is a diagram showing a method of training an artificial intelligence model for determining a device for performing an inference task performed by an electronic device according to an embodiment of the present disclosure.
[0020] Figure 13 It is a diagram showing a method of generating a task execution policy for a new task performed by an electronic device according to an embodiment of the present disclosure.
[0021] Figure 14 It is a diagram showing a method of controlling split inference in an indoor environment performed by an electronic device according to an embodiment of the present disclosure.
[0022] Figure 15 It is a block diagram showing the configuration of an electronic device according to an embodiment of the present disclosure. Detailed Description
[0023] In the present disclosure, the expression "at least one of a, b, or c" may refer to "a", "b", "c", "a and b", "a and c", "b and c", "all of a, b, and c", or a variant thereof.
[0024] The terms used in the present disclosure are selected from commonly used terms as much as possible while considering the functions of the present disclosure, but these may vary according to the intentions of engineers working in the field, precedents, the emergence of new technologies, etc. Additionally, in certain cases, there are terms arbitrarily selected by the applicant, and in such cases, their meanings can be understood through the corresponding explanatory parts. Therefore, the terms used in the present disclosure should be defined based on the meanings of the terms and the overall details of the present disclosure, rather than simply based on the names of the terms.
[0025] In this disclosure, unless the context clearly indicates otherwise, singular expressions may include plural expressions. Terms used in this disclosure that include ordinal numbers (such as "first" or "second") may be used to describe various components, but the components should not be limited by the terms. These terms are only used to distinguish one component from another.
[0026] Unless otherwise specifically stated, when it is said in this disclosure that a part "includes" a specific component, this does not exclude other components, but may include other components. In this disclosure, terms such as "unit" and "module" represent units that process at least one function or operation, which may be implemented by hardware or software, or by a combination of hardware and software.
[0027] Depending on the context, the expression "configured to" used in this disclosure may be interchangeably used with, for example, "suitable for", "capable of...", "designed to", "adapted to", "manufactured to", or "able to". The term "configured to" may not necessarily mean "specially designed in hardware". Optionally, in some contexts, the expression "a system configured to..." may refer to the system "being able to" do something in combination with other devices or components. For example, the phrase "a processor configured (or set) to execute A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for executing those operations, or a general-purpose processor (e.g., a CPU or an application processor) that can execute those operations by running one or more software programs stored in a memory.
[0028] When a component is referred to as "connected" or "coupled" to another component in this disclosure, it should be understood that the component may be directly connected or directly coupled to the other component, but unless otherwise specifically stated, it should also be understood that the component may be connected or coupled via yet another component therebetween.
[0029] When describing this disclosure, descriptions of technical details that are well known in the technical field to which this disclosure pertains and are not directly relevant to this disclosure may be omitted. This is to more clearly convey the gist of this disclosure by omitting unnecessary explanations without obscuring it. To clearly describe this disclosure in the drawings, parts that are not relevant to the description are omitted, and similar components are given similar reference numerals throughout the specification. The size of each component does not fully reflect its actual size. In each drawing, the same or corresponding components are given the same reference numerals.
[0030] Advantages and features of the present disclosure, and methods for realizing them, will become apparent by referring to embodiments described in detail below in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and can be implemented in various different forms. The disclosed embodiments are provided so that the present disclosure will be comprehensive and complete, and will fully convey the scope of the present disclosure to those skilled in the art to which the present disclosure pertains. Embodiments of the present disclosure may be defined by the claims.
[0031] In the present disclosure, each block of the flowchart can be executed by computer program instructions and combinations of the flowchart. The computer program instructions can be embedded in a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, and when the instructions are executed by the processor of the computer or other programmable data processing devices, methods for performing the functions described in the flowchart blocks can be created. The computer program instructions can also be stored in a computer-usable memory or a computer-readable memory, which can instruct the computer or other programmable data processing means to implement functions in a specific manner, and the instructions stored in the computer-usable memory or the computer-readable memory can also produce an article of manufacture including an instruction method for performing the functions described in the flowchart blocks. The computer program instructions can also be embedded in a computer or other programmable data processing device.
[0032] In the present disclosure, each block of the flowchart can represent a module, a code segment, or a portion of code including one or more executable instructions for performing a specified logical function. In an embodiment, the functions mentioned in the block may not occur in sequence. For example, two consecutive blocks shown may be executed substantially simultaneously, or may be executed in the reverse order according to their functions.
[0033] The term "~ unit" used in embodiments of the present disclosure may represent a software or hardware component, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC), and the "~ unit" can perform a specific role. At the same time, the "~ unit" is not limited to software or hardware. The "~ unit" can be configured to reside on an addressable storage medium and can be configured to cause one or more processors to reproduce the same content. In an embodiment, the "~ unit" can include components (such as software components, object-oriented software components, class components, and task components), processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided by specific components or specific "parts" can be combined to reduce their number or divided into additional components. Additionally, in an embodiment, the "~ unit" can include one or more processors.
[0034] The artificial intelligence-related functions according to the present disclosure are operated by a processor and a memory. The processor may consist of one or more processors. The one or more processors may be a general-purpose processor (such as a CPU, an AP, or a digital signal processor (DSP)), a pure graphics processor (such as a GPU or a vision processing unit (VPU)), or an artificial intelligence dedicated processor (such as an NPU). The one or more processors are controlled according to predefined operation rules or artificial intelligence models stored in the memory to process input data. Optionally, if the one or more processors are artificial intelligence dedicated processors, the artificial intelligence dedicated processors may be designed to have a hardware structure dedicated to processing a specific AI model.
[0035] The predefined operation rules or artificial intelligence models are characterized in that they are generated through learning. Here, being generated through learning means using a large amount of learning data and a learning algorithm to learn a basic artificial intelligence model, so as to generate predefined operation rules or artificial intelligence models that are set to perform desired characteristics (or purposes). Such learning may be performed on the device itself that executes the artificial intelligence according to the present disclosure, or may be performed by a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0036] The artificial intelligence model may consist of multiple neural network layers. Each neural network layer in the multiple neural network layers has multiple weight values, and performs neural network operations through the operation between the operation result of the previous layer and the multiple weight values. The multiple weight values of the multiple neural network layers may be optimized through the learning result of the artificial intelligence model. For example, the multiple weight values may be updated such that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), for example, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or a deep Q network, but is not limited to the above examples.
[0037] In the method for controlling the execution of an inference task through split inference of an artificial neural network by an electronic device according to the present disclosure, an artificial intelligence model can be used as a method for an inference or prediction task execution strategy or a device for performing split inference. The artificial intelligence model can be generated through learning. Here, generating through learning means training a basic artificial intelligence model using a large amount of learning data with a learning algorithm, thereby generating predefined operation rules or an artificial intelligence model set to perform desired characteristics (or purposes). The artificial intelligence model can be composed of multiple neural network layers. Each neural network layer among the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weight values.
[0038] Inference prediction is a technique for logically reasoning and predicting by judging information, including knowledge / probability-based reasoning, optimization prediction, preference-based planning, and recommendation.
[0039] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure can be implemented in many different forms and is not limited to the embodiments described herein. To clearly describe the present disclosure in the drawings, parts irrelevant to the explanation are omitted, and similar parts are given the same reference numerals throughout the specification. Additionally, the reference symbols used in each drawing are only intended to describe each drawing, and different reference symbols used in different drawings do not indicate different elements. The present disclosure will be described in detail below with reference to the accompanying drawings.
[0040] In the present disclosure, split inference may refer to using an artificial neural network to split an inference task and execute the inference task on multiple devices. For example, when performing an inference task called object detection, the calculations of the artificial neural network can be split and executed on multiple devices. The multiple devices can process the calculations on the artificial neural network in at least one of a serial manner or a parallel manner. For example, referring to Figure 9 , split inference can be processed such that the first device 910 performs operations on a specific layer of the artificial neural network, and the second device 920 performs operations on the next layer of the artificial neural network based on the operation results of the first device 910. Alternatively, split inference can be processed such that the first device 910 and the second device 920 respectively perform operations on the same layer of the artificial neural network.
[0041] In the present disclosure, a task execution strategy may refer to rules or guidelines for determining a device for performing split inference. The task execution strategy may be referred to as an execution strategy, a strategy, or an inference task strategy. According to an embodiment of the present disclosure, the task execution strategy may include at least one of device conditions or the number of devices for split inference. According to an embodiment of the present disclosure, the device conditions may include at least one of a processor type of the device, a remaining memory capacity, a heat level, a remaining battery capacity, and the number of running applications. According to an embodiment of the present disclosure, the task execution strategy may include a priority of the device conditions. For example, the task execution strategy may include information about the priority of the device conditions.
[0042] In the present disclosure, an available device may refer to a device capable of performing split inference by being connected to an electronic device that controls the execution of split inference. The electronic device may be controlled to identify at least some of the plurality of available devices and perform split inference by using the identified devices. The available device may be used interchangeably with the available equipment. According to an embodiment of the present disclosure, the available device may be a device that is connected to the electronic device via a network and exists within a certain distance from the electronic device.
[0043] Figure 1 is a diagram illustrating a method of controlling the execution of split inference by using one or more devices, performed by an electronic device according to an embodiment of the present disclosure.
[0044] Referring to Figure 1 , the electronic device 100 may control the execution of an inference task by using at least some of the available devices 130.
[0045] According to an embodiment of the present disclosure, the electronic device 100 may include, but is not limited to, a server, a mobile phone, a robotic vacuum cleaner, a television, an augmented reality (AR) device, or a virtual reality (VR) device, and may be another device that performs an inference task by using an artificial neural network.
[0046] According to an embodiment of the present disclosure, it may be necessary for the electronic device 100 to perform a task related to an inference operation by using an artificial neural network. For example, the electronic device 100 may need to perform functions such as estimating a user's pose by using a camera, generating a spatial map of a space within a home, detecting an object included in an image, or predicting a risk regarding a detected situation. According to an embodiment of the present disclosure, the electronic device 100 may share an inference task with another device or control another device to perform an inference task.
[0047] According to an embodiment of the present disclosure, the electronic device 100 may obtain a plurality of task execution policies from the task execution policy DB 110. The task execution policy DB 110 may be stored in the memory of the electronic device 100 or in an external server. The task execution policy may include at least one of a priority considering one or more device conditions for selecting a device for performing split inference and the number of devices for split inference. Refer to Figure 4 A task execution policy according to an embodiment of the present disclosure will be described in more detail.
[0048] According to an embodiment of the present disclosure, the electronic device 100 may determine the priorities of a plurality of task execution policies. The electronic device 100 may identify at least one of the requirements of the inference task and a correction index representing the failure rate of each task execution policy. The electronic device 100 may prioritize the plurality of task execution policies based on at least one of the requirements of the inference task and the correction index. The electronic device 100 may determine one or more task execution policies based on the priorities. Refer to Figure 5 A method for the electronic device 100 to determine a task execution policy according to an embodiment of the present disclosure will be described in more detail.
[0049] According to an embodiment of the present disclosure, the electronic device 100 may identify available devices 130. For example, the electronic device 100 may determine, as available devices 130, devices among the devices connected to the electronic device 100 via a network that consent or do not refuse to perform split inference.
[0050] According to an embodiment of the present disclosure, the electronic device 100 may determine, based on the determined task execution policy, one or more devices for performing split inference of an artificial neural network among the available devices 130. The electronic device 100 may determine a first device group 140 including one or more devices 145 based on a first-priority task execution policy. The electronic device 100 may determine a second device group 150 including one or more devices 155 based on a second-priority task execution policy. The electronic device 100 may determine the first device group 140 and the second device group 150 before performing split inference, but is not limited thereto, and the electronic device 100 may determine the second device group 150 when split inference using one or more devices 145 included in the first device group 140 fails. For ease of description, the second device group 150 is shown, but is not limited thereto, and lower-priority device groups and one or more devices may be determined based on lower-priority policies. Refer to Figure 6 A process in which the electronic device 100 determines devices based on a policy according to an embodiment of the present disclosure will be described in more detail.
[0051] According to an embodiment of the present disclosure, the electronic device 100 may be controlled to perform split inference using one or more devices 145 included in the first device group 140. The electronic device 100 may obtain information about whether the split inference fails and the execution record of the split inference from one or more devices 145. The electronic device 100 may store the information about the execution record of the split inference in the failure log DB 120. The information about the execution record of the split inference may include information about the cause of failure of the split inference.
[0052] If the split inference fails, the electronic device 100 may control the execution of the split inference by using one or more devices 155 included in the second device group 150. The electronic device 100 may be controlled to perform split inference using one or more devices in a lower priority device group until the split inference is successful. Reference will be made to Figures 7 to 9 A control process of performing split inference by using one or more devices executed by the electronic device 100 according to an embodiment of the present disclosure will be described in more detail.
[0053] According to an embodiment of the present disclosure, the electronic device 100 may update a correction index corresponding to a task execution policy based on the result of whether the split inference fails. For example, if the split inference fails, the electronic device 100 may decrease the value of the correction index for the task execution policy, thereby determining a device for split inference. The correction index may be stored in the task execution policy DB 110 corresponding to the task execution policy, or may be stored in the electronic device 100. Reference will be made to Figure 10 A method of updating a correction index by an electronic device according to an embodiment of the present disclosure will be described in more detail.
[0054] According to an embodiment of the present disclosure, the electronic device 100 may update the task execution policy by using the obtained execution record of the split inference. The information about the execution record of the split inference may include information about the cause of failure of the split inference. The electronic device 100 may update the task execution policy based on specific conditions. For example, the electronic device 100 may update the task execution policy during a period when split inference is not performed (e.g., during the early morning period). The electronic device 100 may obtain the execution record of the split inference from the failure log DB 120 and use the obtained execution record to update the task execution policy. Reference will be made to Figure 11 A method of updating a task execution policy executed by an electronic device according to an embodiment of the present disclosure will be described in more detail.
[0055] Figure 2 It is a flowchart of a method for controlling split inference according to an embodiment of the present disclosure.
[0056] Reference will be made to Figure 2, The method 200 for controlling split inference may start from operation S210. The method 200 for controlling split inference according to an embodiment of the present disclosure may be executed by the electronic device 100.
[0057] In operation S210, the electronic device 100 may determine one policy from multiple task execution policies based on at least one of the requirements of the inference task and a correction index determined based on the failure rate of each task execution policy. Each task execution policy among the multiple task execution policies may include at least one of a priority for considering one or more device conditions for selecting a device for performing split inference and the number of devices for split inference. The requirements may include at least one of the type, importance, maximum inference time, and required memory of the inference task. The device conditions may include at least one of the type of the device's processor, remaining memory capacity, heating level, remaining battery capacity, and the number of running applications.
[0058] According to an embodiment of the present disclosure, the electronic device 100 may determine a first score for the multiple task execution policies based on the requirements of the inference task, and determine a second score for the multiple task execution policies based on the first score and the correction index of the multiple task execution policies. The electronic device 100 may determine the execution policy with the highest second score as the policy. The electronic device 100 may determine the execution policy with the second-highest second score as the alternative policy. The electronic device 100 may determine one or more alternative devices for performing split inference based on the alternative policy. The electronic device 100 may control one or more alternative devices to perform split inference based on the failure of the split inference.
[0059] In operation S220, the electronic device 100 may determine one or more devices for performing split inference of the artificial neural network based on the policy.
[0060] In operation S230, the electronic device 100 may update the correction index corresponding to the policy based on the result of whether the split inference performed by one or more devices fails. The electronic device 100 may decrease the correction index based on the failure of the split inference. The electronic device 100 may maintain or increase the correction index based on the success of the split inference.
[0061] In operation S240, the electronic device 100 may update the policy by using the execution record of the split inference obtained from one or more devices. The execution record may contain information about the cause of the failure of the split inference. The electronic device 100 may identify the number of times of split inference failure for one or more device conditions based on the execution record of the split inference. The electronic device 100 may update the priority or update the number of devices for split inference so that the device conditions with a high number of split inference failures have a higher priority.
[0062] Figure 3It is a block diagram showing a process of controlling split inference performed by an electronic device according to an embodiment of the present disclosure.
[0063] Referring to Figure 3 , the electronic device 100 may include a policy determination unit 310, a device determination unit 320, a split inference control unit 330, and an update unit 340. The electronic device 100 may identify requirements of an inference task.
[0064] According to an embodiment of the present disclosure, the policy determination unit 310 may determine a first policy based on the requirements of the inference task. The policy determination unit 310 may determine priorities of multiple task execution policies based on the requirements of the inference task. The policy determination unit 310 may determine a higher priority for a task execution policy having a lower failure probability relative to the requirements of the inference task. As an example, the policy determination unit 310 may determine the highest priority task execution policy as the first policy. Additionally, the policy determination unit 310 may determine lower priority policies (e.g., a second policy, a third policy, etc.) according to the priorities.
[0065] According to an embodiment of the present disclosure, the device determination unit 320 may determine one or more devices for performing split inference based on the policy. For example, the device determination unit 320 may determine a first device group 140 including one or more devices based on the first policy. The device determination unit 320 may determine one or more devices satisfying the first policy among available devices. Similarly, the device determination unit 320 may determine a second device group 150 including one or more devices based on the second policy.
[0066] According to an embodiment of the present disclosure, the split inference control unit 330 may control to perform split inference using the one or more devices determined by the device determination unit 320. For example, the split inference control unit 330 may control one or more devices included in the first device group 140 to perform split inference. The split inference control unit 330 may send a command to the first device group 140 to perform split inference, and receive at least one of result information of the split inference and information about a reason for inference failure. The split inference control unit 330 may, in the case where the split inference of the first device group 140 fails, control one or more devices included in the second device group 150 to perform split inference.
[0067] According to an embodiment of the present disclosure, the update unit 340 may update a correction index or a task execution strategy based on at least one of result information regarding whether a received split inference fails and an execution record of the split inference. The update unit 340 may update a correction index corresponding to a task execution strategy based on the failure of the split inference. For example, if the split inference fails, the correction index may be updated such that the probability of being determined as a task execution strategy is reduced. The update unit 340 may update a task execution strategy based on an execution record of the split inference. For example, if there are many reasons for inference failure due to heat in the execution record of the split inference, the priority of heat may be increased.
[0068] Figure 4 is a diagram for describing a task execution strategy according to an embodiment of the present disclosure.
[0069] Referring to Figure 4 , there may be multiple different task execution strategies for an inference task. Multiple strategies 410, 420, 430 may be stored in the Figure 1 task execution strategy DB. When an initial strategy is generated by a user, strategy A 410 may be a task execution strategy generated by considering the importance of speed. Similarly, strategy B 420 may be a memory-centered task execution strategy, and strategy C 430 may be a hybrid task execution strategy that comprehensively considers multiple factors.
[0070] Strategy A 410 according to an embodiment of the present disclosure may include device conditions, priorities of the conditions, the number of devices, and a correction index. The device conditions may include the type of CPU or GPU, available memory, heat, battery capacity, and the number of running applications. The priority of the device conditions may indicate the importance of the device conditions. For example, the electronic device 100 may determine a device for performing split inference by determining a large weight for the first-priority device condition and a small weight for the low-priority device condition. For example, if an inference task requires a high-specification processor, the type of CPU or GPU may have a high priority. As an example, if an inference task uses a large amount of data, the memory capacity may have a relatively high priority. For example, if an inference task requires reliable results rather than speed, heat or the number of running applications may be prioritized. As an example, if it is expected that an inference task runs for a long time, the battery capacity may be a high priority.
[0071] Policy A 410 may include the type of CPU or GPU as a first-priority device condition, the remaining memory capacity as a second-priority device condition, the heat generation as a third-priority device condition, the battery capacity as a fourth-priority device condition, and the number of running applications as a fifth-priority device condition. The number of devices may indicate a condition for the number of devices to be used for split inference. Policy A 410 may perform split inference by using 4 or more devices and 6 or fewer devices. As the number of devices increases, the computational burden on each device may be reduced. However, in the case of performing split inference using a large number of devices, the number of available devices for performing other split inferences decreases. Therefore, an appropriate number of devices need to be used to perform split inference. The correction index may indicate the failure rate of the task execution policy. When split inference fails, the electronic device 100 may control the selection frequency of the task execution policy to decrease by reducing the correction index. The correction index of Policy A 410 is 8763 / 10000, but the correction index does not remain constant and may be repeatedly updated according to the execution of split inference.
[0072] Policy B 420 according to an embodiment of the present disclosure may include the remaining memory capacity as a first-priority device condition, the heat generation as a second-priority device condition, the number of running applications as a third-priority device condition, the type of CPU or GPU as a fourth-priority device condition, and the battery capacity as a fifth-priority device condition. Policy B 420 may perform split inference by using 6 or more devices and 10 or fewer devices. The correction index of Policy B 420 is 7253 / 10000, but the correction index may not be maintained and may be repeatedly updated according to the execution of split inference.
[0073] Policy C 430 according to an embodiment of the present disclosure may include the remaining memory capacity as a first-priority device condition, the type of CPU or GPU as a second-priority device condition, the battery capacity as a third-priority device condition, the heat generation as a fourth-priority device condition, and the number of running applications as a fifth-priority device condition. Policy C 430 may perform split inference by using 5 or more devices and 8 or fewer devices. The correction index of Policy C 430 is 9581 / 10000, but the correction index may not be maintained and may be repeatedly updated according to the execution of split inference.
[0074] Figure 5 is a block diagram showing a process in which an electronic device determines a task execution policy according to an embodiment of the present disclosure.
[0075] Refer to Figure 5, according to an embodiment of the present disclosure, the policy determination unit 310 may include a first score determination unit 510 and a second score determination unit 520. The first score determination unit 510 may determine a first score based on the requirements of the inference task. As an example, the requirements of the inference task may include, but are not limited to, the type, importance, required memory, and inference time of the task, and some requirements may be omitted or other requirements may be added.
[0076] The type of the inference task may refer to the type of the task for which split inference will be performed. The type of the inference task may include, but is not limited to, segmentation, classification, object detection, sentence completion, and spatial graph generation, and may include other inference tasks using artificial neural networks. The importance of the task may indicate the priority of the inference task. For example, an inference task with urgent importance may be executed before an inference task with basic importance. The importance of the task may include, but is not limited to, urgent, important, and basic, and may be represented as priority 1, priority 2, etc. The required memory may refer to the minimum memory or recommended memory required to execute the inference task. For example, the required memory may be represented in megabytes (MB), and 20000 MB may be required. The maximum inference time may refer to the time limit required to complete the inference. For example, the inference task may need to be completed within 3 seconds.
[0077] According to an embodiment of the present disclosure, the first score determination unit 510 may include an artificial neural network that takes the requirements of the inference task as input and outputs a first score of the task execution policy. The artificial neural network may be trained to output a larger value for a task execution policy that has a high success probability for the inference task relative to the requirements of the inference task. For example, based on the requirements of the inference task, the first score determination unit 510 may determine the first score of task execution policy A as 0.6211, the first score of task execution policy B as 0.1222, the first score of task execution policy C as 0.2548, and the first score of task execution policy D as 0.0019. In this example, policy A with the highest first score may indicate the task execution policy that will be regarded as the first priority by the first score determination unit 510, and policy C with the second highest first score may indicate the task execution policy that will be regarded as the second priority by the first score determination unit 510. Refer to Figure 12a A method for training the first score determination unit 310 according to an embodiment of the present disclosure will be described in more detail.
[0078] According to an embodiment of the present disclosure, the second score determination unit 520 may determine a second score based on a first score and a correction index. For example, the second score may be determined as the product of the first score and the correction index. For example, if the correction indices corresponding to Policy A, Policy B, Policy C, and Policy D are 0.7218, 0.7421, 0.2798, and 0.0572, respectively, the second scores may be determined as the products of the first score and the correction indices: 0.4483, 0.0906, 0.0713, and 0.0001. In this example, Policy A with the highest second score may indicate the task execution policy regarded as the first by the second score determination unit 520, and Policy B with the second highest second score may indicate the task execution policy regarded as the second by the second score determination unit 520. The task execution policy may be generated without considering the user environment. Therefore, even if the task execution policy has a high probability of success in other environments, it may have a low probability of success in the user environment. Since the electronic device 100 continuously requires split inference, it may be difficult to update the split inference policy in real time. In this case, even before updating the task execution policy according to the cause of failure, the electronic device 100 may reduce the probability of selecting a task execution policy with a high number of selection failures by updating the correction index.
[0079] Figure 6 is a block diagram showing a process of an electronic device determining a device according to an embodiment of the present disclosure.
[0080] Referring to Figure 6 , according to an embodiment of the present disclosure, the device determination unit 320 may determine one or more devices to be used for split inference based on device selection conditions. The device selection conditions may include the type of policy, the number of available devices, the current state of the device, and the expected state of the device at the inference time point. The type of policy may refer to the type of policy determined by the policy determination unit 310. When the policy determination unit 310 determines multiple task execution policies, the device determination unit 320 determines the devices corresponding to the determined policies. For example, devices corresponding to the first priority policy and the second priority policy determined by the policy determination unit 310 may be determined respectively. The number of available devices may be identified by the electronic device 100 as described in Figure 1 . The current state of the device may include the heat generation of the device, the remaining memory capacity, and the number of running applications.
[0081] According to an embodiment of the present disclosure, the device determination unit 320 may include an artificial neural network that takes the device selection conditions as input and outputs the fitness for the device. The artificial neural network may be trained to output a larger value for a device with a high probability of success in an inference task for the device selection conditions. Referring to Figure 12b A method of training the device determination unit 320 according to an embodiment of the present disclosure will be described in more detail.
[0082] The device determination unit 320 according to an embodiment of the present disclosure may include an algorithm designed to determine fitness based on a task execution policy with a device selection condition as an input. The device determination unit 320 may quantify the device selection condition and determine one or more devices that satisfy the quantified condition. For example, the device determination unit 320 may quantify the state of the device during inference and determine the fitness for the device by differently weighting the quantified values based on policy-based priorities.
[0083] The electronic device 100 according to an embodiment of the present disclosure may determine one or more devices to be used for split inference based on the number of policy-based devices and the output value of the device determination unit 320. For example, when the number of policy-based devices is 4, the electronic device 100 may determine devices 1, 3, 4, and 5 with high output values as the devices for split inference.
[0084] Figure 7 is a block diagram showing a process in which an electronic device according to an embodiment of the present disclosure determines an inference ratio.
[0085] Referring to Figure 7 , Figure 1 the electronic device 100 may include an inference ratio determination unit 710. The inference ratio determination unit 710 may include a state inference unit 720 and an inference ratio calculation unit 730. The state inference unit 720 according to an embodiment of the present disclosure may include a state inference model, which is an artificial neural network model separate from the artificial neural network for split inference and is trained to predict the state information of a device after the input time when the state information of the device is input.
[0086] The state inference unit 720 may receive first state information of each device at a predetermined first time point from a plurality of devices connected to the network to perform split inference of the artificial neural network. Additionally, the state inference model may infer second state information of each device at a second time point by inputting the first state information of each device, and the second time point is a predetermined time interval after the first time point.
[0087] Additionally, the state inference unit 720 may additionally receive the state information of each device before the first time point and infer the second state information of each device at the second time point.
[0088] Here, the first state information and the second state information of each device may be state information related to the computable amount of each device for inference using an artificial neural network. For example, the first state information and the second state information according to an embodiment may include at least one of the usage rate of a central processing unit (CPU), the usage rate of a graphics processing unit (GPU), the temperature of the CPU, the temperature of the GPU, the number of running applications, and the elapsed time of each device.
[0089] Here, the elapsed time may refer to the reciprocal of floating-point operations per second (FLOPS), which is a unit indicating the computing speed of a computer in terms of the number of instructions that can be processed per unit time. Additionally, in this specification, the elapsed time may refer to the expected computing time of each block of a deep neural network model. That is, the elapsed time may be a criterion indicating the degree to which a device can process an artificial neural network.
[0090] The state inference unit 720 according to an embodiment may infer the elapsed time at a second time point by inputting the first state information of a given device, or may calculate the elapsed time at the second time point by inferring the usage rate of the central processing unit (CPU), the usage rate of the graphics processing unit (GPU), the temperature of the CPU, the temperature of the GPU, and the number of running applications of a given device as the second state information and using the inferred second state information.
[0091] According to an embodiment, in addition to the first state information, the state inference unit 720 may further receive third state information including at least one of whether a given application is running, whether the screen is on, and whether the camera is running. Since it can be expected that the CPU and GPU usage rates on the device will increase when a given application is running, the screen is on, and the camera is running, additional input for this may be received to infer the elapsed time at the second time point.
[0092] The inference ratio calculation unit 730 may receive the second state information of each device from the state inference unit 720 and calculate the inference allocation ratio of the artificial neural network. The inference ratio calculation unit 730 according to an embodiment may normalize the reciprocal of the elapsed time of each device as in the following Mathematical Expression 1. Additionally, the reciprocal of the normalized elapsed time may be determined as the inference allocation ratio (r i ) of the artificial neural network.
[0093] [Mathematical Expression 1]
[0094]
[0095] Here, t i may refer to the elapsed time at the second time point of inference of the i-th device, and n may refer to the total number of multiple devices.
[0096] For example, if the elapsed time of the predicted first device 910 is 0.5, the elapsed time of the predicted second device 920 is 0.4, the elapsed time of the predicted third device 930 is 0.4, and the elapsed time of the predicted electronic device 100 is 0.1, the inference allocation ratio of the first device 910 may be determined to be 0.1176, the inference allocation ratios of the second device 920 and the third device 930 may be determined to be 0.1471, and the inference allocation ratio of the electronic device 100 may be determined to be 0.5882.
[0097] The electronic device 100 according to an embodiment of the present disclosure may send the inference allocation ratio of each determined device and the starting point of the inference process of the artificial neural network. In this case, the plurality of devices may store the entire structure of the artificial neural network. For example, if the determined inference allocation ratio of the first device 910 is 0.1176, the electronic device 100 may send the determined inference allocation ratio 0.1176 and the starting point to the first device 910. Additionally, if the determined inference allocation ratio of the second device 920 is 0.1471, the electronic device 100 may send the determined inference allocation ratio 0.1471 and the point at 11.76% of the entire artificial neural network, which is the starting point. By allocating the inference allocation ratio and the starting point to each device in this way, each device may perform the split inference process of the artificial neural network.
[0098] According to an embodiment of the present disclosure, the state inference model included in the state inference unit 720 may be implemented as a recurrent neural network. The state inference model according to an embodiment of the present disclosure may be a state inference model that is regression-trained by inputting the learned state information at a predetermined third time point and the true state information at a fourth time point (which is a predetermined time interval after the predetermined third time point). In other words, the state inference model may be trained to input the learned state information, calculate a loss function by comparing the inferred state information with the true state information as real data, and reduce the output value of the calculated loss function. The learned state information and the true state information may each include at least one of whether an application is running, whether the screen is on, and whether the camera is running.
[0099] Figure 8 is a block diagram showing a process of an electronic device inferring device state information according to an embodiment of the present disclosure.
[0100] Refer to Figure 8, the state inference unit 720 may receive first state information 810, 820, 830. Additionally, the state inference unit 720 may further receive third state information 850. The third state information 850 received at the first time point T may include whether a specific application (App 1) is running and whether the screen is on. In this case, the pre-trained state inference model may infer more CPU usage, more GPU usage, higher CPU temperature, and higher GPU temperature from the second state information 840.
[0101] Additionally, the state inference unit 720 may receive third state information 850 and infer the state after a long elapsed time. Figure 7 The inference ratio calculation unit 730 that receives the second state information 840 from the state inference unit 720 may determine a lower inference allocation ratio of the first device 910 based on the inferred elapsed time according to the above mathematical expression 1.
[0102] In Figure 8 , it is described that the third state information of the device includes whether a specific application is running and whether the screen is on, but this is only an example, and the third state information may be the state information of a known device in which an environment that can consume the GPU or CPU can be generated. For example, the third state information may further include whether the camera is on.
[0103] Since the state inference unit 720 infers the second state information by additionally obtaining third state information that can significantly change the CPU or GPU usage of the device in addition to obtaining the first state information including the CPU or GPU usage, various effects may exist, including the effect of more accurately predicting that the CPU or GPU usage will significantly increase due to the execution of a specific application, which can be reflected in the inference allocation ratio.
[0104] According to an embodiment of the present disclosure, in order to save the storage space of multiple devices included in the split inference system, the artificial neural network may not be stored in multiple devices. In this case, the electronic device 100 may store the artificial neural network, and according to the inference allocation ratio determined by the electronic device 100, a part of the artificial neural network required for the inference process of the artificial neural network to be executed by each device may be sent.
[0105] Figure 9 is a diagram for describing a method by which an electronic device controls split inference according to an inference ratio according to an embodiment of the present disclosure.
[0106] Referring to Figure 9 , the electronic device 100 may determine the inference allocation ratios of the first device 910, the second device 920, the third device 930, and the electronic device 100, and send theFigure 2 A part of the artificial neural network. In this case, the first device 910, the second device 920, and the third device 930 may not store the artificial neural network for inference, and the electronic device 100 may store the artificial neural network.
[0107] When the inference allocation ratio of the first device 910 is determined to be 0.25, the electronic device 100 may send the first process 915 corresponding to 25% from the start of the entire artificial neural network to the first device 910. In addition, when the inference allocation ratio of the second device 920 is determined to be 0.1, the electronic device 100 may send the second process 925 corresponding to 10% of the entire artificial neural network starting from the 25% point of the entire artificial neural network to the second device 920, and when the inference allocation ratio of the third device 930 is determined to be 0.25, the electronic device 100 may send the third process 935 corresponding to 25% of the entire artificial neural network starting from the 35% point of the entire artificial neural network to the third device 930. In this case, the electronic device 100 may execute the inference process corresponding to 40% of the entire artificial neural network starting from the 60% point of the entire artificial neural network.
[0108] Hereinafter, the inference process and the inference allocation ratio reset process after sending the inference processes to be executed by each device will be described. The first device 910 may execute the first process 915 by inputting the input value of the artificial neural network to perform inference and obtain the first intermediate result value. Then, the first device 910 may send the obtained first intermediate result value and the state information of the first device 910 when executing the first process 915 to the second device 920.
[0109] The second device 920 may obtain the second intermediate result value by using the received first intermediate result value as the input of the second process 925. Then, the second device 920 may send the obtained second intermediate result value, the state information of the second device 920 when executing the second process 925, and the received state information of the first device 910 to the third device 930.
[0110] The third device 930 may obtain the third intermediate result value by using the received second intermediate result value as the input of the third process 935. Then, the third device 930 may send the obtained third intermediate result value, the state information of the third device 930 when executing the third process 935, and the received state information of the first device 910 and the second device 920 to the electronic device 100.
[0111] The electronic device 100 may obtain a final inference result by using the received third intermediate result value as an input for the remaining process of the artificial neural network, and may reset the inference allocation ratio by inferring the state information of each device at a later time point based on the received state information of each device and the state information of the electronic device 100 at the time point when the final inference result is obtained.
[0112] The electronic device 100 that determines a preselected split inference ratio among multiple devices according to an embodiment of the present disclosure may be a device with a good network connection or a large amount of computing power compared to other devices. In other words, the electronic device 100 may be a device determined based on the network state of each device among the multiple devices.
[0113] Here, the network state may be the amount of network input / output (I / O) data packets of each device, which is based on test information received by a first device 910 randomly selected from among the multiple devices from each device different from the first device 910.
[0114] In addition, the electronic device 100 may be at least one candidate device among the multiple devices whose network I / O data packet amount is less than or equal to a predetermined data packet amount, and may be one candidate device among the at least one candidate device connected to a wired network.
[0115] The electronic device 100 according to an embodiment may be a candidate device among the at least one candidate device having the highest GPU processing power.
[0116] As described above, the first device 910 randomly selected from among the multiple devices may select the electronic device 100 that determines the inference allocation ratio.
[0117] Figure 10 It is a diagram showing a method of updating a correction index executed by an electronic device according to an embodiment of the present disclosure.
[0118] Refer to Figure 10 , the task execution policy may include a correction index. The electronic device 100 according to an embodiment of the present disclosure may initialize the correction index when generating the task execution policy or when updating the task execution policy. For example, the correction index may be an initial value of 10000 / 10000. If one or more devices determined based on Policy A do not complete the split inference, the electronic device 100 may decrease the correction index. For example, if the split inference is repeated, the correction index may be reduced from the initial value of 10000 / 10000 to 8763 / 10000. As the number of times the split inference fails increases, the electronic device 100 may significantly decrease the correction index. For example, if Policy A repeatedly fails or the number of failures is large, the reduction amplitude of the correction index may be increased.
[0119] According to an embodiment of the present disclosure, the electronic device 100 may select an appropriate strategy when determining a strategy for the next inference task by reducing the correction index of a failed segmentation strategy. Since the electronic device 100 has time limitations in analyzing reasons and updating strategies every time it performs split inference, the possibility of selecting a failed strategy next time can be reduced by changing the correction index of the strategy.
[0120] Figure 11 is a diagram showing a method of updating a strategy executed by an electronic device according to an embodiment of the present disclosure.
[0121] Referring to Figure 11 , according to an embodiment of the present disclosure, the electronic device 100 may update a task execution strategy by using an execution record of split inference. The electronic device 100 may obtain the execution record from the Figure 1 failure log DB 120.
[0122] The execution record may include records of inference failures due to overheating, inference failures due to insufficient memory capacity, inference failures due to insufficient battery power, and inference failures due to exceeding the maximum time. According to an embodiment of the present disclosure, the electronic device 100 may classify the execution record of split inference. The electronic device 100 may update the strategy by adjusting the priority for many failure reasons or by generating learning data for the failure reasons to train an artificial neural network.
[0123] According to an embodiment of the present disclosure, the electronic device 100 may change the priority of a strategy based on the execution record of split inference. For example, based on the execution record of split inference that is most frequently abnormally terminated due to a throttling problem, the electronic device 100 may increase the priority of overheating. The electronic device 100 may increase the priority of the memory capacity in the case of an inference failure due to insufficient memory capacity, may increase the priority of the remaining battery capacity in the case of an inference failure due to insufficient battery capacity, and may increase the priority of the CPU or GPU type in the case of an inference failure due to exceeding the maximum time.
[0124] According to an embodiment of the present disclosure, the electronic device 100 may generate learning data for split inference failure cases and use the generated learning data to train an artificial neural network. The electronic device 100 may train the artificial neural network to determine fewer strategies with many failure reasons. Specifically, the Figure 3 device determination unit 320 may be trained to output a lower value for a strategy for which the task execution strategy has failed.
[0125] An electronic device 100 according to an embodiment of the present disclosure may increase the number of devices. For example, the electronic device 100 may increase at least one of a lower limit of the number of devices and an upper limit of the number of devices. For example, the number of devices in the electronic device 100 may be increased from a lower limit of four to five. If the failure reason occurs on average rather than under specific conditions, the electronic device 100 may determine to increase the number of devices.
[0126] An electronic device 100 according to an embodiment of the present disclosure may perform at least one operation of changing a priority, changing the number of devices, and initializing a correction index. For example, a correction index of an updated policy A may be 10000 / 10000.
[0127] Figure 12a is a diagram illustrating a method of training an artificial intelligence model for determining an inference task policy performed by an electronic device according to an embodiment of the present disclosure.
[0128] Refer to Figure 12a , a system 1200 for training an artificial intelligence model for determining an inference task policy according to an embodiment of the present disclosure is described. The system 1200 may be included in the electronic device 100 or may be included in an external server. The system 1200 may include a critic 1210, one or more agents 1220-1,..., 1220-n, and a virtual device 1250. Figure 12a FIG. shows a system 1200 according to an embodiment of the present disclosure, which may be processed to train a policy selection model instead of a device selection model.
[0129] Reinforcement learning may be used to train an artificial intelligence model for determining an inference task policy according to an embodiment of the present disclosure. For example, multi-agent reinforcement learning may be used to train the artificial intelligence model.
[0130] The critic 1210 according to an embodiment of the present disclosure provides a reward based on an inference result. For example, if the split inference is successful, a reward of +5 may be given to the policy selection model, and if the split inference fails, a reward of -30 may be given to the policy selection model. Additionally, the critic 1210 may give an additional reward based on a ratio of an actual inference time to an inference request time. For example, if a maximum split time is 2 seconds and the actual inference is 1.2 seconds, the critic 1210 may give an additional reward of 0.6 (=1.2 / 2) to the agent. Additionally, an additional reward may be given based on a value of a reward given to each agent.
[0131] According to an embodiment of the present disclosure, one or more agents 1220-1, ..., 1220-n may include a policy selection model 1230-1, ..., 1230-n and a device selection model 1240-1, ..., 1240-n. The policy selection model 1230-1, ..., 1230-n may be included in Figure 3 the policy determination unit 310 of Figure 3 . The device selection model 1240-1, ..., 1240-n may be included in Figure 12a the device determination unit 320 of. Referring to Figure 12a , the system 1200 may fix the device selection model so that it is not trained, but rather train the policy selection model. By using the policy selection model 1230-1, ..., 1230-n and the device selection model 1240-1, ..., 1240-n, one or more agents 1220-1, ..., 1220-n may determine one or more devices in the virtual device 1250 that will be used to perform split inference, and control the use of the determined one or more devices to perform split inference. The critic 1210 may obtain from one or more devices the result on whether the split inference fails and the status information of the one or more devices that perform the inference.
[0132] As the device selection model, a heuristic model may be used when learning the initial policy model. According to an embodiment of the present disclosure, the system 1200 may train the policy selection model 1230-1, ..., 1230-n by using various learning data, while randomly changing the status of the devices during repeated learning.
[0133] Figure 12b is a diagram showing a method of training an artificial intelligence model for determining a device for performing an inference task by an electronic device according to an embodiment of the present disclosure.
[0134] Referring to Figure 12b , according to an embodiment of the present disclosure, the system 1200 may train the device selection model 1240-1, ..., 1240-n. Figure 12b The critic 1210, one or more agents 1220-1, ..., 1220-n, and the virtual device 1250 of Figure 12a may be understood as corresponding to those of Figure 12a . The policy selection model 1230-1, ..., 1230-n may have been trained in Figure 12a .
[0135] An artificial intelligence model for determining a device to be used for an inference task according to an embodiment of the present disclosure may be trained by reinforcement learning. For example, multi-agent reinforcement learning may be used to train the artificial intelligence model.
[0136] According to an embodiment of the present disclosure, the method by which the critic 1210 gives rewards may be understood as being the same asFigure 12a The method is the same. Additionally, reference may be made to Figure 12a to understand the method for determining one or more devices to be used for split inference and controlling these devices to perform split inference, which is executed by one or more agents 1220-1,..., 1220-n according to an embodiment of the present disclosure.
[0137] The system 1200 according to an embodiment of the present disclosure may train the device selection models 1240-1,..., 1240-n by using various learning data while randomly changing the states of the devices during repeated learning.
[0138] Figure 13 is a diagram showing a method for generating a task execution strategy for a new task, which is executed by an electronic device according to an embodiment of the present disclosure.
[0139] Reference is made to Figure 13 , and the electronic device 100 may generate a task execution strategy for a new task. The electronic device 100 may identify the type of an existing task. For example, the electronic device 100 may obtain information about the type of the existing task from the segmentation policy DB 110.
[0140] The electronic device 100 according to an embodiment of the present disclosure may determine the task in the type of the existing task that is most similar to the new task. As an example, the electronic device 100 may provide information about the existing task to the user 1310 and obtain user input for selecting a similar task. As an example, the electronic device 100 may determine a task with similar characteristics among the existing tasks based on the characteristics of the task.
[0141] The electronic device 100 according to an embodiment of the present disclosure may generate a new task policy based on the task execution policy of the determined similar task. For example, the electronic device 100 may generate a new task policy that is the same as the task execution policy of the similar task.
[0142] Figure 14 is a diagram showing a method for controlling split inference in an indoor environment, which is executed by an electronic device according to an embodiment of the present disclosure.
[0143] Reference is made to Figure 14 , and the AR / VR device 1410, the mobile device 1420, the robotic vacuum cleaner 1430, and the smart TV 1440 may be in the same space.
[0144] The electronic device 100 according to an embodiment of the present disclosure may be at least one of an AR / VR device 1410, a mobile device 1420, a robotic vacuum cleaner 1430, and a smart TV 1440. Additionally, an available device capable of performing split inference according to an embodiment of the present disclosure may be at least one of an AR / VR device 1410, a mobile device 1420, a robotic vacuum cleaner 1430, and a smart TV 1440. In other words, the AR / VR device 1410, the mobile device 1420, the robotic vacuum cleaner 1430, and the smart TV 1440 may all be devices that request split inference and may be devices that perform split inference.
[0145] The AR / VR device 1410 according to an embodiment of the present disclosure may generate a virtual image (e.g., placing a virtual apple on a real table) based on inference using an artificial neural network. The AR / VR device 1410 may use split inference to perform an inference task. The robotic vacuum cleaner 1430 according to an embodiment of the present disclosure may use an artificial neural network to generate a spatial map of the space in a home. The robotic vacuum cleaner 1430 may use split inference to perform an inference task. The smart TV 1440 according to an embodiment of the present disclosure may evaluate a user's exercise posture by using a camera. The smart TV 1440 may use split inference to perform an inference task.
[0146] Figure 15 is a block diagram showing the configuration of an electronic device according to an embodiment of the present disclosure.
[0147] Referring to Figure 15 , the electronic device 100 according to an embodiment may include a memory 1520 and a processor 1510. According to an embodiment of the present disclosure, the configuration of the electronic device 100 is not limited to Figure 15 the configuration shown in Figure 15 and may additionally include Figure 15 components not shown in Figure 15 , or omit some of the components shown in
[0148] For example, although not shown in Figure 15 , the electronic device 100 may further include a transceiver capable of transmitting and receiving information by communicating with another device, an input unit capable of receiving an artificial neural network and input data, and an output unit capable of outputting a result.
[0149] The memory 1520 may be electrically connected to the processor 1510 and may store commands or data related to operations of components included in the electronic device. According to an embodiment of the present disclosure, the memory 1520 may store instructions for operations for performing inference on a task execution policy obtained using a transceiver, first state information of each device, third state information, an artificial neural network model, and a state inference model.
[0150] According to an embodiment of the present disclosure, in an embodiment, when at least some of the modules included in each unit that conceptually divides the functions of the electronic device 100 are implemented as software run by the processor 1510, the memory 1520 may store instructions for running the software modules.
[0151] The processor 1510 may be electrically connected to components included in the electronic device and may perform operations or data processing related to control and / or communication of components included in the electronic device. According to an embodiment of the present disclosure, the processor 1510 may load commands or data received from at least one of other components into the memory 1520, process them, and store the resulting data in the memory 1520.
[0152] In addition, in Figure 15 , for ease of description, the processor 1510 is shown as operating as a single processor 1510, but at least one function that conceptually separates the learning model and the functions of the electronic device described below may be implemented as multiple processors. In this case, the processor 1510 may not operate as a single processor 1510 but may be implemented as multiple processors, where the multiple processors are implemented as separate hardware to perform each operation. The present disclosure is not limited thereto.
[0153] The transceiver may support establishing a wired communication channel or a wireless communication channel between the electronic device and another external electronic device and support performing communication through the established communication channel.
[0154] In addition, according to various embodiments, the transceiver may include a wireless communication module (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (e.g., a local area network (LAN) communication module or a power line communication module), and may communicate with an external electronic device through a short-range communication network (e.g., Bluetooth, WiFi Direct, or Infrared Data Association (IrDA)) or a long-range communication network (e.g., a cellular network, the Internet, or a computer network (e.g., LAN or WAN)) by using the corresponding communication module.
[0155] Figure 3The multiple devices for split inference of an artificial neural network may each include components that perform the same functions as the memory 1520, the processor 1510, and the transceiver of the above-described electronic device 100. Since the functions of each component are as described above, detailed descriptions thereof will be omitted.
[0156] According to an embodiment of the present disclosure, a method for controlling the execution of an inference task through split inference of an artificial neural network is provided. The method may include determining one strategy from multiple task execution strategies based on the requirements of the inference task and a correction index indicating the failure rate of each task execution strategy. The method may include determining one or more devices for performing split inference of the artificial neural network based on the strategy. The method may include updating the correction index corresponding to the strategy based on the result of whether the split inference performed by the one or more devices fails. The method may include updating the strategy by using the execution record of the split inference obtained from the one or more devices, where the execution record includes information about the cause of failure of the split inference. Each task execution strategy among the multiple task execution strategies may include considering at least one of the priority of one or more device conditions for selecting a device for performing split inference and the number of devices for split inference.
[0157] According to an embodiment of the present disclosure, the operation of determining the strategy may include determining a first score for the multiple task execution strategies based on the requirements of the inference task. When determining the strategy, a second score for the multiple task execution strategies may be determined based on the first score and the correction index determined based on the failure rate of the task execution strategy. The operation of determining the strategy may include determining the execution strategy with the highest second score as the strategy.
[0158] According to an embodiment of the present disclosure, the method may include determining the execution strategy with the second highest second score as an alternative strategy. The method may include determining one or more alternative devices for performing split inference of the artificial neural network based on the alternative strategy. The method may include controlling the one or more alternative devices to perform split inference based on the failure of the split inference performed by the one or more devices.
[0159] According to an embodiment of the present disclosure, the operation of updating the correction index may include decreasing the correction index based on the failure of the split inference. The operation of updating the correction index may include maintaining or increasing the correction index based on the success of the split inference.
[0160] According to an embodiment of the present disclosure, the operation of updating the strategy may include: identifying the number of times of split inference failure for each device condition among the one or more device conditions based on the execution record of the split inference. The operation of updating the strategy may include updating the priority or updating the number of devices for split inference such that the device condition with a high number of split inference failures has a higher priority.
[0161] According to an embodiment of the present disclosure, the method may further include training at least one of the first artificial intelligence model and the second artificial intelligence model by using at least one of an updated calibration index corresponding to a policy and an updated policy. The first artificial intelligence model may be an artificial intelligence model trained to determine one or more policies among a plurality of task execution policies based on requirements of an inference task and a calibration index. The second artificial intelligence model may be an artificial intelligence model trained to determine one or more devices for performing split inference of an artificial neural network based on the determined one or more policies.
[0162] According to an embodiment of the present disclosure, the method may include identifying a new inference task different from the inference task. The method may include controlling one or more devices to perform split inference, wherein one or more devices are determined to perform the new inference task based on an updated plurality of task execution policies.
[0163] According to an embodiment of the present disclosure, the method may include: displaying a plurality of types of inference tasks corresponding to a plurality of task execution policies stored in a database based on that a task execution policy corresponding to the type of the inference task is not stored in the database. The method may include obtaining a user input for selecting one inference task from the plurality of inference tasks. The method may include generating a task execution policy for an inference task identical to the one or more task execution policies corresponding to the inference task input by the user.
[0164] According to an embodiment of the present disclosure, the requirements may include at least one of a type of an inference task, importance, a maximum inference time, and required memory.
[0165] According to an embodiment of the present disclosure, conditions of one or more devices may include at least one of a processor type of the device, a remaining memory capacity, a heat generation level, a remaining battery capacity, and a number of running applications.
[0166] According to an embodiment of the present disclosure, there is provided an apparatus for controlling the execution of an inference task through split inference of an artificial neural network. The apparatus may include a memory and at least one processor, the memory including one or more instructions. The at least one processor may determine a policy from a plurality of task execution policies based on the requirements of the inference task and a correction index indicating the failure rate of each task execution policy. The at least one processor may determine one or more devices for performing split inference of the artificial neural network based on the policy. The at least one processor may update the correction index corresponding to the policy based on the result of whether the split inference performed by the one or more devices fails. The at least one processor may update the policy using the execution record of the split inference obtained from the one or more devices. The execution record may include information about the cause of failure of the split inference. Each task execution policy among the plurality of task execution policies may include at least one of a priority for considering one or more device conditions for selecting a device for performing split inference and the number of devices for split inference.
[0167] According to an embodiment of the present disclosure, the at least one processor may determine a first score for a plurality of task execution policies based on the requirements of the inference task. The at least one processor may determine a second score for the plurality of task execution policies based on the first score and a correction index determined based on the failure rate of the task execution policy. The at least one processor may determine the execution policy having the highest second score as the policy.
[0168] According to an embodiment of the present disclosure, the at least one processor may determine the execution policy having the second highest second score as an alternative policy. The at least one processor may determine one or more alternative devices for performing split inference of the artificial neural network based on the alternative policy. The at least one processor may control the one or more alternative devices to perform split inference based on the failure of the split inference performed by the one or more devices.
[0169] According to an embodiment of the present disclosure, the at least one processor may decrease the correction index based on the failure of the split inference. The at least one processor may maintain or increase the correction index based on the success of the split inference.
[0170] According to an embodiment of the present disclosure, the at least one processor may identify the number of times of failure of the split inference for each of the one or more device conditions based on the execution record of the split inference. The at least one processor may update the priority or update the number of devices for split inference such that the device condition with a high number of times of failure of the split inference has a higher priority.
[0171] According to an embodiment of the present disclosure, at least one processor may train at least one of a first artificial intelligence model and a second artificial intelligence model by using at least one of an updated calibration index corresponding to a policy and an updated policy. The first artificial intelligence model may be an artificial intelligence model trained to determine one or more policies among a plurality of task execution policies based on requirements of an inference task and a calibration index. The second artificial intelligence model may be an artificial intelligence model trained to determine one or more devices for performing split inference of an artificial neural network based on the determined one or more policies.
[0172] According to an embodiment of the present disclosure, at least one processor may identify a new inference task different from an inference task. The at least one processor may control one or more devices to perform split inference, wherein one or more devices are determined to perform the new inference task based on an updated plurality of task execution policies.
[0173] According to an embodiment of the present disclosure, if a task execution policy corresponding to a type of an inference task is not stored in a database, the at least one processor may display a plurality of types of inference tasks corresponding to a plurality of task execution policies stored in the database. The at least one processor may obtain a user input for selecting one inference task among the plurality of inference tasks. The at least one processor may generate a task execution policy for an inference task identical to the one or more task execution policies corresponding to the inference task input by the user.
[0174] According to an embodiment of the present disclosure, the requirements may include at least one of a type, importance, maximum inference time, and required memory of an inference task.
[0175] According to an embodiment of the present disclosure, there is provided a computer-readable recording medium having recorded thereon a program for running the method on a computer.
[0176] Embodiments of the present disclosure may also be implemented in the form of a recording medium including computer-executable instructions, such as program modules run by a computer. The computer-readable medium may be any available medium accessible by a computer and includes volatile and non-volatile media, removable and non-removable media. The computer-readable medium may include computer storage media and communication media. The computer storage media includes volatile and non-volatile media, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The communication media typically may include other data in a modulated data signal, such as computer-readable instructions, data structures, or program modules.
[0177] A computer-readable storage medium according to an embodiment of the present disclosure may be provided in the form of a non-transitory storage medium. Here, the "non-transitory storage medium" is a tangible device and refers to a device that does not include signals (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium. For example, the "non-transitory storage medium" may include a buffer for temporarily storing data.
[0178] A method according to an embodiment of the present disclosure may be provided to be included in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online via an application store (e.g., by downloading or uploading) or directly between two user devices (e.g., a smart phone). In the case of online distribution, at least a part of the computer program product (e.g., a downloadable application) may be temporarily stored or temporarily generated in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediate server.
[0179] The description of the present disclosure is for illustrative purposes only, and those of ordinary skill in the art to which the present disclosure pertains will understand that the present disclosure can be easily modified into other specific forms without changing the technical idea or basic features of the present disclosure. Therefore, it should be understood that the above embodiments are examples in all aspects and not restrictive. For example, each component described as a single entity may be implemented in a distributed manner, and similarly, components described as distributed may be implemented in a combined manner.
[0180] The scope of the present disclosure is indicated by the claims described below rather than the above detailed description, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be construed as being included within the scope of the present disclosure.
Claims
1. A method for controlling the execution of an inference task through split inference of an artificial neural network, the method comprising: Determining a policy from a plurality of task execution policies based on at least one of the requirements of the inference task and a correction index, wherein the correction index is determined based on the failure rate of each task execution policy (S210); Determining one or more devices for performing split inference of the artificial neural network based on the policy (S220); Updating the correction index corresponding to the policy based on the result of whether the split inference performed by the one or more devices fails (S230); and Updating the policy by using the execution record of the split inference obtained from the one or more devices (S240), wherein the execution record includes information about the cause of failure of the split inference, wherein each task execution policy among the plurality of task execution policies includes at least one of a priority for considering one or more device conditions for selecting a device for performing split inference and the number of devices for split inference.
2. The method according to claim 1, Among them, The operation of determining the policy (S210) includes: Determining a first score for the plurality of task execution policies based on the requirements of the inference task; Determining a second score for the plurality of task execution policies based on the first score and a correction index determined based on the failure rate of the task execution policy; Determining the execution policy with the highest second score as the policy.
3. The method according to claim 2, further comprising: Determining the execution policy with the highest second score as an alternative policy; Determining one or more alternative devices for performing split inference of the artificial neural network based on the alternative policy; Controlling the one or more alternative devices to perform split inference based on the failure of the split inference performed by the one or more devices.
4. The method according to any one of claims 1 to 3, Among them, The operation of updating the correction index (S230) includes: Reducing the correction index based on the failure of the split inference; and Maintaining or increasing the correction index based on the success of the split inference.
5. The method according to any one of claims 1 to 4, Among them, The operation of updating the policy (S240) includes: Identifying the number of times of split inference failure for each device condition among the one or more device conditions based on the execution record of the split inference; Updating the priority or updating the number of devices for split inference such that the device condition with a high number of split inference failures has a higher priority.
6. The method according to any one of claims 1 to 5, Further comprising training at least one of a first artificial intelligence model and a second artificial intelligence model by using at least one of the updated correction index and the updated policy corresponding to the policy, Among them, The first artificial intelligence model is an artificial intelligence model trained to determine one or more policies among a plurality of task execution policies based on the requirements of the inference task and the correction index, and The second artificial intelligence model is an artificial intelligence model trained to determine one or more devices for performing split inference of the artificial neural network based on the determined one or more policies.
7. The method according to any one of claims 1 to 6, comprising: identifying a new inference task different from the inference task; and controlling one or more devices to perform split inference, wherein the one or more devices are determined to perform the new inference task based on an updated plurality of task execution policies.
8. The method according to any one of claims 1 to 7, comprising: displaying a plurality of types of inference tasks corresponding to a plurality of task execution policies stored in the database based on that a task execution policy corresponding to the type of the inference task is not stored in the database; obtaining user input for selecting one inference task from the plurality of inference tasks; and generating a task execution policy for the inference task that is the same as one or more task execution policies corresponding to the inference task of the user input.
9. The method according to any one of claims 1 to 8, Among them, wherein the claims include at least one of the type, importance, maximum inference time, and required memory of the inference task.
10. The method according to any one of claims 1 to 9, Among them, wherein the conditions of the one or more devices include at least one of the type of the processor of the device, the remaining memory capacity, the heat generation level, the remaining battery capacity, and the number of running applications.
11. An apparatus 100 for controlling the execution of an inference task by split inference of an artificial neural network, the apparatus comprising: a memory 1520 including one or more instructions; and at least one processor 1510, wherein the at least one processor 1510 is configured to: determine a policy from a plurality of task execution policies based on the requirements of the inference task and a correction index indicating the failure rate of each task execution policy; determine one or more devices for performing split inference of the artificial neural network based on the policy; update the correction index corresponding to the policy based on the result of whether the split inference performed by the one or more devices fails; update the policy by using an execution record of the split inference obtained from the one or more devices, wherein the execution record includes information on the cause of failure of the split inference, wherein each task execution policy in the plurality of task execution policies includes at least one of a priority for considering one or more device conditions for selecting a device for performing split inference and the number of devices for split inference.
12. The apparatus according to claim 11, Among them, wherein the at least one processor 1510 is further configured to: determine a first score for the plurality of task execution policies based on the requirements of the inference task; determine a second score for the plurality of task execution policies based on the first score and a correction index determined based on the failure rate of the task execution policy; and determine the execution policy with the highest second score as the policy.
13. The apparatus according to claim 12, Among them, wherein the at least one processor 1510 is further configured to: determine an execution policy having the highest second score as an alternative policy, determine one or more alternative devices for performing split inference of the artificial neural network based on the alternative policy, control the one or more alternative devices to perform split inference based on a failure of the split inference performed by the one or more devices.
14. The apparatus according to any one of claims 11 to 13, Among them, wherein the at least one processor 1510 is further configured to: reduce the correction index based on a failure of split inference, and maintain or increase the correction index based on a success of split inference.
15. A computer-readable recording medium having recorded thereon a program for performing the method according to any one of claims 1 to 10 on a computer.