Electronic device, method, server, and storage medium for scaling instance of artificial intelligence model

By employing a response time AI model and loss function optimization to determine optimal AI instance counts, the method addresses inefficiencies in scaling AI models, achieving efficient and timely user request processing with reduced resource consumption.

WO2025147105A1PCT designated stage expired Publication Date: 2025-07-10SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/000056
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-24
Filing Date
2025-01-02
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing systems face challenges in efficiently scaling instances of artificial intelligence models to optimize response times and resource utilization in processing user requests, leading to suboptimal performance and resource inefficiencies.

Method used

A method and system for determining an optimal number of instances of AI models based on a response time AI model, using a loss function optimization procedure to balance response time and resource usage, which can be executed on electronic devices or external servers, utilizing GPUs and neural processing units.

Benefits of technology

This approach allows for rapid user request processing with minimal resources by dynamically adjusting the number of AI model instances, optimizing response times and resource allocation, thereby enhancing service efficiency and reducing latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025000056_10072025_PF_FP_ABST
    Figure KR2025000056_10072025_PF_FP_ABST
Patent Text Reader

Abstract

A method for executing an instance of at least one AI model may be provided, according to one embodiment. The method may comprise an operation of identifying a response time AI model for at least one AI model associated with a first service. The response time AI model may be configured to receive the number of user requests for the first service and the number of instances of each of the at least one AI model as input values, and output a response time. The method may comprise an operation of identifying a loss function set on the basis of the number of instances of the at least one AI model, the response time AI model, and an allowable response time. The method may comprise an operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization process on the loss function using the number of user requests for the first service. The method may comprise an operation of performing at least one operation for executing instances of the at least one AI model according to the at least one optimal number of instances. Other various embodiments are possible.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices, methods, servers, and storage media for scaling instances of artificial intelligence models

[0001] The present disclosure relates to electronic devices, methods, servers, and storage media for scaling instances of artificial intelligence (AI) models.

[0002] To provide a specific service, multiple artificial intelligence (AI) models can be operated in a linked form. For example, AI models are being utilized in diverse fields such as content streaming, translation, photo editing, finance, new drug development, law, and the military. At least some of the multiple AI models for a single service can be implemented as generative AI models. For example, a service providing the content of relatively large PDF files may include an embedding AI model for vectorizing the PDF (Portable Document Format) file and a generative AI language model for processing vector search results. Furthermore, research is actively underway on frameworks for linking multiple AI models (e.g., LangChain).

[0003] Services that include AI models can generally be performed using graphics processing units (GPUs). GPUs are specialized for parallel processing and are therefore suitable for processing AI models with structures such as deep neural networks (DNNs). A service provider can acquire and / or lease multiple GPUs to execute AI models to provide services. For example, a service provider can experimentally verify information related to user requests processed on a single GPU for a specific AI model and, based on the verification results, determine the level of GPU acquisition, lease, and / or activation.

[0004] The above information may be provided as background information to aid in understanding this document. None of the above is claimed to be prior art related to this document or can be used to determine prior art.

[0005] According to one embodiment, a method for executing instances of at least one AI model may be provided. The method may include an operation of identifying a response time AI model for at least one AI model associated with a first service. The response time AI model may be configured to receive, as inputs, a number of user requests for the first service and a number of instances of each of the at least one AI model, and output a response time. The method may include an operation of identifying a loss function established based on the number of instances of the at least one AI model, the response time AI model, and an acceptable response time. The method may include an operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function using the number of user requests for the first service. The method may include an operation of performing at least one operation for executing instances of the at least one AI model according to the at least one optimal number of instances.

[0006] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by at least a portion of at least one processor of an electronic device, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of identifying a response time AI model for at least one AI model associated with a first service. The response time AI model may be configured to receive as input values ​​a number of user requests for the first service and a number of instances of each of the at least one AI model, and output a response time. The at least one operation may include an operation of identifying a loss function set based on the number of instances of the at least one AI model, the response time AI model, and an acceptable response time. The at least one operation may include an operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function using the number of user requests for the first service. The at least one operation may include performing at least one operation to execute instances of the at least one AI model according to the at least one optimal number of instances.

[0007] According to one embodiment, an electronic device may include at least one processor. The electronic device may include a memory that stores at least one instruction. The at least one instruction, when executed by at least a portion of the at least one processor, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of identifying a response time AI model for at least one AI model associated with a first service. The response time AI model may be configured to receive, as input values, a number of user requests for the first service and a number of instances of each of the at least one AI model, and output a response time. The at least one operation may include an operation of identifying a loss function set based on the number of instances of the at least one AI model, the response time AI model, and an acceptable response time. The at least one operation may include an operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function using the number of user requests for the first service. The at least one operation may include performing at least one operation to execute instances of the at least one AI model according to the at least one optimal number of instances.

[0008] According to one embodiment, a method for generating training data for training an AI model of a service's response time associated with at least one AI model may be provided. The method may include an operation of setting a number of user requests. The method may include an operation of setting a number of instances of the at least one AI model. The method may include an operation of executing an instance of the at least one AI model according to the number of instances of the at least one AI model. The method may include an operation of generating user requests according to the number of user requests. The method may include an operation of checking responses corresponding to each of the user requests based on processing the user requests based on instances of the at least one AI model. The method may include an operation of generating training data including the number of user requests and the number of instances of the at least one AI model as input values ​​and including response times corresponding to the responses as output values.

[0009] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by at least a portion of at least one processor of an electronic device, may cause the electronic device to perform at least one operation. The at least one operation may include setting a number of user requests. The at least one operation may include setting a number of instances of at least one AI model. The at least one operation may include executing an instance of the at least one AI model according to the number of instances of the at least one AI model. The at least one operation may include generating user requests according to the number of user requests. The at least one operation may include checking responses corresponding to each of the user requests based on processing the user requests based on instances of the at least one AI model. The at least one operation may include generating training data including the number of user requests and the number of instances of the at least one AI model as input values ​​and response times corresponding to the responses as output values.

[0010] According to one embodiment, an electronic device may include at least one processor. The electronic device may include a memory storing at least one instruction. The at least one instruction, when executed by at least a portion of the at least one processor, may cause the electronic device to perform at least one operation. The at least one operation may include setting a number of user requests. The at least one operation may include setting a number of instances of at least one AI model. The at least one operation may include executing an instance of the at least one AI model according to the number of instances of the at least one AI model. The at least one operation may include generating user requests according to the number of user requests. The at least one operation may include checking responses corresponding to each of the user requests based on processing the user requests based on the instances of the at least one AI model. The at least one operation may include generating training data including the number of user requests and the number of instances of the at least one AI model as input values ​​and response times corresponding to the responses as output values.

[0011] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0012] FIG. 1a is a diagram illustrating processing of a user request by an instance of an AI model according to one embodiment.

[0013] Figure 1b is a diagram for explaining resource utilization according to one embodiment.

[0014] FIG. 1c is a diagram illustrating processing of a user request by an instance of an AI model according to one embodiment.

[0015] FIG. 1d is a diagram for explaining resource utilization according to one embodiment.

[0016] FIG. 2A illustrates a block diagram of an electronic device according to one embodiment.

[0017] FIG. 2b illustrates a block diagram of an electronic device according to one embodiment.

[0018] FIG. 2c illustrates a block diagram of an electronic device according to one embodiment.

[0019] FIG. 3 is a flowchart illustrating a method for executing an instance of an AI model according to one embodiment.

[0020] FIG. 4 is a flowchart illustrating a method for training a response time AI model according to one embodiment.

[0021] FIG. 5a is a flowchart illustrating a method for generating a training data set of a response time AI model according to one embodiment.

[0022] FIG. 5b is a flowchart illustrating a method for generating a training data set of a response time AI model according to one embodiment.

[0023] FIG. 6A illustrates a block diagram of an electronic device according to one embodiment.

[0024] FIG. 6b illustrates a block diagram of an electronic device according to one embodiment.

[0025] Figure 7 is a flowchart illustrating a service execution method according to one embodiment.

[0026] Figure 8 is a flowchart for explaining a service execution method according to one embodiment.

[0027] FIG. 9 is a diagram for explaining a generative artificial intelligence system according to one embodiment.

[0028] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0029] FIG. 1A is a diagram illustrating the processing of a user request by an instance of an AI model according to one embodiment. The embodiment of FIG. 1A will be described with reference to FIG. 1B. FIG. 1B is a diagram illustrating resource utilization according to one embodiment.

[0030] According to one embodiment, a service may be performed by linking at least one AI model. For example, a service that provides the contents of a PDF file may receive a PDF file as input and provide contents corresponding to the PDF file (e.g., text, but not limited thereto) as output. The service may include an embedding AI model for vectorizing the PDF file and a generative AI language model for processing vector search results. In the embodiment of FIG. 1A, the service may include (or link to) at least one AI model (e.g., the first AI model to the Mth AI model). To perform the service, the electronic device (101) to be described with reference to FIG. 2A may execute an instance (12a, 13a, 15a) of at least one AI model. An instance may be, for example, an object corresponding to a program (or application) such as an AI model, and may be named a replica, a pod, a container, or a virtual machine, without limitation in its name. The number of instances may correspond to the size of the resource (e.g., GPU), and thus the number of instances may be used interchangeably with the size of the resource, or instances may be used interchangeably with the resource.

[0031] In the embodiment of FIG. 1A, for example, it is assumed that one instance is executed for each of at least one AI model. For example, a first instance (12a) corresponding to a first AI model, a first instance (13a) corresponding to a second AI model, and a first instance (15a) corresponding to an M AI model may be executed. As the first instance (12a) corresponding to the first AI model, the first instance (13a) corresponding to the second AI model, and the first instance (15a) corresponding to the M AI model are executed, a relatively small amount of resources may be utilized. For example, FIG. 1B illustrates resources (50) available to an electronic device (101). The resources (50) may include, for example, at least one unit resource (50a to 50o). The at least one unit resource (50a to 50o) may be, for example, a GPU, but is not limited thereto. For example, those skilled in the art will appreciate that at least one unit resource (50a to 50o) may include an entity for processing (e.g., but not limited to, a central processing unit (CPU), a GPU, and / or a neural processing unit (NPU)) including at least one circuit and / or an entity for storage (e.g., but not limited to, a volatile memory and / or a non-volatile memory). For example, one unit resource (e.g., unit resource (50a)) may correspond to one GPU or a specified unit size, but this is merely for illustration and there is no limitation on the size and / or expression method of one unit resource.

[0032] For example, in the example of FIG. 1b, it is assumed that the execution of the first instance (12a) corresponding to the first AI model may require one unit resource (e.g., unit resource (50a)), the execution of the first instance (13a) corresponding to the second AI model may require two unit resources (e.g., unit resources (50f, 50g)), and the execution of the first instance (15a) corresponding to the M AI model may require one unit resource (e.g., unit resource (50k)). Meanwhile, the number and / or size of resources required for each instance of the above-described AI model are merely exemplary. Accordingly, when executing one instance for each of at least one AI model, four unit resources may be required. The remaining idle resources may, for example, be unused or may be used for other services.

[0033] Meanwhile, referring back to FIG. 1a, N user requests (11a, 11b, 11c, ..., 11n) may be input to the electronic device (101). The user requests (11a, 11b, 11c, ..., 11n) may be associated with, for example, a service. For example, the user request (11a) may be processed by the first instance (12a) of the first AI model, and a first processing result may be provided from the first instance (12a) of the first AI model. The first processing result may be processed by the first instance (13a) of the second AI model, and thus, a second processing result may be provided by the first instance (13a) of the second AI model. By the chain processing of the processing results, the first instance (15a) of the M AI model may receive the (N-1) processing result and process it. The first instance (15a) of the M AI model can provide the Nth processing result as a response (16a). Accordingly, a response (16a) corresponding to a user request (11a) can be provided. Based on the above-described process, responses (16a, 16b, 16c, ..., 16n) corresponding to a plurality of user requests (11a, 11b, 11c, ..., 11n) can be provided, respectively. Meanwhile, since processing must be performed by instances (12a, 13a, 15a), the time (which can be referred to as a response time) for providing responses (16a, 16b, 16c, ..., 16n) corresponding to a plurality of user requests (11a, 11b, 11c, ..., 11n), respectively, may take a relatively long time. To reduce response time, the electronic device (101) may increase the number of instances of at least one AI model, which may also be referred to as scaling out. Hereinafter, the increase in the number of instances will be described with reference to FIGS. 1c and 1d.

[0034] FIG. 1C is a diagram illustrating the processing of a user request by an instance of an AI model according to one embodiment. The embodiment of FIG. 1C will be described with reference to FIG. 1D. FIG. 1D is a diagram illustrating resource utilization according to one embodiment.

[0035] According to one embodiment, the electronic device (101) described later in FIG. 2a can execute two instances for each of at least one AI model. For example, the electronic device (101) can execute two instances (12a, 12b) corresponding to a first AI model, two instances (13a, 13b) corresponding to a second AI model, and two instances (15a, 15b) corresponding to an M AI model. As multiple instances (12a, 12b, 13a, 13b, 15a, 15b) are executed, user requests (11a, 11b, 11c, ..., 11n) can be distributedly processed. For example, one instance group (12a, 13a, 15a) may process some of the user requests (11a, 11b, 11c, ... , 11n), and another instance group (12b, 13b, 15b) may process the remaining some of the user requests (11a, 11b, 11c, ... , 11n). Accordingly, the response time for one instance group (12a, 13a, 15a) to provide responses (16a, 16b, 16c, ... , 16n) corresponding to the user requests (11a, 11b, 11c, ... , 11n) may be shorter than the response time of the embodiment of Fig. 1a. On the other hand, by executing relatively more instances, relatively larger unit resources may be required. For example, referring to FIG. 1d, two unit resources (50a, 50b) may be used for the execution of each of two instances (12a, 12b) corresponding to the first AI model. Four unit resources (50f, 50g, 50h, 50i) may be used for the execution of each of two instances (13a, 13b) corresponding to the second AI model. Two unit resources (50k, 50l) may be used for the execution of each of two instances (15a, 15b) corresponding to the M AI model.As shown in FIGS. 1b and 1d, for doubling the number of instances, doubling the number of unit resources may be required, and there may be a trade-off between the response time and the resource size. Accordingly, an optimal number of instances that takes the trade-off into account is required. In addition, an optimal number of instances that takes the allowable response time into account may be required. For example, different optimal numbers of instances may be required when the allowable response time is set to a relatively large value and when the allowable response time is set to a relatively small value. According to one embodiment, the electronic device (101) can determine the optimal number of instances based on the number of user requests and the allowable response time, which will be described later.

[0036] FIG. 2A illustrates a block diagram of an electronic device according to one embodiment.

[0037] Referring to FIG. 2A, an electronic device (101) may include a processor (120), a memory (130), and / or a communication interface (190). The electronic device (101) may be a device for providing a service linked to at least one AI model. The electronic device (101) may be implemented as, for example, a single entity, or may be implemented as a plurality of entities. For example, the electronic device (101) may be implemented as a cloud server, and there is no limitation on the form of implementation. The resource (50) described in FIGS. 1B and 1D may refer to, for example, a portion of the processor (120), the memory (130), and / or the communication interface (190). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more external electronic devices. For example, when an electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service on its own, request one or more external electronic devices to execute at least a portion of the function or service. The one or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technologies may be utilized, for example. The electronic device (101) may provide an ultra-low latency service by utilizing, for example, distributed computing or mobile edge computing.The electronic device (101) may be an intelligent server using machine learning and / or neural networks.

[0038] The processor (120) may, for example, execute software (e.g., a program) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operation, the processor (120) may store a command or data received from another component (e.g., a communication interface (190)) in a memory (130) (e.g., a volatile memory, but not limited thereto), process the command or data stored in the volatile memory, and store the resulting data in the memory (130) (e.g., a non-volatile memory, but not limited thereto). According to one embodiment, the processor (120) may include, but is not limited to, a central processing unit (CPU), a GPU, and / or a neural processing unit (NPU) including circuitry.

[0039] The memory (130) may store various data used by at least one component (e.g., the processor (120) and / or the communication interface (190)) of the electronic device (101). The data may include, for example, software (e.g., a program) and input data or output data for commands related thereto. The memory (130) may include volatile memory and / or non-volatile memory. The memory (130) may include a hard disk, ROM, RAM, cache memory, and / or registers, and the implementation thereof is not limited thereto. Some of the above-described entities (e.g., registers, but not limited thereto) may be implemented as part of the processor (120), and the implementation form thereof is not limited thereto. The memory (130) may store at least one AI model for instance execution.

[0040] The processor (120) may execute, for example, at least one instruction stored in the memory (130). The memory (130) may store at least one instruction, and the at least one instruction may be executed by the processor (120). The at least one instruction, when executed by the processor (120), may cause the electronic device (101) to perform at least one operation. For example, as the at least one instruction is executed, at least one other component may be controlled, and / or various data processing or operations may be performed. As at least a part of the data processing or operations, the processor (120) may store commands or data received from other components in at least a part of the memory (130), process the commands or data stored in the memory (130), and store result data in the memory (130). A processor (120) performing an operation may mean, for example, that the operation is performed by (or under the control of) one entity included in the processor (120) (e.g., but not limited to, the main processor). For example, a processor performing an operation may mean, for example, that a specific operation is performed by (or under the control of) multiple entities (e.g., multiple processors). For example, a plurality of operations may mean, for example, that all of the plurality of operations are performed by (or under the control of) one entity (e.g., but not limited to, the main processor). For example, a plurality of operations may mean that some of the plurality of operations are performed by at least one entity, and some of the remaining operations are performed by at least one other entity. For example, at least one instruction that causes one or more actions to be performed may be stored in one memory, or may be distributed across each of a plurality of memories.

[0041] The communication interface (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and the external electronic devices (106a, 106b, 106c, ..., 106n), and the performance of communication through the established communication channel. The communication interface (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication, but is not limited thereto. According to one embodiment, the communication interface (190) may include an Ethernet-based wired communication module (e.g., a local area network (LAN) communication module, or a power line communication module). The communication interface (190) may also include a wireless communication module, but is not limited thereto. At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0042] FIG. 2b illustrates a block diagram of an electronic device according to one embodiment.

[0043] According to one embodiment, the electronic device (101) may include and / or execute a frontend module (210) and / or a service processing module (220). The frontend module (210) and / or the service processing module (220) may be executed by, for example, the processor (120), or may be included as at least a part of the processor (120) or another entity. At least some of the operations performed by the frontend module (210) and / or the service processing module (220) in the present disclosure may be understood to be performed by, for example, the processor (120) and / or another entity under the control of the processor (120).

[0044] The front-end module (210) can perform at least one operation for exchanging data with, for example, external electronic devices (106a, 106b, 106c, ... , 106n). For example, the front-end module (210) can provide data that can configure a UI that can input user input from the external electronic devices (106a, 106b, 106c, ... , 106n). For example, the front-end module (210) can be implemented as a web server. The external electronic devices (106a, 106b, 106c, ... , 106n) can access the front-end module (210) based on a URL entered into a web browsing application, or can access the front-end module (210) by executing an application, but there is no limitation. The front-end module (210) can provide processing for a user request to the service processing module (220). The service processing module (220) can perform a service using a user request and may also be referred to as a backend module. The service processing module (220) can provide a response corresponding to the user request to the frontend module (210). The frontend module (210) can provide the response received from the service processing module (220) to an external electronic device (106a, 106b, 106c, ..., 106n).

[0045] The service processing module (220) may include, for example, a user request confirmation module (221), an optimization module (222), a loss function management module (223), an AI model management module (224), a response time AI model management module (225), a policy management module (226), and / or a service execution module (227).

[0046] The user request verification module (221) can verify information associated with a user request provided from an external electronic device (106a, 106b, 106c, ..., 106n). The information associated with the user request can be expressed as, for example, the number of user requests for a certain period of time and / or the size of the user request, but there is no limitation thereto. For example, the user request verification module (221) can count the number of user requests provided from the external electronic devices (106a, 106b, 106c, ..., 106n) and / or monitor the size of the content included in the user request (for example, text and / or graphic objects, but there is no limitation thereto), but there is no limitation on the type of information associated with the user request and / or the verification method.

[0047] The optimization module (222) may provide at least one optimal number of instances for each of at least one AI model associated with the service. The loss function management module (223) may store and / or manage (for example, including but not limited to, adding, deleting, and / or updating) the loss function. The AI ​​model management module (224) may store and / or manage (for example, including but not limited to, adding, deleting, and / or updating) the AI ​​model associated with the service. The response time AI model management module may store and / or manage (for example, including but not limited to, adding, deleting, and / or updating) the AI ​​model trained to provide the response time by the service. The policy management module (226) may store and / or manage (for example, including but not limited to, adding, deleting, and / or updating) the allowable response time, for example. For example, the management device (104) may verify (e.g., receive or determine) at least one input for determining an acceptable response time and provide it to the electronic device (101). For example, the manager may input information about an acceptable response time for the corresponding service into the management device (104), but this is exemplary and there is no limitation on the method by which the acceptable response time is verified. The response time AI model may provide a response time based on, for example, information associated with a user request and the number of instances. The loss function may provide a loss value based on, for example, the response time provided from the response time AI model and the acceptable response time managed by the policy management module (226). The optimization module (222) may provide an optimal number of instances based on performing an optimization process (e.g., gradient descent, but without limitation) on the loss value.

[0048] The service execution module (227) can execute, for example, at least one instance group (231, 232, 233) corresponding to at least one AI model linked to the service. The service execution module (227) can execute, for example, at least one instance group (231, 232, 233) corresponding to at least one AI model according to the optimal number of instances provided by the optimization module (222). The service execution module (227) can process a user request based on the at least one instance group (231, 232, 233) executed, and can provide a response according to the processing result to an external electronic device (106a, 106b, 106c, ..., 106n) through the front-end module (210). As described with reference to FIG. 1c, when there are multiple instances being executed, the user request can be distributedly processed by the multiple instances.

[0049] The electronic device (101) can determine the optimal number of instances of each of at least one AI model associated with the service based on information (e.g., number and / or size) and policies (e.g., allowed response time) associated with the provided user request that is confirmed in real time. As described above, the number of instances of the AI ​​model may be in a trade-off relationship with the number and / or size of resources. The electronic device (101) can determine the optimal number of instances of each of at least one AI model that can process information (e.g., number and / or size) associated with the user request according to policies (e.g., allowed response time). Accordingly, relatively rapid user request processing based on relatively few resources may be possible.

[0050] FIG. 2c illustrates a block diagram of an electronic device according to one embodiment.

[0051] According to one embodiment, the external service execution module (229) may not be included in the electronic device (101) or may be executed by a device other than the electronic device (101). For example, the electronic device (101) may provide a user request provided from an external electronic device (106a, 106b, 106c, ..., 106n) to the external service execution module (229). Meanwhile, the electronic device (101) may determine an optimal number of instances of each of at least one AI model capable of processing information (e.g., number and / or size) associated with the user request according to a policy (e.g., allowed response time), as described above. The electronic device (101) may provide the determined optimal number of instances of each of at least one AI model to the external service execution module (229). The external service execution module (229) may confirm the optimal number of instances of each of at least one AI model from the electronic device (101). The external service execution module (229) can execute at least one instance group (231, 232, 233) corresponding to at least one AI model according to the optimal number of instances. The external service execution module (229) can process a user request based on the executed at least one instance group (231, 232, 233) and provide a response according to the processing result to the external electronic device (106a, 106b, 106c, ..., 106n) through the front-end module (210). Meanwhile, as in FIG. 2c, the electronic device (101) relaying the user request and response is merely exemplary, and those skilled in the art will understand that the user request and response may be directly transmitted / received between the external electronic device (106a, 106b, 106c, ..., 106n) and the external service execution module (229).

[0052] Meanwhile, it is exemplary that at least one instance of the entire AI model is executed by the electronic device (101) as in FIG. 2b, or at least one instance of the entire AI model is executed by an external entity (e.g., an external service execution module (229)) as in FIG. 2c. Those skilled in the art will understand that, depending on the implementation, some instances of at least one AI model may be executed by the electronic device (101), and some other instances may be executed by an external entity (e.g., an external service execution module (229)). For example, those skilled in the art will understand that some of the multiple instances corresponding to one AI model may be executed by the electronic device (101), and some other instances may be executed by an external entity (e.g., an external service execution module (229)).

[0053] FIG. 3 is a flowchart illustrating a method for executing an instance of an AI model according to one embodiment.

[0054] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0055] According to one embodiment, the electronic device (101) may, in operation 301, check a response time AI model for at least one AI model associated with the first service. The response time AI model may be configured to receive, for example, the number of user requests and the number of instances as input values, and output the response time of the first service. For example, the response time of the first service may mean, but is not limited to, the time taken until a response is output based on the user request being processed by at least one AI model associated with the first service when a user request is input for the first service as described above. The response time AI model may use, for example, a fully-connected layer (FC) or a graph neural network (GNN) that can effectively express a workload in a graph form, such as an AI pipeline, but this is exemplary and the type of the model is not limited. The response time AI model may be named a response time function, and there is no limitation on the name, as it provides an output value (e.g., response time) corresponding to input values ​​(e.g., number of user requests and number of instances). The number of user requests may be, for example, the number of user requests confirmed per unit time, but this is exemplary and there is no limitation on the confirmation method. Those skilled in the art will understand that, depending on the implementation, the number of user requests may be replaced with the total size of user requests, or both may be used. The generation of a training data set for the response time AI model will be described with reference to FIGS. 5A and 5B . The electronic device (101) may generate, receive, and / or load a trained response time AI model.

[0056] The electronic device (101) can, in operation 303, determine a loss function based on the number of instances of at least one AI model, the response time AI model, and the allowable response time. The loss function can be configured to include, for example, an increasing component when the number of instances of the AI ​​model increases, and a penalty component having a relatively large value when the response time is relatively greater than the allowable response time, but there is no limitation on the configuration. The loss function can be set based on, for example, the number of instances of at least one AI model, the response time AI model, and the allowable response time. Mathematical expression 1 is an example of a loss function.

[0057] <Mathematical Formula 1>

[0058]

[0059] The above mathematical formula 1 is merely an example to aid understanding, and embodiments of the present disclosure may not be limited thereto. For example, the above mathematical formula 1 may be modified, applied, or expanded in various ways.

[0060] In mathematical expression 1, Loss can mean a loss function. , may mean the number of instances of at least one AI model. The dimension of may correspond to the number of at least one AI model, for example. If there is at least one AI model, can also be expressed as a scalar. The SLO can be, for example, an acceptable response time. The acceptable response time can be, for example, a tail value of a distribution of response times, but this is not limited to an example. For example, a tail value can mean a response time at a point that is greater than or equal to a specified threshold percentage of a distribution of multiple response times. For example, if the threshold percentage is 95% and the tail value is about 100ms, it can mean that 95% of the multiple response times must have a value of about 100ms or less. The acceptable response time can be set, for example, by an administrator, but there is no limitation on how it can be set. can be the number of user requests per unit time. For example, the number of user requests can be categorized by type, Each component of may be the number of user requests for each type. If the number of user requests is not managed by type, can be a scalar. , may be a response time function. As described above, the response time function is the number of instances of at least one AI model ( ) and number of user requests ( ) as an input value and can be configured to output a response time. The φ function in mathematical expression 1 can be the same as mathematical expression 2.

[0061] <Mathematical Formula 2>

[0062]

[0063] The above mathematical formula 2 is merely an example to aid understanding, and embodiments of the present disclosure may not be limited thereto. For example, the above mathematical formula 2 may be modified, applied, or expanded in various ways.

[0064] As described above, the φ function may be a function configured to output the larger value between 0 and the difference between two input values. ρ may be a constant and may correspond to, for example, a penalty level. Meanwhile, the Max function, as in Equation 2, is merely exemplary, and those skilled in the art will understand that there are no limitations as long as it provides a relatively large value when the response time is greater than the allowable response time.

[0065] Referring to mathematical expression 1, the loss function is the first component may include. The first component comprises the number of instances of at least one AI model ( ) can be the sum of. For example, if the number of instances of the first AI model is 1 and the number of instances of the second AI model is 2, the first component can be 3, which is the sum of 1 and 2. For example, the number of instances of at least one AI model ( ) increases, the first component may increase.

[0066] The loss function is the second component For example, if the response time is relatively larger than the allowable response time, the second component can have a relatively large value. If the response time is relatively smaller than the allowable response time, the second component can have 0, and thus the second component can be named a penalty component. It can be understood that the weight for the penalty component is set relatively large as ρ is set relatively large.

[0067] The electronic device (101), in operation 305, can perform an optimization procedure for a loss function using the number of user requests to determine at least one optimal number of instances for each of at least one AI model. For example, the electronic device (101) can input the number of user requests into a loss function such as Equation 1 and determine an output value of the loss function. The electronic device (101) can perform an optimization procedure based on the output value of the loss function, and thus determine at least one optimal number of instances for each of at least one AI model. For example, the electronic device (101) can perform gradient descent as an optimization procedure, but this is exemplary and there is no limitation on the optimization procedure. For example, the electronic device (101) can perform an optimization procedure for the loss function. The result of partial differentiation for , Based on the value and step size at the current step, the next step The electronic device (101) can perform an optimization procedure until the difference between the output value of the loss function at the current step and the output value of the loss function at the previous step becomes within the critical difference, and accordingly The optimal value for can be identified, which can be named the optimal number of instances. For example, the optimization process can be as shown in mathematical expression 3.

[0068] <Mathematical Formula 3>

[0069]

[0070] The above mathematical formula 3 is merely an example to aid understanding, and embodiments of the present disclosure may not be limited thereto. For example, the above mathematical formula 3 may be modified, applied, or expanded in various ways.

[0071] In mathematical expression 3, is the vector corresponding to the next step, can be a vector corresponding to the current step. α can be the step size, It may be a result of inputting information about the current step into the result of partial differentiation of the loss function. The electronic device (101) can confirm the output value of the loss function at the current step as the optimal value when the difference between the output value of the loss function at the current step and the output value of the loss function at the previous step is within a critical difference, and this can be named the optimal number of instances. There is no restriction on the initial value of . For example, the loss function can be a convex function and can be a function with a single global minimum, so there may be no restriction on the initial value.

[0072] The electronic device (101), in operation 307, may perform at least one operation to execute an instance corresponding to at least one AI model according to at least one optimal instance number. For example, when the electronic device (101) directly executes an instance, the instance corresponding to at least one AI model may be executed according to at least one optimal instance number. Alternatively, the electronic device (101) may provide information about the optimal instance number to the external service execution module (229), so that the external service execution module (229) may execute an instance corresponding to at least one AI model according to at least one optimal instance number. For example, when the service is running in Kubernetes, the electronic device (101) may issue a command to an API (application programming interface) server of Kubernetes to execute an instance corresponding to at least one AI model according to at least one optimal instance number, but there is no limitation.

[0073] The electronic device (101) may, for example, periodically check the optimal number of instances, or may check the optimal number of instances based on an event such as an increase or decrease in the number of user requests.

[0074] FIG. 4 is a flowchart illustrating a method for training a response time AI model according to one embodiment.

[0075] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0076] According to one embodiment, the electronic device (101) may, in operation 401, verify a training data set. The electronic device (101), in operation 403, may train a response time AI model based on the training data set. The electronic device (101), in operation 405, may provide the trained response time AI model. For example, the response time AI model may be configured to receive the number of instances of at least one AI model and the number of user requests as input values, and output a response time, as described above. For example, each training data set may include the number of instances of at least one AI model and the number of user requests as input values, and include a response time as an output value. A method of generating a training data set will be described with reference to FIGS. 5A and 5B . As described above, the response time AI model may use, for example, a fully connected neural network or a graph neural network, but this is exemplary and the type of the model is not limited. The electronic device (101) can train a response time AI model, for example, through supervised learning on a training data set, but there are no limitations on the training method. For example, the electronic device (101) can perform training based on a different loss function for the response time AI model using the training data set. If the loss value according to the different loss function is below a threshold loss value, it can be determined that training has been sufficient, but there are no limitations.

[0077] FIG. 5a is a flowchart illustrating a method for generating a training data set of a response time AI model according to one embodiment.

[0078] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0079] According to one embodiment, the electronic device (101) may set the number of user requests in operation 501. There is no limitation on the initial value of the number of user requests. The electronic device (101) may set the number of instances corresponding to each of at least one AI models in operation 503. There is no limitation on the initial value of the number of instances. The electronic device (101) may execute at least one AI model according to the set number of instances in operation 505. For example, the electronic device (101) may execute at least one AI model according to the set number of instances based on AI model-related information, driving engine program information, resource information per instance, and / or resource-related information, but there is no limitation. The AI ​​model-related information may include information on an AI pipeline that constitutes a service and / or at least one AI model. For example, the information on at least one AI model may include parameters of each AI model, but there is no limitation. The information on the pipeline may include information on the order in which each AI model constitutes a service, but there is no limitation. Information about at least one AI model may be provided by the AI ​​model developer, and / or information about the pipeline structure may be provided by the developer of an application that utilizes the AI ​​model, but there is no limitation on who provides such information. The driving engine program information may include programming code capable of driving at least one AI model and / or the pipeline. For example, if a pipeline structure is implemented, an AI model may be driven using GPU resources by calling a GPU-related library, and the driving engine program information may include such code. The code may be provided by the application developer, but there is no limitation on who provides such code.Per-instance resource information refers to the amount of resources (e.g., the number of GPUs, but there is no limit) that each AI model can use at one time. For example, if the number of GPUs per instance is 1, 1 GPU may be required to execute 1 instance of the AI ​​model. Per-instance resource information may be provided, for example, by an infrastructure operator's device, but there is no limit to the entity providing the information. Resource-related information refers to information about the resource (e.g., GPU) used by the AI ​​model, and may include, for example, the manufacturer and / or information of the GPU model. Per-instance resource information may also be determined based on resource-related information. For example, if the processing capacity of the GPU is relatively large, the number of GPUs per instance may be reduced. Resource-related information may be provided, for example, by an AI infrastructure operator's device, but there is no limit to the entity providing the information.

[0080] The electronic device (101) may generate a load according to the set number of user requests in operation 507. Here, the load may refer to a user request. The electronic device (101) may generate a load, for example, based on a user request format. The user request format may refer to a format for an input value to a service based on at least one AI model. For example, the user request format may follow the Hypertext Transfer Protocol (HTTP) or Google Remote Procedure Call (gRPC) protocol format, but this is exemplary and not limiting. The user request format may include information about at least one parameter that can be used for input to the service, for example. The user request format may be provided by the developer of an application that utilizes the AI ​​model, but there is no limitation on the providing entity. The electronic device (101) may input the generated load to at least one AI model and verify training data based on the execution result of at least one AI model. For example, the electronic device (101) can check the response time required to process the generated load. For example, the training data may include the number of user requests and the number of instances corresponding to each of at least one AI models as input values, and may include the response time as an output value.

[0081] The electronic device (101) can, in operation 509, determine whether the degree of response time reduction is below a critical reduction degree. The critical reduction degree can be set to a value at which, for example, a reduction in response time due to a change in the number of instances becomes insignificant, and there is no limitation on the setting criteria. If the degree of response time reduction is not below the critical reduction degree (operation 509 - No), the electronic device (101) can, in operation 511, change the number of instances (for example, it can be increased, but there is no limitation). After changing the number of instances, the electronic device (101) can additionally check the load generation and training data. The electronic device (101) can, for example, collect training data while changing the number of instances until the degree of response time reduction is below the critical reduction degree. If the degree of response time reduction is below the critical reduction degree (operation 509 - Yes), the electronic device (101) can, in operation 513, determine whether the number of user requests is greater than or equal to the critical number. The threshold number can be set based on the service usage history, but there is no limit. If the number of user requests is not greater than the threshold number (operation 513 - No), the electronic device (101) can change the number of user requests (for example, it can be increased, but there is no limit) in operation 515. After changing the number of user requests, the electronic device (101) can additionally perform load generation. The electronic device (101) can collect training data while changing the number of user requests, for example, until the number of user requests is greater than the threshold number. Through the above-described process, a training data set can be collected, and the electronic device (101) can train a response time AI model, as shown in FIG. 4, based on the collected training data set. Meanwhile, those skilled in the art will understand that some of the collected training data set may be used for training, and the remaining part may be used for verification.

[0082] The electronic device (101) can train a response time AI model according to a training data set in the same manner as in Fig. 5a. For example, as mentioned in mathematical expression 1, the input to the response time AI model is can also be expressed as a scalar.

[0083] FIG. 5b is a flowchart illustrating a method for generating a training data set of a response time AI model according to one embodiment.

[0084] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0085] According to one embodiment, the electronic device (101) may set the number of user requests in operation 501. There is no limitation on the initial value of the number of user requests. The electronic device (101) may set the number of instances corresponding to each of at least one AI models in operation 503. There is no limitation on the initial value of the number of instances. The electronic device (101) may set the type of the user request in operation 504. There is no limitation on the initial type of the type of the user request. The electronic device (101) may execute at least one AI model according to the set number of instances in operation 505. The electronic device (101) may generate a load according to the set number of user requests in operation 507. The electronic device (101) may determine whether the degree of reduction in response time is less than or equal to a critical reduction degree in operation 509. The critical reduction degree may be set to a value at which, for example, a reduction in response time due to a change in the number of instances is insignificant, and there is no limitation on the setting criteria. If the degree of response time reduction is not below the critical reduction degree (Operation 509 - No), the electronic device (101) may change the number of instances (for example, increase, but there is no limit) in operation 511. After changing the number of instances, the electronic device (101) may additionally check the load generation and training data. The electronic device (101) may collect training data while changing the number of instances, for example, until the degree of response time reduction is below the critical reduction degree. If the degree of response time reduction is below the critical reduction degree (Operation 509 - Yes), the electronic device (101) may check whether the number of user requests is greater than or equal to the critical number in operation 513. The critical number may be set based on the usage history of the service, but there is no limit.If the number of user requests is not greater than or equal to the threshold number (Operation 513 - No), the electronic device (101) may, in operation 515, change the number of user requests (e.g., increase but without limitation). After changing the number of user requests, the electronic device (101) may additionally perform load generation. The electronic device (101) may, for example, collect training data while changing the number of user requests until the number of user requests becomes greater than or equal to the threshold number. If the number of user requests is greater than or equal to the threshold number (Operation 513 - Yes), the electronic device (101) may, in operation 517, check whether training data collection has been completed for all user request types. If training data collection has not been completed for all user request types (Operation 517 - No), the electronic device (101) may, in operation 519, change the type of user request. After changing the type of user request, the electronic device (101) may additionally perform load generation. The electronic device (101) can collect training data while changing the user request type until, for example, training data collection is completed for all user request types.

[0086] FIG. 6A illustrates a block diagram of an electronic device according to one embodiment.

[0087] According to one embodiment, the electronic device (101) may include and / or execute a load generation module (641), a training data verification module (642), a training module (643), and / or a service execution module (227). The load generation module (641), the training data verification module (642), the training module (643), and / or the service execution module (227) may be executed by, for example, the processor (120), or may be included as at least a part of the processor (120) or another entity. At least some of the operations performed by the load generation module (641), the training data verification module (642), the training module (643), and / or the service execution module (227) in the present disclosure may be understood to be performed by, for example, the processor (120) and / or another entity under the control of the processor (120). The electronic device (101) may include a training DB (630). The training DB (630) may be stored (or included) in, for example, memory (130), but there is no limitation.

[0088] The training DB (630) may include AI model-related information (631), driving engine program information (632), instance-specific resource information (633), user request format (634), and / or resource-related information (635). The AI ​​model-related information (631) may include information about an AI pipeline that constitutes a service and / or at least one AI model. For example, the information about at least one AI model may include parameters of each AI model, but is not limited thereto. The information about the pipeline may include information about the order in which each AI model constitutes a service, but is not limited thereto. The information about at least one AI model may be provided by an AI model developer, and / or the information about the pipeline structure may be provided by a developer of an application that uses the AI ​​model, but is not limited thereto. The AI ​​model developer may be, for example, an entity that develops and / or trains an AI model, but is not limited thereto. An application developer may be an entity that develops an application that utilizes an AI model, for example, and may develop a chatbot, such as a generative AI model, using a language model developed by an AI model developer, but there are no limitations.

[0089] The driving engine program information (632) may include programming code capable of driving at least one AI model and / or pipeline. For example, if a pipeline structure is implemented, an AI model may be driven using GPU resources by calling a GPU-related library, and the driving engine program information (632) may include the corresponding code. The code may be provided by an application developer, but there is no limitation on the entity providing the code. The resource information per instance (633) per instance indicates the amount of resources that each AI model will use at one time (for example, the number of GPUs may be available, but there is no limitation). The resource information per instance (633) may be provided by, for example, an infrastructure operator's device, but there is no limitation on the entity providing the resource information. The infrastructure operator's device may be, for example, an entity that manages and operates the infrastructure. The infrastructure operator's device may periodically monitor the infrastructure and perform tasks to respond to problems, but there is no limitation. The infrastructure operator's device may perform infrastructure management based on the deployment of applications developed by application developers and / or the requirements of the service operator's device. A service operator's device may be an entity that operates and / or manages AI model-based applications. The service operator's device may operate AI model-based applications under operational requirements and, for example, perform initial response tasks in the event of a failure.

[0090] The user request format (634) may refer to a format for input values ​​to a service based on at least one AI model. For example, the user request format (634) may follow an HTTP or gRPC protocol format, but this is exemplary and not limiting. The user request format (634) may include information about at least one parameter that can be used as input to the service, for example. The user request format (634) may be provided by a developer of an application utilizing the AI ​​model, but there is no limitation on the entity providing the user request format.

[0091] Resource-related information (635) refers to information about resources (e.g., GPUs) utilized by the AI ​​model, and may include, for example, the manufacturer and / or information of the GPU model. Resource-related information (635) may also determine resource information for each instance. For example, if the processing capacity of the GPU is relatively large, the number of GPUs per instance may be reduced. Resource-related information (635) may be provided, for example, by a device of an AI infrastructure operator, but there are no restrictions on the provider.

[0092] The load generation module (641) can generate loads, for example, based on a set number of user requests. As described above, the loads may refer to user requests. For example, the load generation module (641) can generate loads based on a user request format (634).

[0093] The service execution module (227) can execute instances (661, 662, 663) corresponding to at least one AI model. The service execution module (227) can execute instances (661, 662, 663) based on, for example, AI model-related information (631), driving engine program information (632), instance-specific resource information (633), and / or resource-related information (635), but is not limited thereto. The service execution module (227) can process user requests provided from the load generation module (641) using the instances (661, 662, 663). The training data verification module (642) can verify the response time required to process the user request. For example, the training data verification module (642) can verify the response time based on the tail value of the response times at which responses corresponding to multiple user requests are provided. For example, the training data verification module (642) can verify the tail value corresponding to the lower approximately 95% of multiple response times as the response time, but there is no limitation on the method of verifying the response time. The training data verification module (642) can verify the number of instances and the number of user requests as input values, and verify the corresponding response time as an output value. As described with reference to FIGS. 5A and 5B , multiple training data sets can be provided as the number of instances, the number of user requests, and / or the type of user requests are changed, and thus a training data set can be provided. The training module (643) can train a response time AI model based on the training data set, for example, as described in FIG. 4 .

[0094] FIG. 6b illustrates a block diagram of an electronic device according to one embodiment.

[0095] According to one embodiment, the external service execution module (229) may not be included in the electronic device (101) or may be executed by a device other than the electronic device (101). For example, the electronic device (101) may provide an instance execution request according to a set number of instances to the external service execution module (229). The external service execution module (229) may execute instances (661, 662, 663) according to the received set number of instances. The electronic device (101) may provide a user request provided by the load generation module (641) to the external service execution module (229). The external service execution module (229) may process the user request based on the instances (661, 662, 663) and may provide the electronic device (101) with a response time corresponding to each user request. The training data verification module (642) may verify the number of instances and the number of user requests as input values, and verify the response time as an output value. Meanwhile, as in FIG. 6A, it is exemplary that at least one instance of the entire AI model is executed by the electronic device (101), or as in FIG. 6B, at least one instance of the entire AI model is executed by an external entity (e.g., an external service execution module (229)). Those skilled in the art will understand that, depending on the implementation, some instances of at least one AI model may be executed by the electronic device (101), and some other instances may be executed by an external entity (e.g., an external service execution module (229)). For example, those skilled in the art will understand that some of the multiple instances corresponding to one AI model may be executed by the electronic device (101), and some other instances may be executed by an external entity (e.g., an external service execution module (229)).

[0096] Figure 7 is a flowchart illustrating a service execution method according to one embodiment.

[0097] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0098] According to one embodiment, the electronic device (101) may execute a first service based on a first resource in operation 701. For example, the electronic device (101) may identify an optimal number of instances of each of at least one AI model associated with the first service based on the number of user requests and the allowed response time. The electronic device (101) may execute instances of each of at least one AI model based on the optimal number of instances, and thus, the first service may be executed based on the first resource corresponding to the optimal number of instances. The electronic device (101) may newly identify the number of user requests in operation 703. The electronic device (101) may identify a second resource corresponding to the newly identified number of user requests in operation 705. For example, the electronic device (101) may identify a new optimal number of instances of each of at least one AI model associated with the first service based on the newly identified number of user requests and the allowed response time. The electronic device (101) may identify a second resource for executing at least one instance of each AI model based on the new optimal instance count. For example, the second resource may be assumed to be greater than the first resource. In operation 707, the electronic device (101) may execute the first service based on the second resource to maintain an acceptable response time. For example, to maintain an acceptable response time, the electronic device (101) may increase the resource used for the first service from the first resource to the second resource.

[0099] Figure 8 is a flowchart for explaining a service execution method according to one embodiment.

[0100] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0101] According to one embodiment, the electronic device (101) may execute a first service based on a first resource in operation 801. For example, the electronic device (101) may identify an optimal number of instances of each of at least one AI model associated with the first service based on the number of user requests and the allowed response time. The electronic device (101) may execute instances of each of at least one AI model based on the optimal number of instances, and thus, the first service may be executed based on the first resource corresponding to the optimal number of instances. The electronic device (101) may newly identify the number of user requests in operation 803. The electronic device (101) may identify a second resource corresponding to the newly identified number of user requests in operation 805. For example, the electronic device (101) may identify a new optimal number of instances of each of at least one AI model associated with the first service based on the newly identified number of user requests and the allowed response time. The electronic device (101) may identify a second resource for executing at least one instance of each AI model based on the new optimal number of instances. For example, the second resource may be assumed to be greater than the first resource. In operation 807, the electronic device (101) may adjust (e.g., increase) the allowed response time to maintain the resources being used. For example, the electronic device (101) may identify the allowed response time for processing the newly identified number of user requests while maintaining the first resource. The electronic device (101) may output the identified allowed response time so that an administrator can recognize it, and the administrator may operate the electronic device (101) to maintain the resources if the identified allowed response time is determined to be acceptable. Alternatively, the administrator may input an acceptable allowed response time into the electronic device (101).In this case, the electronic device (101) can adjust the SLO in mathematical expression 1 to the input value and can also identify resources based on the optimal number of instances. The electronic device (101) can also execute the first service using resources corresponding to the identified optimal number of instances.

[0102] FIG. 9 is a diagram for explaining a generative artificial intelligence system according to one embodiment.

[0103] According to one embodiment, a user query / response interface (910) may receive a user's input. The user's input may be in the form of natural language, images, and / or videos, but is not limited thereto. Furthermore, context information may also be transmitted when the user's input is transmitted. The context information may include various additional information at the time of the user's input. For example, the additional information may include information about the application currently being used by the user or information about the user's location. Furthermore, the user's input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Furthermore, the user's input may also be in a non-natural language form, such as selecting a menu. The user query / response interface (910) may output the results of the generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The user query / response interface (910) may output the results of the generative artificial intelligence system to the user. The output can be in natural language form, in the form of specific content, or in the form of actions requested by the user.

[0104] The AI ​​framework (940) can receive user input and coordinate and control each component necessary to perform the user's intention based on the user's query.

[0105] User input received from the user query / response interface (910) can be transmitted to a prompt design component (941). The prompt design component (941) can be used to generate prompts suitable for inputting user input into a large language model (LLM) or a large multimodal model (LMM). The prompt design component (941) can be an AI component that uses a machine learning algorithm or a neural network to develop better prompts over time. The prompt design component (941) can access a knowledge component including user preference data, a prompt library, and prompt examples based on the user input to generate prompts, and transmit the generated prompts to the LLM or LMM.

[0106] The API / Plug-in management component (942) can communicate with external information when there is a request for additional information when passing user input as input to the generative model. The API / Plug-in management component (942) can establish a channel for communicating with the outside of the AI ​​Interface through the API, and can enable access to various data sources (e.g., knowledge repositories (920)) through the established channel. In addition, the API / Plug-in management component (942) can request the application / service component (930) through the API for an action that ultimately performs the user input, rather than an intermediate result, when the action needs to be performed in the application or service. Information obtained from the outside can be used to generate a prompt in the prompt design component (941) together with the user input, or can be passed as an input to the generative model.

[0107] The output modification component (also called a refiner component) (943) can fine-tune the output from the generative model. For example, the output modification component (943) can verify that the content generated through the LLM and / or LMM is not irrelevant, does not contain biased content, or does not contain harmful content. In addition, the output modification component (943) can determine to what extent it matches the result desired by the user and, if necessary, can perform additional processes. The output modification component (943) can additionally configure and provide the user with hints to avoid unwanted output.

[0108] A generative AI model (960) may generally refer to an artificial intelligence neural network that creates new types of data based on user input information. The generative AI model (960) may include an image-generating model and / or a language-generating model. Representative models for generating images include a generative adversarial network (GAN) and a variational auto encoder (VAE), and examples include a VAE and a Diffusion-based generative model that uses a Transformer structure. A language-generating model is a model trained to statistically output the most appropriate output based on input values, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there are also LMMs (large multimodal models) that can recognize various types of data input, such as text, images, and voice, and generate new data corresponding to them.

[0109] According to one embodiment, a method for executing instances of at least one AI model may be provided. The method may include an operation of identifying a response time AI model for at least one AI model associated with a first service. The response time AI model may be configured to receive, as inputs, a number of user requests for the first service and a number of instances of each of the at least one AI model, and output a response time. The method may include an operation of identifying a loss function established based on the number of instances of the at least one AI model, the response time AI model, and an acceptable response time. The method may include an operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function using the number of user requests for the first service. The method may include an operation of performing at least one operation for executing instances of the at least one AI model according to the at least one optimal number of instances.

[0110] According to one embodiment, the loss function may include a component based on the difference between the number of user requests input into the response time AI model and the allowed response time.

[0111] In one embodiment, the loss function may include a component based on the difference and a sum of the number of instances of the at least one AI model.

[0112] According to one embodiment, the loss function is based on Equation 1, wherein Equation 1 is It can be. Above The number of instances of at least one AI model may be the number of instances of the at least one AI model. The SLO may be the acceptable response time. , may be the number of types of the above user requests. , may be the response time function. , may be the sum of the number of instances of at least one AI model. The φ function is a component based on the difference, based on mathematical expression 2, wherein mathematical expression 2 is It can be. The above ρ can be a constant.

[0113] According to one embodiment, the operation of determining at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function may include determining at least one optimal number of instances for each of the at least one AI model by performing the optimization procedure for the loss function with respect to the number of instances of the at least one AI model.

[0114] In one embodiment, the optimization procedure may include gradient descent.

[0115] According to one embodiment, the operation of performing at least one operation for executing instances of the at least one AI model according to the at least one optimal number of instances may include an operation of executing instances of the at least one AI model according to the at least one optimal number of instances.

[0116] According to one embodiment, the operation of performing at least one operation for executing instances of the at least one AI model according to the at least one optimal number of instances may include an operation of requesting an external electronic device to execute instances of the at least one AI model according to the at least one optimal number of instances.

[0117] According to one embodiment, the method may further include receiving a command setting the allowable response time.

[0118] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by at least a part of at least one processor (120) of an electronic device, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of identifying a response time AI model for at least one AI model associated with a first service. The response time AI model may be configured to receive as input values ​​the number of user requests for the first service and the number of instances of each of the at least one AI model, and output a response time. The at least one operation may include an operation of identifying a loss function set based on the number of instances of the at least one AI model, the response time AI model, and an acceptable response time. The at least one operation may include an operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function using the number of user requests for the first service. The at least one operation may include performing at least one operation to execute instances of the at least one AI model according to the at least one optimal number of instances.

[0119] According to one embodiment, the loss function may include a component based on the difference between the number of user requests input into the response time AI model and the allowed response time.

[0120] In one embodiment, the loss function may include a component based on the difference and a sum of the number of instances of the at least one AI model.

[0121] According to one embodiment, the loss function is based on Equation 1, wherein Equation 1 is It can be. Above The number of instances of at least one AI model may be the number of instances of the at least one AI model. The SLO may be the acceptable response time. , may be the number of types of the above user requests. , may be the response time function. , may be the sum of the number of instances of at least one AI model. The φ function is a component based on the difference, based on mathematical expression 2, wherein mathematical expression 2 is It can be. The above ρ can be a constant.

[0122] According to one embodiment, the operation of determining at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function may include determining at least one optimal number of instances for each of the at least one AI model by performing the optimization procedure for the loss function with respect to the number of instances of the at least one AI model.

[0123] In one embodiment, the optimization procedure may include gradient descent.

[0124] According to one embodiment, the operation of performing at least one operation for executing instances of the at least one AI model according to the at least one optimal number of instances may include an operation of executing instances of the at least one AI model according to the at least one optimal number of instances.

[0125] According to one embodiment, the operation of performing at least one operation for executing instances of the at least one AI model according to the at least one optimal number of instances may include an operation of requesting an external electronic device to execute instances of the at least one AI model according to the at least one optimal number of instances.

[0126] According to one embodiment, the at least one operation may further comprise receiving a command setting the allowable response time.

[0127] According to one embodiment, an electronic device may include at least one processor (120). The electronic device may include a memory (130) that stores at least one instruction. The at least one instruction, when executed by at least a portion of the at least one processor (120), may cause the electronic device to perform at least one operation. The at least one operation may include an operation of identifying a response time AI model for at least one AI model associated with a first service. The response time AI model may be configured to receive, as input values, the number of user requests for the first service and the number of instances of each of the at least one AI model, and output a response time. The at least one operation may include an operation of identifying a loss function set based on the number of instances of the at least one AI model, the response time AI model, and an acceptable response time. The at least one operation may include an operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function using the number of user requests for the first service. The at least one operation may include performing at least one operation to execute instances of the at least one AI model according to the at least one optimal number of instances.

[0128] According to one embodiment, the loss function may include a component based on the difference between the number of user requests input into the response time AI model and the allowed response time.

[0129] In one embodiment, the loss function may include a component based on the difference and a sum of the number of instances of the at least one AI model.

[0130] According to one embodiment, the loss function is based on Equation 1, wherein Equation 1 is It can be. Above The number of instances of at least one AI model may be the number of instances of the at least one AI model. The SLO may be the acceptable response time. , may be the number of types of the above user requests. , may be the response time function. , may be the sum of the number of instances of at least one AI model. The φ function is a component based on the difference, based on mathematical expression 2, wherein mathematical expression 2 is It can be. The above ρ can be a constant.

[0131] According to one embodiment, the operation of determining at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function may include determining at least one optimal number of instances for each of the at least one AI model by performing the optimization procedure for the loss function with respect to the number of instances of the at least one AI model.

[0132] In one embodiment, the optimization procedure may include gradient descent.

[0133] According to one embodiment, the operation of performing at least one operation for executing instances of the at least one AI model according to the at least one optimal number of instances may include an operation of executing instances of the at least one AI model according to the at least one optimal number of instances.

[0134] According to one embodiment, the operation of performing at least one operation for executing instances of the at least one AI model according to the at least one optimal number of instances may include an operation of requesting an external electronic device to execute instances of the at least one AI model according to the at least one optimal number of instances.

[0135] According to one embodiment, the at least one operation may further comprise receiving a command setting the allowable response time.

[0136] According to one embodiment, a method for generating training data for training an AI model of a service's response time associated with at least one AI model may be provided. The method may include an operation of setting a number of user requests. The method may include an operation of setting a number of instances of the at least one AI model. The method may include an operation of executing an instance of the at least one AI model according to the number of instances of the at least one AI model. The method may include an operation of generating user requests according to the number of user requests. The method may include an operation of checking responses corresponding to each of the user requests based on processing the user requests based on instances of the at least one AI model. The method may include an operation of generating training data including the number of user requests and the number of instances of the at least one AI model as input values ​​and including response times corresponding to the responses as output values.

[0137] According to one embodiment, a storage medium storing at least one computer-readable instruction may be provided. The at least one instruction, when executed by at least a portion of at least one processor (120) of an electronic device, may cause the electronic device to perform at least one operation. The at least one operation may include an operation of setting a number of user requests. The at least one operation may include an operation of setting a number of instances of at least one AI model. The at least one operation may include an operation of executing an instance of the at least one AI model according to the number of instances of the at least one AI model. The at least one operation may include an operation of generating user requests according to the number of user requests. The at least one operation may include an operation of checking responses corresponding to each of the user requests based on processing the user requests based on instances of the at least one AI model. The at least one operation may include an operation of generating training data including the number of user requests and the number of instances of the at least one AI model as input values ​​and response times corresponding to the responses as output values.

[0138] According to one embodiment, an electronic device may include at least one processor (120). The electronic device may include a memory (130) that stores at least one instruction. The at least one instruction, when executed by at least a portion of the at least one processor (120), may cause the electronic device to perform at least one operation. The at least one operation may include an operation of setting a number of user requests. The at least one operation may include an operation of setting a number of instances of at least one AI model. The at least one operation may include an operation of executing an instance of the at least one AI model according to the number of instances of the at least one AI model. The at least one operation may include an operation of generating user requests according to the number of user requests. The at least one operation may include an operation of checking responses corresponding to each of the user requests based on processing the user requests based on instances of the at least one AI model. The at least one operation may include generating training data that includes the number of user requests and the number of instances of the at least one AI model as input values ​​and response times corresponding to the responses as output values.

[0139] The embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0140] The term "module" used in the embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0141] One embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0142] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0143] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. A method for executing an instance of at least one artificial intelligence (AI) model, An operation of determining a response time AI model for at least one AI model associated with a first service, wherein the response time AI model is configured to receive, as input values, a number of user requests for the first service and a number of instances of each of the at least one AI model, and output a response time; An operation of checking the number of instances of at least one AI model, the response time AI model, and a loss function set based on the acceptable response time; An operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function using the number of user requests for the first service; and An action for performing at least one operation for executing instances of at least one AI model according to at least one optimal number of instances. A method including:

2. In paragraph 1, A method wherein the loss function includes a component based on the difference between the number of user requests input into the response time AI model and the allowed response time.

3. In paragraph 1 or 2, A method wherein the loss function comprises a component based on the difference and a sum of the number of instances of the at least one AI model.

4. In any one of paragraphs 1 to 3, The above loss function is based on mathematical expression 1, and mathematical expression 1 is as follows: Above is the number of instances of at least one AI model, The above SLO is the above acceptable response time, Above is the number of each type of user request, Above is the response time function, Above is the sum of the number of instances of at least one AI model, The above φ function is a component based on the above difference, and is based on mathematical expression 2, and mathematical expression 2 is as follows: The above method where ρ is a constant.

5. In any one of paragraphs 1 to 4, The operation of determining at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the above loss function is as follows: An operation of performing the optimization procedure for the number of instances of the at least one AI model with respect to the loss function to determine at least one optimal number of instances for each of the at least one AI model. A method including:

6. In any one of paragraphs 1 to 5, The above optimization procedure is a method including gradient descent.

7. In any one of paragraphs 1 to 6, An operation of performing at least one operation for executing instances of at least one AI model according to at least one optimal number of instances, An operation of executing instances of at least one AI model according to the at least one optimal instance number. A method including:

8. In any one of paragraphs 1 to 7, An operation of performing at least one operation for executing instances of at least one AI model according to at least one optimal number of instances, An action of requesting an external electronic device to execute at least one optimal number of instances of instances of at least one AI model. A method including:

9. In any one of paragraphs 1 to 8, An action for receiving a command to set the above allowable response time. How to include more.

10. In a storage medium storing at least one computer-readable instruction, The at least one instruction, when executed individually or collectively by at least a portion of one or more processors (120) comprising processing circuitry of the electronic device, causes the electronic device to perform at least one operation; At least one of the above actions: An operation of determining a response time AI model for at least one artificial intelligence (AI) model associated with a first service, wherein the response time AI model is configured to receive, as input values, a number of user requests for the first service and a number of instances of each of the at least one AI model, and output a response time; An operation of checking the number of instances of at least one AI model, the response time AI model, and a loss function set based on the acceptable response time; An operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function using the number of user requests for the first service; and An action for performing at least one operation for executing instances of at least one AI model according to at least one optimal number of instances. A storage medium containing.

11. In paragraph 10, A storage medium wherein the loss function includes a component based on the difference between the number of user requests input into the response time AI model and the allowed response time.

12. In clause 10 or 11, A storage medium wherein the loss function comprises a component based on the difference and a sum of the number of instances of the at least one AI model.

13. In any one of paragraphs 10 to 12, The above loss function is based on mathematical expression 1, and mathematical expression 1 is as follows: Above is the number of instances of at least one AI model, The above SLO is the above acceptable response time, Above is the number of each type of user request, Above is the response time function, Above is the sum of the number of instances of at least one AI model, The above φ function is a component based on the above difference, and is based on mathematical expression 2, and mathematical expression 2 is as follows: The above ρ is a storage medium where the constant is.

14. In any one of paragraphs 10 to 13, The operation of determining at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the above loss function is as follows: An operation of performing the optimization procedure for the number of instances of the at least one AI model with respect to the loss function to determine at least one optimal number of instances for each of the at least one AI model. A storage medium containing.

15. In electronic devices, One or more processors (120) including processing circuitry; and Contains a memory (130) that stores instructions, The instructions, when individually or collectively executed by at least some of the one or more processors, cause the electronic device to perform at least one operation; At least one of the above actions: An operation of determining a response time AI model for at least one artificial intelligence (AI) model associated with a first service, wherein the response time AI model is configured to receive as input values ​​a number of user requests for the first service and a number of instances of each of the at least one AI model, and output a response time; An operation of checking the number of instances of at least one AI model, the response time AI model, and a loss function set based on the acceptable response time; An operation of identifying at least one optimal number of instances for each of the at least one AI model by performing an optimization procedure for the loss function using the number of user requests for the first service; and An action for performing at least one operation for executing instances of at least one AI model according to at least one optimal number of instances. An electronic device comprising:

Citation Information

Patent Citations

  • Method for service capacity expansion and shrinkage and related equipment

    CN112000459A

  • Cloud-based deep learning task execution time prediction system and method

    KR102504939B1

  • Method and apparatus to operate search system through response time using machine learning

    KR102525918B1

  • An exterior module of wall

    KR102705193B1