Device and method for distributing ai computation task among ai data center, edge ai infrastructure, and terminal in wireless communication system

The method optimizes AI computational task distribution among AI data centers, edge AI infrastructure, and user equipment by considering resource information and computational capabilities, addressing the overabundance of AI computing resources and minimizing service delays.

WO2026054444A1PCT designated stage Publication Date: 2026-03-12SK TELECOM CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Conventional AI computing resources are expanding independently in each field, leading to overabundance and waste, and there is a need for efficient distribution of AI computational tasks among AI data centers, edge AI infrastructure, and user equipment in wireless communication systems.

Method used

A method and device for distributing AI computational tasks by obtaining resource information, confirming available AI computation capabilities, and distributing tasks based on these capabilities and necessary resource information among AI data centers, edge AI infrastructure, and user equipment, considering time requirements and computational efficiencies.

Benefits of technology

Enables efficient distribution of AI computational tasks, optimizing resource utilization and minimizing service delays by leveraging the strengths of each component in the wireless communication system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025013384_12032026_PF_FP_ABST
    Figure KR2025013384_12032026_PF_FP_ABST
Patent Text Reader

Abstract

According to an embodiment of a first aspect of the present invention, a method for distributing artificial intelligence (AI) computation performed by an AI data center (DC) comprises the steps of: acquiring resource information required for the AI computation; identifying an available AI computation capability for each of an edge AI infrastructure and an AI terminal connected thereto, the edge AI infrastructure being disposed at an edge node of a network including the AI data center; and distributing the AI computation to each of the AI data center, the edge AI infrastructure, and the AI terminal by comparing the identified available AI computation capability for each of the edge AI infrastructure and the AI terminal, an available AI computation capability of the AI data center, and the required resource information.
Need to check novelty before this filing date? Find Prior Art

Description

Device and method for distributing AI computational tasks among AI data centers, edge AI infrastructure, and terminals in a wireless communication system.

[0001] The present disclosure relates generally to a wireless communication system, and more particularly to a device and method for distributing AI computational tasks between an artificial intelligence (AI) data center, an edge AI infrastructure, and user equipment (UE) in a wireless communication system.

[0002] For reference, this application claims priority from Korean Patent Application No. 10-2024-0119262, filed September 3, 2024. The entire contents of that application, which serves as the basis for this priority claim, are incorporated herein by reference.

[0003] With the rapid growth of AI (artificial intelligence) technology and services, the need for AI data centers (DCs) is also increasing, and interest in Edge AI infrastructure to relieve the centralization of AI operations and minimize AI service delays is also growing.

[0004] In addition, AI terminals capable of AI computation are being released to reduce dependence on AI data centers or edge AI infrastructure for AI computation and to improve AI service quality.

[0005] Therefore, AI computing resources are expanding and becoming more widespread, and they are expected to become more diverse and ubiquitous. However, conventional technologies are expanding and disseminating independently in each field, ultimately resulting in an overabundance and waste of AI computing resources.

[0006] Based on the discussion described above, the present disclosure provides a device and method for distributing AI computational tasks between an AI (artificial intelligence) data center, an edge AI infrastructure, and a user equipment (UE) in a wireless communication system.

[0007] Additionally, the present disclosure provides a device and method for interconnection and authentication between an AI data center, an edge AI infrastructure, and a terminal in a wireless communication system.

[0008] A distribution method for AI computation performed by an AI (artificial intelligence) data center (DC) according to one embodiment of the first aspect of the present invention includes the steps of: obtaining resource information necessary for the AI ​​computation; confirming available AI computation capabilities for each of an edge AI infrastructure and an AI terminal connected to the edge AI infrastructure located at an edge of a network including the AI ​​data center; and distributing the AI ​​computation to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal based on the confirmed available AI computation capabilities for each of the edge AI infrastructure and the AI ​​terminal, the available AI computation capabilities in the AI ​​data center, and the necessary resource information.

[0009] The above-mentioned necessary resource information may include the prediction time required for the AI ​​operation. At this time, in the process of distributing the AI ​​operation, the AI ​​operation may be distributed based on at least one of the time required to transmit the AI ​​model for the AI ​​operation from the AI ​​data center to the edge AI infrastructure or the AI ​​terminal, whether or not information required for the AI ​​operation is transmitted to the edge AI infrastructure or the AI ​​terminal, and the time required to transmit the required information.

[0010] In the process of distributing the AI ​​operation, the AI ​​operation may be distributed in consideration of the time required for each of the AI ​​data center and the edge AI infrastructure to perform the AI ​​operation, depending on the distribution of the AI ​​operation, but the AI ​​operation may be distributed such that the time required for the AI ​​data center to perform the AI ​​operation is greater by a preset amount of time compared to the edge AI infrastructure.

[0011] In the process of distributing the above AI operation, the AI ​​operation can be distributed by taking into consideration the time required for distributing the AI ​​operation and the time for performing the AI ​​operation after distributing the AI ​​operation.

[0012] In the process of distributing the above AI operations, the distribution order may be such that the AI ​​data center comes first, followed by the edge AI infrastructure and the AI ​​terminal sequentially.

[0013] Some of the AI ​​operations distributed to the edge AI infrastructure may be distributed to the above AI terminal.

[0014] In the process of redistributing to the AI ​​terminals, the AI ​​operations are redistributed in consideration of the time required for each of the edge AI infrastructure and the AI ​​terminals to perform the AI ​​operations, and the AI ​​operations can be distributed such that the time required for the edge AI infrastructure to perform the AI ​​operations is greater by a preset amount of time compared to the AI ​​terminals.

[0015] The required resource information may be provided to the edge AI infrastructure or each AI terminal according to the distribution of the AI ​​operation. At this time, when the AI ​​service corresponding to the AI ​​operation is terminated on the AI ​​terminal, the required resource information provided to the edge AI infrastructure or each AI terminal may be deleted.

[0016] A method for distributing AI computations performed by edge AI (artificial intelligence) infrastructure according to another embodiment of the first aspect of the present invention includes the steps of transmitting available AI computational capacity for the edge AI infrastructure to an AI data center (DC) connected to the edge AI infrastructure, and the steps of distributing the AI ​​computations from the AI ​​data center. At this time, the distributed AI computations are distributed to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminals based on the available AI computational capacity confirmed for each of the edge AI infrastructure and the AI ​​terminals connected to the edge AI infrastructure, the available AI computational capacity in the AI ​​data center, and resource information required for the AI ​​computations.

[0017] According to one embodiment of the second aspect of the present invention, an AI (artificial intelligence) data center (DC) includes a transceiver and a control unit operably connected to the transceiver, wherein the control unit obtains resource information necessary for AI computation, verifies available AI computation capabilities for each of an edge AI infrastructure disposed at an edge of a network including the AI ​​data center and AI terminals connected to the edge AI infrastructure, and distributes the AI ​​computation to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminals based on the verified available AI computation capabilities for each of the edge AI infrastructure and the AI ​​terminals, the available AI computation capabilities in the AI ​​data center, and the necessary resource information.

[0018] According to another embodiment of the second aspect of the present invention, an edge AI (artificial intelligence) infrastructure comprises a transceiver and a control unit operably connected to the transceiver, wherein the control unit transmits available AI computational power for the edge AI infrastructure to an AI data center (DC) connected to the edge AI infrastructure, and distributes AI computations from the AI ​​data center, wherein the distributed AI computations are distributed to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal, based on the available AI computational power confirmed for each of the edge AI infrastructure and the AI ​​terminal connected to the edge AI infrastructure, the available AI computational power in the AI ​​data center, and resource information required for the AI ​​computation.

[0019] A non-transitory computer-readable recording medium storing at least one computer-executable instruction according to one embodiment of the third aspect of the present invention, wherein the at least one instruction, when executed by a processor, causes the processor to perform a method including: obtaining resource information necessary for the AI ​​operation; confirming available AI operation capabilities for each of an edge AI infrastructure and an AI terminal connected to the edge AI infrastructure located at the edge of a network including the AI ​​data center; and distributing the AI ​​operation to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal based on the confirmed available AI operation capabilities for each of the edge AI infrastructure and the AI ​​terminal, the available AI operation capabilities at the AI ​​data center, and the necessary resource information.

[0020] A non-transitory computer-readable recording medium storing at least one computer-executable instruction according to another embodiment of the third aspect of the present invention, wherein the at least one instruction, when executed by a processor, comprises the steps of: transmitting available AI computational power for an edge AI infrastructure to an AI data center (DC) connected to an edge AI infrastructure; and distributing the AI ​​computation from the AI ​​data center, wherein the distributed AI computation is distributed to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal by comparing the available AI computational power confirmed for each of the edge AI infrastructure and the AI ​​terminal connected to the edge AI infrastructure with the available AI computational power in the AI ​​data center and resource information required for the AI ​​computation, wherein the processor performs a method of distributing the AI ​​computation to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal.

[0021] According to one embodiment of the fourth aspect of the present invention, a computer program stored in a non-transitory computer-readable recording medium, wherein the computer program, when executed by a processor, comprises instructions for causing the processor to perform a method including the steps of: obtaining resource information necessary for the AI ​​operation; confirming available AI operation capabilities for each of an edge AI infrastructure and an AI terminal connected to the edge AI infrastructure located at the edge of a network including the AI ​​data center; and distributing the AI ​​operation to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal based on the confirmed available AI operation capabilities for each of the edge AI infrastructure and the AI ​​terminal, the available AI operation capabilities in the AI ​​data center, and the necessary resource information.

[0022] According to another embodiment of the fourth aspect of the present invention, a computer program stored in a non-transitory computer-readable recording medium, wherein the computer program, when executed by a processor, comprises the steps of: transmitting available AI computational power for an edge AI infrastructure to an AI data center (DC) connected to an edge AI infrastructure; and receiving the AI ​​computation from the AI ​​data center, wherein the distributed AI computation comprises instructions for causing the processor to perform a method of distributing the AI ​​computation to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal by comparing the available AI computational power confirmed for each of the edge AI infrastructure and the AI ​​terminal connected to the edge AI infrastructure with the available AI computational power in the AI ​​data center and resource information required for the AI ​​computation.

[0023] Devices and methods according to various embodiments of the present disclosure enable distribution of AI computational tasks among an AI data center, an edge AI infrastructure, and a user equipment (UE) by interconnecting and authenticating between the AI ​​data center, an edge AI infrastructure, and a user equipment (UE) in a wireless communication system.

[0024] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from the description below.

[0025] FIG. 1 illustrates an example of a network structure between an AI data center, an edge AI infrastructure, and an AI terminal according to various embodiments of the present disclosure.

[0026] FIG. 2 illustrates a signal flow diagram for interconnection and authentication between an AI data center, an edge AI infrastructure, and a terminal according to one embodiment of the present disclosure.

[0027] FIG. 3 illustrates a signal flow diagram for distributing AI computational tasks among an AI data center, an edge AI infrastructure, and a terminal, according to one embodiment of the present disclosure.

[0028] FIG. 4 illustrates an operation method of an AI data center according to one embodiment of the present disclosure.

[0029] FIG. 5 illustrates a configuration diagram of an AI data center, or edge AI infrastructure, according to various embodiments of the present disclosure.

[0030] The terms used in this disclosure are used only to describe specific embodiments and may not be intended to limit the scope of other embodiments. The singular expression may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this disclosure. Terms defined in general dictionaries among the terms used in this disclosure may be interpreted as having the same or similar meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this disclosure. In some cases, even if a term is defined in this disclosure, it cannot be interpreted to exclude embodiments of the present disclosure.

[0031] The various embodiments of the present disclosure described below illustrate a hardware-based approach as an example. However, since the various embodiments of the present disclosure include techniques utilizing both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.

[0032] Additionally, in the detailed description and claims of the present disclosure, “at least one of A, B, and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” Additionally, “at least one of A, B, or C” or “at least one of A, B, and / or C” can mean “at least one of A, B, and C.”

[0033] The present disclosure relates to a device and method for distributing AI computational tasks among an AI (artificial intelligence) data center (DC), an edge AI infrastructure, and user equipment (UE) in a wireless communication system. Specifically, the present disclosure describes a technology for distributing AI computational tasks among an AI data center, an edge AI infrastructure, and a UE by interworking and authenticating the AI ​​data center, the edge AI infrastructure, and the UE in a wireless communication system.

[0034] The terms used in the following description, including terms referring to signals, channels, control information, network entities, and device components, are provided for convenience of explanation. Therefore, the present disclosure is not limited to the terms described below, and other terms with equivalent technical meanings may be used.

[0035] Additionally, while this disclosure describes various embodiments using terminology used in certain communication standards (e.g., 3rd Generation Partnership Project (3GPP)), these are merely illustrative examples. The various embodiments of this disclosure can be easily modified and applied to other communication systems.

[0036] FIG. 1 illustrates an example of a network structure between an AI data center, an edge AI infrastructure, and an AI terminal according to various embodiments of the present disclosure.

[0037] Referring to Figure 1, a data center (e.g., an AI data center) is connected to multiple edge AI infrastructures. Furthermore, each of the multiple edge AI infrastructures is connected to multiple terminals (e.g., AI terminals).

[0038] Through this network structure, Figure 1 illustrates how AI computational tasks can be distributed to data centers, edge AI infrastructure, and AI terminals.

[0039] FIG. 2 illustrates a signal flow diagram for interconnection and authentication between an AI data center, an edge AI infrastructure, and a terminal according to one embodiment of the present disclosure.

[0040] AI data centers can be interconnected with multiple edge AI infrastructures, and for this purpose, the AI ​​data centers and each edge AI infrastructure can send and receive messages for interconnection and authentication.

[0041] Referring to FIG. 2, the edge AI infrastructure can attempt to connect to the AI ​​data center by radiating its own information (e.g., ID, IP address, etc.) to the connected AI data center (201).

[0042] In one embodiment, edge data centers, edge AI infrastructure, and AI terminals may each be assigned a unique ID, which is globally unique. For example, the unique ID may use an international mobile subscriber identity (IMSI), mobile subscriber ISDN number (MSISDN), or the like.

[0043] According to one embodiment, in the case of operation (201), if there is an address (e.g., IP address, etc.) of an AI data center stored in the edge AI infrastructure, an attempt may be made to connect to the address of the AI ​​data center.

[0044] In another embodiment, in the case of operation (201), if there is no address (e.g., IP address, etc.) of the AI ​​data center stored in the edge AI infrastructure, the edge AI infrastructure may attempt to connect by radiating its own information to all communication network paths to which it is connected.

[0045] When the AI ​​data center receives a connection attempt by operation (201), it can respond to it within a certain period of time (203).

[0046] In some embodiments, the scheduled time can be set in seconds and can be set for each edge AI infrastructure. For example, the scheduled time can be set in units of 1, 2, 5, or 10 seconds.

[0047] In one embodiment, the edge AI infrastructure may retry the linkage attempt of operation (201) if it does not receive a response to operation (203) from the AI ​​data center within a certain amount of time.

[0048] In one embodiment, if a response is not received from the AI ​​data center after multiple connection attempts, the edge AI infrastructure may query surrounding edge AI infrastructures for the address of the AI ​​data center. For example, the number of connection attempts may be set to 2, 4, 8, 10, or 20.

[0049] In one embodiment, if the edge AI infrastructure queries the surrounding edge AI infrastructures for the AI ​​data center address but does not receive a response, operation (201) may be performed again.

[0050] An AI terminal may attempt to connect to an edge AI infrastructure by radiating its own information (e.g., ID, IP address, etc.) to the connected edge AI infrastructure (207).

[0051] According to one embodiment, in the case of operation (207), if there is an address (e.g., IP address, etc.) of the AI ​​edge infrastructure stored in the AI ​​terminal, an attempt may be made to link with the address of the corresponding AI data center.

[0052] In another embodiment, in the case of operation (207), if there is no address (e.g., IP address, etc.) of the AI ​​edge infrastructure stored in the AI ​​terminal, the AI ​​terminal may attempt to connect by radiating its own information to all communication network paths to which the AI ​​terminal is connected.

[0053] When the edge AI infrastructure receives a connection attempt by operation (207), it can respond to it within a certain period of time (209).

[0054] In one embodiment, the scheduled time can be set in seconds and can be set for each AI terminal. For example, the scheduled time can be set in units of 1, 2, 5, or 10 seconds.

[0055] In one embodiment, if the AI ​​terminal does not receive a response to operation (209) from the edge AI infrastructure within a certain period of time, the AI ​​terminal may retry the linking attempt of operation (207).

[0056] In one embodiment, if a response is not received from the edge AI infrastructure even after multiple connection attempts, the AI ​​terminal may query surrounding AI terminals for the address of the edge AI infrastructure. For example, the number of connection attempts may be set to 2, 4, 8, 10, or 20.

[0057] In one embodiment, if an AI terminal queries surrounding AI terminals for an edge AI infrastructure address but does not receive a response, operation (207) may be performed again.

[0058] FIG. 3 illustrates a signal flow diagram for distributing AI computational tasks among an AI data center, an edge AI infrastructure, and a terminal, according to one embodiment of the present disclosure.

[0059] To distribute AI computation among AI data centers, edge AI infrastructure, and terminals, each device's AI computational power needs to be shared. However, AI computation is distributed in the following order: AI data centers, edge AI infrastructure, and then AI terminals. Therefore, AI computational power can be shared in the following order: AI terminals, edge AI infrastructure, and then AI data centers.

[0060] Edge AI infrastructure can share computational power with AI data centers (301). Computational power can refer to the amount of computational power available and the power consumed by computing equipment.

[0061] The units of acceptable computational capacity can be i) bit per second (bit operations per second), ii) FLOPS (floating point operations per second), iii) GFLOPS (Giga Floating point OPerations per Second), iv) TFLOPS (TeraFloating point Operations per Second), and TOPS (TeraOperations per Second).

[0062] FLOPS (Floating Point Operations Per Second) is a unit commonly used to quantify computer performance. It stands for floating point operations per second, and is based on the number of floating point operations a computer can perform per second. Operations in FLOPS include arithmetic operations, as well as operations such as root, log, and exponential, each of which can be calculated as a single operation.

[0063] TOPS (Tera Operations Per Second) is a unit used to numerically express the performance of a computer, similar to FLOPS. If FLOPS is a unit for floating point operations, TOPS can be a unit based on integer processing.

[0064] The power consumption of computing equipment is expressed in units of power consumption, such as kW (kilo watt) and MW (mega watt).

[0065] Edge AI infrastructure assigns numbers to computing power in units of a certain number to share computing power between devices, and AI terminals can report the computing power corresponding to that number.

[0066] For example, as an example of a unit of computational power, an AI terminal can report [Calculation level 1: 100 GFLOPS], [Calculation level 2: 500 GFLOPS], [Calculation level 3: 100 TFLOPS], [Calculation level 4: 500 TFLOPS], [Calculation level 5: 100 TOPS], etc.

[0067] As another example, the AI ​​terminal can inform [Calculation level 1: 100KW], [Calculation level 2: 500 KW], [Calculation level 3: 100 MW], [Calculation level 4: 500 MW]… etc.

[0068] AI terminals can share computational power with edge AI infrastructure or AI data centers (303, 305).

[0069] In operation (303, 305), when sharing computational power, the AI ​​terminal can share information such as the number of CPU cores, the number of CPU threads, the clock per CPU core (e.g., 2 GHz), the number of GPU stream processors (cores), and the GPU clock (e.g., 1700 MHz) (first method).

[0070] In addition, in operations (303, 305), when sharing computational capabilities, the AI ​​terminal can clearly transmit the computational amount that the AI ​​terminal can accept using computational units. As units of acceptable computational amounts, i) bit per second (bit computation per second), ii) FLOPS (floating point operations per second), iii) GFLOPS (Giga Floating point Operations per Second), iv) TFLOPS (TeraFloating point Operations per Second), and TOPS (TeraOperations per Sencd) can be used (second method).

[0071] In order to share computational power, AI terminals assign numbers to each other by setting the computational power in units of a certain number, and AI terminals can inform each other of the computational power corresponding to that number.

[0072] For example, as an example of a computational power unit, an AI terminal can be configured as follows: [Calculation level 1: CPU 1 core, 2 threads, 1GHz or higher, GPU stream processor 1000, clock 1000MHz or higher], [Calculation level 2: CPU 1 core, 2 threads, 1GHz or higher, GPU stream processor 1500, clock 1500MHz or higher], [Calculation level 3: CPU 2 cores, 4 threads, 1GHz or higher, GPU stream processor 2000, clock 2000MHz or higher], [Calculation level 4: CPU 2 cores, 4 threads, 2GHz or higher, GPU stream processor 4000, clock 2000MHz or higher]… (3rd method).

[0073] For another example, as an example of a unit of computational power, an AI terminal can be configured as [Calculation level 1: computational data 128 Mbit / sec], [Calculation level 2: computational data 256 Mbit / sec], [Calculation level 3: computational data 512 Mbit / sec], [Calculation level 4: computational data 1 Gbit / sec]… (4th method).

[0074] As another example, as an example of a unit of computational power, an AI terminal can be notified as [Calculation level 1: 100 GFLOPS], [Calculation level 2: 500 GFLOPS], [Calculation level 3: 100 TFLOPS], [Calculation level 4: 500 TFLOPS], [Calculation level 5: 100 TOPS], etc. (Fifth method).

[0075] In one embodiment, the edge AI infrastructure can transmit the computational power of the AI ​​terminal to the AI ​​data center when needed.

[0076] Specifically, according to one embodiment, when the edge AI infrastructure transmits the computational power of the AI ​​terminal to the AI ​​data center, the computational power transmitted by the AI ​​terminal defined in operations (303, 305) to the edge AI infrastructure can be transmitted as is to the AI ​​data center.

[0077] In another embodiment, when the edge AI infrastructure transmits the computational power of the AI ​​terminal to the AI ​​data center, if the computational power of the AI ​​terminal is shared from the AI ​​terminal in the first or second manner of operation (303, 305), the computational power of the AI ​​terminal can be transmitted to the AI ​​data center by converting it into the third, fourth, and fifth manners.

[0078] An AI terminal can request an AI service or AI calculation required for an AI service from an AI data center (307).

[0079] The AI ​​data center can calculate the time for AI computation in response to a request for action (307) and decide whether to distribute AI computation to the edge AI infrastructure (309).

[0080] Specifically, the AI ​​data center can calculate, prior to AI computation, i) the time required to transmit an AI model, ii) whether information required for AI computation is transmitted (because the Edge AI infrastructure or AI terminal can directly receive the required information from an external source), and iii) the time required to transmit information required for AI computation. In one embodiment, if the information required for AI computation is located outside the AI ​​data center, the AI ​​data center can transmit a path (e.g., an IP address or URL) through which the information can be received to the Edge AI infrastructure.

[0081] Specifically, the AI ​​data center can perform calculations according to Equations 1 and 2.

[0082]

[0083]

[0084] Here, x is the expected AI computational distribution from the AI ​​data center to the edge AI infrastructure, and y is the time required for the edge AI infrastructure to receive the AI ​​model and information required for AI computation from the AI ​​data center and receive x amount of AI computational volume to complete the computation.

[0085]

[0086] Here, z is the time required for the AI ​​data center to complete the remaining AI operations after distributing x amount of AI computation to the edge AI infrastructure.

[0087]

[0088] In the case of mathematical equations 1 and 2, since the calculation is performed while changing the value of x, y and z have multiple values ​​rather than a single value. At this time, we find x such that y + △ = z.

[0089] △ can be defined as the 'AI computation distribution margin'. △ can be defined arbitrarily, for example, 1ms, 10ms, 100ms, 1s... The reason for defining △ is that AI computation is more efficient when performed in an AI data center than in an edge AI infrastructure, and if a problem occurs in the edge AI infrastructure and computation cannot be performed for x, there is a burden in terms of service delay when AI DC must recompute x. Therefore, this margin can be set. The larger △ is set, the more computational effort AI DC must make, but the more stable the service can be.

[0090] Additionally, separately, the AI ​​data center can also compute Equation 3.

[0091]

[0092] Here, y is the time required for the edge AI infrastructure to receive AI models and information required for AI operations from the AI ​​data center and to distribute x amount of AI operations to complete the operations.

[0093] After calculating a in mathematical expression 3, compare the sizes of a and w, and when a is smaller than w (a <w)에만 AI 데이터 센터는 AI 연산 분배를 수행할 수 있다(311). 이때 w는 'AI 연산 분배비'로 정의할 수 있다. w는 AI 분배를 위해 사전 작업 시간 대비 AI 분배 후 연산시간의 비율로 w를 조절함으로써 AI 분배를 위해 사전 작업이 할애되는 비율을 조절할 수 있다.

[0094] In some embodiments, w can be 1 / 2, 1 / 5, 1 / 10, 1 / 100, 1 / 1000, etc. For example, if w is 1 / 2, it may mean that the AI ​​DC does not perform AI computation distribution if the time it takes for the edge AI infrastructure to distribute AI computations, receive the AI ​​model and the information required for the AI ​​computation, and complete the AI ​​computation is more than half the time it takes to receive the AI ​​model and the information required for the AI ​​computation.

[0095] Edge AI infrastructure can receive AI computation distribution from an AI data center (311) and decide whether to distribute the computation to an AI terminal (313).

[0096] Specifically, the edge AI infrastructure can calculate, before AI computation, i) the time required to transmit an AI model, ii) whether information required for AI computation is transmitted (because the edge AI infrastructure or AI terminal can directly receive the necessary information from an external source), and iii) the time required to transmit information required for AI computation.

[0097] Specifically, the edge AI infrastructure can perform calculations according to Equations 4 and 5.

[0098]

[0099] Here, s is the expected AI computational distribution from edge AI infrastructure to AI terminals.

[0100] r is the time required for an AI terminal to receive information necessary for AI model and AI operation from the edge AI infrastructure, and to complete the operation by distributing s amount of AI operation.

[0101]

[0102] Here, s is the expected AI computational distribution from edge AI infrastructure to AI terminals.

[0103] t is the time required for the edge AI infrastructure to complete the remaining AI operations after distributing s of AI computational capacity to AI terminals.

[0104]

[0105] In the case of mathematical equations 4 and 5, since the calculation is performed while changing the value of s, r and t have multiple values ​​rather than a single value. At this time, s is found so that r + △ = t.

[0106] △ can be defined as the 'AI operation distribution margin'. △ can be defined arbitrarily, for example, 1ms, 10ms, 100ms, 1s... The reason for defining △ is that AI operation is more efficient when performed on the edge AI infrastructure than on the AI ​​terminal, and if a problem occurs in the AI ​​terminal and x cannot be calculated, there is a burden in terms of service delay when x must be re-calculated on the edge AI infrastructure. Therefore, this margin can be set. The larger △ is set, the more computational effort the edge AI infrastructure must make, but the more stable the service can be provided.

[0107] Alternatively, θ, the 'AI service margin (tentative name)' can be defined. Based on this, it is possible to determine whether to distribute computation from the Edge AI infrastructure to the AI ​​terminal. In the case of AI terminals, while using AI services, they may also use other services at the same time, so using all of the computational capacity of the AI ​​terminal only for the AI ​​service may not be good in terms of service quality. Accordingly, by defining the 'AI service margin (tentative name)', if r is greater than θ, s can be calculated so that r becomes less than or equal to θ, and the computational load can be distributed to the AI ​​terminal. According to one embodiment, θ can be defined as 1ms, 10ms, 100ms, 1s... etc.

[0108] In one embodiment, an AI data center can determine whether to simultaneously distribute AI computations to edge AI infrastructure and AI terminals. To do so, the AI ​​data center can calculate the computational capacity ratio between the edge AI infrastructure and AI terminals. For example, the computational capacity ratio ε can be defined as Equation 6.

[0109]

[0110] Specifically, if an AI terminal shares its computing power solely with the edge AI infrastructure, the edge AI infrastructure can calculate the computing power ratio and report it to the AI ​​data center. Furthermore, if the AI ​​terminal shares its computing power with the AI ​​data center, the AI ​​data center can calculate the computing power ratio.

[0111] To explain more specifically using formulas, AI data centers can calculate i) the time required to transmit AI models (the longer time between the edge AI infrastructure and the time it takes for the AI ​​terminal to receive the AI ​​model), ii) whether information required for AI computation is transmitted (because the edge AI infrastructure or the AI ​​terminal can directly receive the necessary information from an external source), and iii) the time required to transmit information required for AI computation (the longer time between the edge AI infrastructure and the time it takes for the AI ​​terminal to receive information).

[0112] The AI ​​Day Center can calculate the computational allocation to edge AI infrastructure and AI terminals based on mathematical equations 7 and 8 below.

[0113]

[0114] Here, g = expected AI computational distribution from AI DC to Edge AI infra and AI terminals, e is the time required for Edge AI infra and AI terminals to receive AI models and information necessary for AI computation from AI DC and receive AI computational volume equivalent to g to complete the computation.

[0115]

[0116] Here, f is the time required for AI DC to complete the remaining AI operations after distributing the AI ​​computation amount of g to the Edge AI infrastructure and AI terminals.

[0117]

[0118] In the case of mathematical expressions 7 and 8, since the calculation is performed while changing the value of g, e and f have multiple values ​​rather than a single value. At this time, we find x such that e + △(AI operation distribution margin) = f.

[0119] Additionally, AI data centers can also compute Equation 9.

[0120]

[0121] Here, e is the time required for Edge AI infrastructure and AI terminals to receive AI models and information required for AI operations from AI DC and to distribute AI operations equivalent to g to complete the operations.

[0122] According to mathematical formula 9, after calculating h, compare the sizes of h and w, and when h is smaller than w (h <w)에만 AI 연산 분배를 수행할 수 있다. 이때, w는 'AI 연산 분배비'로 w는 AI 분배를 위해 사전 작업 시간 대비 AI 분배 후 연산시간의 비율로 w를 조절함으로써 AI 분배를 위해 사전 작업이 할애되는 비율을 조절할 수 있다. 일 실시 예에 따라, w는 1 / 2, 1 / 5, 1 / 10, 1 / 100, 1 / 1000… 등을 사용할 수 있다. 수학식 9의 설명에 기재한 w는 수학식 3의 설명에 기재한 w와 동일한 의미일 수 있다.

[0123] The AI ​​terminal can terminate the use of AI services (317).

[0124] In response to operation (317), the AI ​​terminal may delete the AI ​​model and information required for AI operations (319) on the AI ​​terminal. Specifically, when the AI ​​terminal has completed using the AI ​​service and has received AI operations from an AI data center or edge AI infrastructure, the AI ​​terminal may delete the AI ​​model and information required for AI operations and then reply to the AI ​​data center regarding whether the deletion has been completed (not shown).

[0125] The AI ​​terminal can notify the AI ​​data center (321) whether or not to terminate the use of the AI ​​service in response to the action (317).

[0126] When the AI ​​data center distributes AI operations to the edge AI infrastructure in response to the action (321), it may request the edge AI infrastructure to delete the relevant AI model and information required for the AI ​​operation (323).

[0127] Edge AI infrastructure can delete information required for AI models and AI operations (325).

[0128] The edge AI infrastructure can respond to the action (325) by returning the deletion result to the AI ​​data center (327).

[0129] FIG. 4 illustrates an operation method of an AI data center according to one embodiment of the present disclosure.

[0130] Referring to FIG. 4, the AI ​​data center can perform linkage and authentication with the edge AI infrastructure and AI terminal based on the identifier (ID) of each of the edge AI infrastructure, AI terminal, and AI DC (401).

[0131] In one embodiment, the AI ​​data center may receive an identifier for the edge AI infrastructure from the AI ​​DC connected to the edge AI infrastructure to perform linkage and authentication with the edge AI infrastructure and AI terminals. Furthermore, upon receiving the identifier for the edge AI infrastructure, the AI ​​data center may respond to the edge AI infrastructure within a certain timeframe.

[0132] Additionally, when the edge AI infrastructure and AI DC are linked, the AI ​​data center can receive information about AI terminals connected to the edge AI infrastructure. The information about the AI ​​terminals may include their identifiers.

[0133] The AI ​​data center can receive AI computing power for AI computation distribution from edge AI infrastructure and AI terminals (403). In one embodiment, the AI ​​computing power can be determined based on the computational capacity that each edge AI infrastructure and AI terminal can accommodate.

[0134] When an AI data center receives an AI operation request from an AI terminal, it can calculate the time for the AI ​​operation (405).

[0135] In one embodiment, when an AI data center receives an AI operation request from an AI terminal, in order to calculate the time for the AI ​​operation, the AI ​​data center may calculate the time required to transmit an AI model, whether information required for the AI ​​operation is transmitted, or the time required to transmit information required for the AI ​​operation before the AI ​​operation is performed.

[0136] In one embodiment, the AI ​​data center may determine the expected AI computational allocation from the AI ​​DC to the edge AI infrastructure based on the calculations described above.

[0137] In one embodiment, the AI ​​data center may include a process for distributing AI computations when the time it takes for the edge AI infrastructure to distribute AI computations, receive AI models and information necessary for AI computations, and complete AI computations compared to the time it takes to receive AI models and information necessary for AI computations exceeds a threshold.

[0138] The AI ​​data center can decide whether to distribute AI computations to edge AI infrastructure based on the calculation of the time for AI computation (407).

[0139] In one embodiment, the AI ​​data center may determine whether to distribute AI computation to the edge AI infrastructure and AI terminals simultaneously based on the ratio of computational capabilities of the edge AI infrastructure and AI terminals, in order to determine whether to distribute AI computation to the edge AI infrastructure.

[0140] The AI ​​data center can receive notification from the AI ​​terminal regarding the termination of AI service use. Upon receiving notification, the AI ​​data center can delete the edge AI infrastructure and any information necessary for AI models and AI operations on the AI ​​terminal.

[0141] FIG. 5 illustrates a configuration diagram of an AI data center, or edge AI infrastructure, according to various embodiments of the present disclosure.

[0142] Referring to FIG. 5, an AI data center or edge AI infrastructure (500) may include at least one processor (510), a memory (520), and a communication device (530) that is connected to a network and performs communication. In addition, the AI ​​data center or edge AI infrastructure may further include an input interface device (540), an output interface device (550), a storage device (560), etc. Each component included in the AI ​​data center or edge AI infrastructure may be connected by a bus (570) and communicate with each other.

[0143] However, each component included in the AI ​​data center or edge AI infrastructure (500) may be connected through individual interfaces or individual buses centered around the processor (510), rather than through a common bus (570). For example, the processor (510) may be connected to at least one of a memory (520), a communication device (530), an input interface device (540), an output interface device (550), and a storage device (560) through a dedicated interface.

[0144] The processor (510) can execute program commands stored in at least one of the memory (520) and the storage device (560). The processor (510) may refer to a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processor in which methods according to embodiments of the present invention are performed. Each of the memory (520) and the storage device (560) may be configured with at least one of a volatile storage medium and a non-volatile storage medium. For example, the memory (520) may be configured with at least one of a read-only memory (ROM) and a random access memory (RAM).

[0145] The methods according to the embodiments described in the claims or specification of the present disclosure may be implemented in the form of hardware, software, or a combination of hardware and software.

[0146] When implemented in software, a computer-readable storage medium storing one or more programs (software modules) may be provided. The one or more programs stored in the computer-readable storage medium are configured for execution by one or more processors within an electronic device. The one or more programs include instructions that cause the electronic device to execute methods according to the embodiments described in the claims or specification of the present disclosure.

[0147] These programs (software modules, software) may be stored in random access memory, non-volatile memory including flash memory, read only memory (ROM), electrically erasable programmable read only memory (EEPROM), magnetic disc storage devices, compact disc-ROM (CD-ROM), digital versatile discs (DVDs) or other forms of optical storage devices, magnetic cassettes, or may be stored in memories formed by a combination of some or all of these. In addition, each configuration memory may include multiple copies.

[0148] Additionally, the program may be stored on an attachable storage device that is accessible via a communication network, such as the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a storage area network (SAN), or a combination thereof. Such a storage device may be connected to a device performing an embodiment of the present disclosure via an external port. Additionally, a separate storage device on the communication network may be connected to a device performing an embodiment of the present disclosure.

[0149] In the specific embodiments of the present disclosure described above, components included in the disclosure are expressed in the singular or plural form, depending on the specific embodiment presented. However, the singular or plural expressions are selected to suit the presented situation for convenience of explanation, and the present disclosure is not limited to singular or plural components. Components expressed in the plural form may be composed of singular elements, or components expressed in the singular form may be composed of plural elements.

[0150] While the detailed description of this disclosure has described specific embodiments, it should be understood that various modifications are possible without departing from the scope of this disclosure. Therefore, the scope of this disclosure should not be limited to the described embodiments, but should be defined not only by the scope of the claims described below, but also by equivalents thereof.

Claims

1. As a distribution method for AI computations performed by an AI (artificial intelligence) data center (DC), The process of acquiring resource information required for the above AI computation, and A process for verifying the available AI computing capacity for each edge AI infrastructure deployed at the edge of a network including the above AI data center and each AI terminal connected to the edge AI infrastructure; A process of distributing the AI ​​operation to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal, based on the available AI operation capability confirmed for each of the edge AI infrastructure and the AI ​​terminal, the available AI operation capability in the AI ​​data center, and the necessary resource information. Distribution method for AI computation.

2. In Claim 1, The above required resource information includes, The prediction time required for the above AI computation is included, In the process of distributing the above AI computation, Distributing the AI ​​computation based on at least one of the time required to transmit an AI model for the AI ​​computation from the AI ​​data center to the edge AI infrastructure or the AI ​​terminal, whether information required for the AI ​​computation is transmitted to the edge AI infrastructure or the AI ​​terminal, and the time required for transmitting the required information. Distribution method for AI computation.

3. In Claim 1, In the process of distributing the above AI computation, According to the distribution of the AI ​​computation, the AI ​​computation is distributed by considering the time required for each of the AI ​​data center and the edge AI infrastructure to perform the AI ​​computation, wherein the AI ​​computation is distributed such that the time required for the AI ​​data center to perform the AI ​​computation is greater than that of the edge AI infrastructure by a preset time. Distribution method for AI computation.

4. In Claim 1, In the process of distributing the above AI computation, Distribute the AI ​​operation by considering the time required to distribute the AI ​​operation and the time for performing the AI ​​operation after distributing the AI ​​operation. Distribution methods for AI operations.

5. In claim 1, The distribution order in the process of distributing the above AI operations is: The above AI data center comes first, followed by the edge AI infrastructure and the AI ​​terminals sequentially. Distribution methods for AI operations.

6. In claim 5, In the above AI terminal, Some of the AI ​​operations distributed to the above edge AI infrastructure are distributed Distribution methods for AI operations.

7. In claim 5, In the process of redistributing to the above AI terminal, According to the redistribution of the AI ​​operation, the AI ​​operation is redistributed in consideration of the time taken by each of the edge AI infrastructure and the AI ​​terminal to perform the AI ​​operation, and the AI ​​operation is distributed so that the time taken by the edge AI infrastructure to perform the AI ​​operation is greater by a preset amount of time compared to the AI ​​terminal. Distribution methods for AI operations.

8. In claim 1, The above required resource information is: Provided to each of the edge AI infrastructure or AI terminals according to the distribution of the above AI operation, When the AI ​​service corresponding to the AI ​​operation is terminated at the AI ​​terminal, the necessary resource information provided to the edge AI infrastructure or each of the AI ​​terminals is deleted. Distribution methods for AI operations.

9. A distribution method for AI operations performed by edge AI (artificial intelligence) infrastructure. A process of transmitting available AI computing power for the edge AI infrastructure to an AI data center (DC) connected to the edge AI infrastructure; Including a process of distributing the AI ​​operation from the AI ​​data center, The above distributed AI operations are: The available AI computing power confirmed for each of the edge AI infrastructure and AI terminals connected to the edge AI infrastructure, the available AI computing power in the AI ​​data center, and the resource information required for the AI ​​computing are distributed to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal. Distribution methods for AI operations.

10. In the AI ​​(artificial intelligence) data center (DC), Transmitter and receiver, A control unit operably connected to the above transmitter and receiver, The above control unit, Obtain the resource information required for AI operations, Check the available AI computing capacity for each AI terminal connected to the edge AI infrastructure and the edge AI infrastructure deployed at the edge of the network including the above AI data center. In comparison with the available AI computing power confirmed for each of the edge AI infrastructure and the AI ​​terminal, the available AI computing power in the AI ​​data center, and the necessary resource information, the AI ​​computing is distributed to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal. AI data center.

11. In edge AI (artificial intelligence) infrastructure, Transmitter and receiver A control unit operably connected to the above transmitter and receiver, The above control unit, Transfer the available AI computing power for the edge AI infrastructure to an AI data center (DC) connected to the edge AI infrastructure, Distribute AI operations from the above AI data center, The above distributed AI operations are: The available AI computing power confirmed for each of the edge AI infrastructure and AI terminals connected to the edge AI infrastructure, the available AI computing power in the AI ​​data center, and the resource information required for the AI ​​computing are distributed to each of the AI ​​data center, the edge AI infrastructure, or the AI ​​terminal. Edge AI infrastructure.

Citation Information

Patent Citations

  • Technologies for distributing iterative computations in heterogeneous computing environments

    US20190220703A1

  • System for deep learning training using edge devices

    US20200272896A1

  • Methods and apparatus to load balance edge device workloads

    US20220327005A1

  • Distributed Machine-Learning Resource Sharing and Request Routing

    US20230020939A1

  • Distributed neural network communication system

    US20230222355A1