Neural network-based method and system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments

The neural network-based system optimizes AI workload placement in hybrid and multi-cloud environments by integrating cloud and network information, addressing inefficiencies in existing systems and reducing costs and time by recommending optimal configurations.

JP2026528667APending Publication Date: 2026-08-25メガゾーン クラウド コーポレーション
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025550930
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2024-08-28
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing systems struggle to optimally place AI workloads in hybrid and multi-cloud environments, leading to wasted resources and increased costs due to the difficulty in finding the optimal cloud location considering workload characteristics, network path information, and user requirements.

Method used

A neural network-based method and system that integrates cloud environment, network path, and user AI workload information to generate optimal execution plans, recommending the best cloud environment and network path for AI workloads, thereby optimizing time, cost, and resource usage.

Benefits of technology

This approach maximizes the efficiency of AI workload execution by selecting the optimal cloud environment and network path, reducing costs and optimizing time required to execute AI workloads, while providing stable operation in hybrid and multi-cloud environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026528667000001_ABST
    Figure 2026528667000001_ABST
Patent Text Reader

Abstract

The present invention relates to a neural network-based method and system for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment. The neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to the present invention may include the steps of: receiving user AI workload definition information and user optimization requirement details information from a user terminal; sampling information for different cloud environments and different network paths to generate a plurality of sample group data including the different cloud environments and different network paths; inputting each of the plurality of sample group data into a neural network and receiving a plurality of predicted values ​​for the plurality of sample group data from the neural network; and using an optimal prediction calculation to identify an optimal predicted value that satisfies the user AI workload definition information and the user optimization requirement details information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a neural network-based method and system for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments.

[0002] The present invention is implemented as part of the "World Class Plus Project Support" research project supported by the Ministry of Industry, Trade and Resources and managed by the Korea Institute of Industrial Technology. The project specific number is 1415189612, and the project number is P0024391. The name of this research project is "Development of Hybrid Cloud-based Intelligent MLOps Workload Management Tool Technology". Megazone Cloud Co., Ltd. participates as the implementing agency, and it is implemented from April 1, 2023 to December 31, 2026.

Background Art

[0003] Recently, the amount of computing resources required for artificial intelligence workloads has been increasing, and it has become important to efficiently secure cloud infrastructure.

[0004] Along with this, the cases of constructing artificial intelligence workloads in hybrid cloud and multi-cloud environments are increasing.

[0005] A hybrid cloud is a cloud environment in which multiple workloads or one workload are mixed in a public cloud and a private cloud. A multi-cloud is a cloud environment configured using cloud computing services of two or more public cloud service providers.

[0006] Hybrid cloud and multi-cloud environments offer advantages in terms of infrastructure scalability and the ability to configure a cloud environment using only the benefits of cloud servers provided by each cloud service provider.

[0007] However, failing to place artificial intelligence workloads in the optimal cloud location can lead to wasted resources and costs. The optimal location, in this context, means the location that minimizes idle resources, costs, and execution time.

[0008] However, because it is difficult for users to directly find and place resources in the optimal location, waste of cloud resources frequently occurs. Furthermore, in order to optimally place workloads, it is important to find and place them in a suitable cloud environment considering the characteristics, cost, and processing speed of the workload, but current systems do not provide the functionality to fulfill these requirements.

[0009] Furthermore, conventional methods have been limited in their ability to adequately reflect information that affects the actual workload execution cost, such as the characteristics of the user AI workload itself and network transmission path information. [Overview of the project] [Problems that the invention aims to solve]

[0010] The present invention provides a neural network-based method and system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments.

[0011] More specifically, the present invention provides a neural network-based method and system for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments, which can calculate an optimal execution plan for AI workloads in hybrid and multi-cloud environments by receiving not only cloud environment information but also user AI workload definition information and network path information as input.

[0012] Furthermore, the present invention provides a neural network-based method and system for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments, which can systematically manage heterogeneous cloud environments based on unified information values.

[0013] Furthermore, the present invention provides a neural network-based method and system for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments, which can recommend the optimal cloud environment that meets the user's requirements. [Means for solving the problem]

[0014] As discussed above, the neural network-based method and system for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments according to the present invention can achieve a more optimal execution point for AI workloads in hybrid and multi-cloud environments and maximize the efficiency of AI workload execution by integrally considering factors that affect the optimal execution of AI workloads (e.g., user AI workload definition information, cloud environment information, network path information, etc.) and user requirements.

[0015] Furthermore, the neural network-based method and system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention allows users to maximize the efficient use of time, cost, and resources required to execute AI workloads by selecting and providing users with the optimal cloud environment and network path that is suitable for the characteristics of the AI ​​workload and the user's requirements.

[0016] More specifically, the neural network-based method and system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention can recommend optimal cloud environment configuration information to users by comparing and analyzing the prices and performance of various cloud service providers. This allows users to reduce the cost burden and further optimize the time required to run AI workloads.

[0017] In other words, the neural network-based method and system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention provides users with convenience in building hybrid and multi-cloud environments, thereby enabling users to operate hybrid and multi-cloud environments stably.

[0018] Furthermore, the neural network-based method and system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention can provide recommendations for optimal cloud environment configuration information that satisfies user AI workload definition information and user optimization requirement details. This allows users to select the optimal cloud environment and network path that simultaneously optimizes the time and cost required to execute the AI ​​workload from various perspectives. [Brief explanation of the drawing]

[0019] [Figure 1] Figure 1 is a conceptual diagram illustrating a neural network-based system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention. [Figure 2] Figure 2 is a conceptual diagram illustrating a neural network-based system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention. [Figure 3] Figure 3 is a flowchart illustrating a neural network-based method for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments according to the present invention. [Figure 4A] Figure 4A is a conceptual diagram illustrating a neural network-based method for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments according to the present invention. [Figure 4B] Figure 4B is a conceptual diagram illustrating a neural network-based method for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention. [Figure 4C] Figure 4C is a conceptual diagram illustrating a neural network-based method for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention. [Figure 5] Figure 5 is a conceptual diagram illustrating a neural network-based method for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments according to the present invention. [Figure 6] Figure 6 is a conceptual diagram illustrating a neural network-based method for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments according to the present invention. [Figure 7]FIG. 7 is a conceptual diagram for explaining a neural network-based method for generating an optimal execution plan for AI workloads in a hybrid and multi-cloud environment according to the present invention. [Figure 8] FIG. 8 is a conceptual diagram for explaining a neural network-based method for generating an optimal execution plan for AI workloads in a hybrid and multi-cloud environment according to the present invention. [Figure 9A] FIG. 9A is a conceptual diagram for explaining a neural network-based method for generating an optimal execution plan for AI workloads in a hybrid and multi-cloud environment according to the present invention. [Figure 9B] FIG. 9B is a conceptual diagram for explaining a neural network-based method for generating an optimal execution plan for AI workloads in a hybrid and multi-cloud environment according to the present invention. [Figure 10] FIG. 10 is a conceptual diagram for explaining a neural network-based method for generating an optimal execution plan for AI workloads in a hybrid and multi-cloud environment according to the present invention. [Figure 11] FIG. 11 is a conceptual diagram for explaining a method for recommending cloud environment setting information to a user in the present invention. [Figure 12] FIG. 12 is a conceptual diagram for explaining a method for recommending cloud environment setting information to a user in the present invention. [Figure 13] FIG. 13 is a conceptual diagram for explaining a method for recommending cloud environment setting information to a user in the present invention.

BEST MODE FOR CARRYING OUT THE INVENTION

[0020] The present invention relates to a neural network-based method and system for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments, which accepts not only cloud environment information but also user AI workload definition information and network path information as input, calculates optimal execution data for AI workloads in hybrid and multi-cloud environments, and can systematically manage heterogeneous cloud environments based on a unified (or consistent) intermediate representation.

[0021] Cloud environment information may include various elements related to the cloud computing infrastructure provided by the cloud service provider. For example, cloud environment information includes cloud service providers (e.g., AWS, Azure, Google Cloud, etc.), cloud service location (or region), cloud service pricing (or cost) policies (e.g., cost information based on resources used, pricing elements, fee plans, discounts and benefits, etc.), cloud service types (e.g., IaaS (Infrastructure as a Service), PaaS (Platform as a Service), SaaS (Software as a Service), etc.), resource configuration (e.g., computing resources (virtual machines, containers, serverless computing, etc.), storage resources (block storage, file storage, object storage, etc.), network resources (virtual networks, load balancers, VPNs, CDNs, etc.), resource status (e.g., status and usage of virtual machines, storage, networks, databases, etc.), resource placement (e.g., location of virtual machines, containers, storage, etc.), security and regulations (e.g., firewalls, IAM (Identity and Access Management), MFA (Multi-Factor Authentication), encryption methods, security authentication, and regulatory compliance (GDPR, HIPAA, SOC), etc.). It may include at least one of the following: (2) operation and management (e.g., monitoring, automation, orchestration, etc.).

[0022] Furthermore, network routing information may include various elements related to network performance and network transmission paths necessary for sending and receiving data in a cloud environment. For example, network routing information may include at least one of the following: network location (or region), network topology (e.g., network diagram, subnet, routing table), routing information (e.g., routing protocol (BGP, OSPF), static routing, etc.), network devices (e.g., router, switch, firewall, etc.), IP address range (e.g., public IP, private IP, CIDR block, etc.), DNS settings (e.g., domain name, DNS server, record type, etc.), network security (e.g., firewall rules, security groups, ACL (Access Control List), etc.), traffic management (e.g., QoS (Quality of Service), traffic shaping, CDN, etc.), latency (e.g., network latency for each segment), bandwidth (e.g., maximum and average bandwidth usage for the network path), packet loss rate (e.g., data packet loss rate for each segment), and route optimization information (e.g., information necessary for route optimization such as network traffic load balancing, congestion prevention, and bypass routes).

[0023] Furthermore, user AI workload definition information may include various elements necessary to perform a specific AI task. For example, user AI workload definition information may include the type of workload (e.g., training, inference), the type of artificial intelligence (AI) model (e.g., CNN, RNN, Transformer), the AI ​​model architecture (e.g., the structure of the AI ​​model (number of layers, number of nodes per layer, number of parameters, etc.), the AI ​​algorithm, etc.), the characteristics of the dataset (e.g., the size (capacity) of the dataset, the format (CSV, image, text, etc.), the data source (data lake, database, API, etc.), the data pipeline (e.g., data preprocessing and postprocessing stages, data augmentation methods, etc.), and the execution environment (e.g., the necessary software). It may include at least one of the following: software (frameworks such as TensorFlow, PyTorch, and Scikit-learn), learning parameters (e.g., batch size, learning rate, number of epochs), computing resources (e.g., resource requirements such as required CPU, GPU, memory, and storage), performance targets for the artificial intelligence model (e.g., accuracy, precision, recall), inference requirements (e.g., real-time inference, batch inference, response time targets), distribution methods (e.g., strategies for distributing the model, such as distribution to a real-time prediction service or distribution to batch work), and monitoring and logging (e.g., model performance monitoring, error logging, and learning process recording).

[0024] However, in the present invention, the elements included in cloud environment information, network path information, and user AI workload definition information are not limited to these, and may further include various other elements besides those described above.

[0025] On the other hand, a heterogeneous cloud can refer to a cloud computing approach that integrates and uses different types of cloud environments and infrastructures. Such heterogeneous clouds can operate in environments where various forms of cloud infrastructure, such as public clouds, private clouds, and on-premises systems, coexist. In other words, heterogeneous clouds ensure interoperability between various platforms and provide the ability to move or integrate data and applications between different environments.

[0026] Furthermore, a hybrid cloud can refer to a cloud environment where multiple workloads, or a single workload, coexist in a mix of public and private clouds (or on-premises infrastructure). A hybrid cloud connects on-premises and public cloud infrastructure, storing sensitive data in the private cloud or on-premises, while leveraging the public cloud for general data processing and workloads requiring scaling. In other words, a hybrid cloud can simultaneously achieve flexible resource scalability, cost efficiency, and security.

[0027] Furthermore, multi-cloud can refer to a cloud environment composed of cloud computing services from two or more public cloud service providers. This means a method of using a combination of cloud services provided by various cloud service providers, rather than relying on a single cloud service provider. Each cloud service provider offers different functions and cost structures, and users can select and combine various cloud services according to their specific requirements.

[0028] On the other hand, neural network-based methods and systems for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments can be implemented in the form of various platforms, such as applications, software, and websites.

[0029] The neural network-based system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention will be described in more detail below with reference to the attached drawings.

[0030] Figures 1 and 2 are conceptual diagrams illustrating a neural network-based system for generating optimal execution plans for AI workloads in hybrid and multi-cloud environments according to the present invention (hereinafter referred to as the "AI Workload Optimal Execution Plan Generation System").

[0031] Referring to Figure 1, the AI ​​workload optimal execution plan generation system 100 according to the present invention may include at least one of the following components: an input unit 110, a display unit 120, a communication unit 130, a storage unit 140, and a control unit 150.

[0032] The input unit 110 can receive user input via the configuration of the input unit provided in the user terminal 10 (for example, a touch screen, virtual key, physical key (or hardware button), input sensor, microphone, etc.).

[0033] Specifically, the input unit 110 may be configured to receive (or select) the user's response to the user AI workload definition information and the user optimization request details information using the configuration of the input unit provided in the user terminal 10. Here, "received as input" may mean that when the user makes an input via the configuration of the input unit provided in the user terminal 10, the input unit 110 receives an input signal (or selection signal or user input) corresponding to such user input. For example, as shown in Figures 4a and 4b, the input unit 110 can receive user AI workload definition information 410 and user optimization request details information 420 that are input from the user via the user terminal 10.

[0034] Furthermore, the display unit 120 can output information via the configuration of the display unit provided in the user terminal 10 (e.g., output unit, touchscreen, speaker, etc.). In this case, the display unit 120 can perform both the role of outputting information and the role of receiving information. For example, as shown in Figures 4a and 4b, the display unit 120 can output a page (or screen) for receiving input from the user regarding user AI workload definition information 410 and user optimization request details information 420.

[0035] In this case, the input unit and the display unit of the user terminal 10 may exist independently of each other, or they may exist as a single unit, such as a touchscreen. When the input unit and the display unit exist as a single unit, such as a touchscreen, the input unit is a sensing unit that senses input via the display unit (for example, touch input or scroll input), and can be understood as a component of the display unit.

[0036] Hereinafter, without distinguishing whether the input and display sections of the user terminal 10 exist independently or as a single unit, a configuration that performs the function of receiving information will be referred to as the input section, and a configuration that performs the function of outputting information will be referred to as the display section.

[0037] The communication unit 130 can be connected wirelessly or via a wired network to the user terminal 10, a cloud service provider (or cloud service provision servers 21, 22, 23), an external server, and one or more networks, and can be configured to send and receive overall data and information necessary for the operation of the AI ​​workload optimization execution plan generation system 100.

[0038] Here, the user terminal 10 may include at least one of the following: mobile phone, smartphone, notebook computer, laptop computer, slate PC, tablet PC, ultrabook, desktop computer, digital broadcasting terminal, PDA (personal digital assistant), PMP (portable multimedia player), navigation system, and wearable device (e.g., smartwatch, smart glasses, HMD (head mounted display)).

[0039] In connection with this, the communication unit 130 can receive user responses (or response data) to the user AI workload definition information 410 and user optimization request details information 420 via the user terminal 10.

[0040] Furthermore, the communication unit 130 is connected to each of the multiple cloud service providers 21, 22, and 23 that provide different cloud environments (e.g., heterogeneous cloud environments), and is able to receive different cloud environment information related to the cloud computing infrastructure provided by each of the multiple cloud service providers 21, 22, and 23.

[0041] Furthermore, the communication unit 130 can support various communication methods depending on the communication standard of the device it communicates with.

[0042] For example, the communication unit 130 supports WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), and Bluetooth. TM It can be configured to communicate with a target using at least one of the following technologies: Frequency Identification (FRF), Infrared Data Association (IrDA), Ultra-Wideband (UWB), ZigBee, Near Field Communication (NFC), Wi-Fi Direct, or Wireless USB (Wireless Universal Serial Bus).

[0043] On the other hand, the storage unit 140 is also called a "database (DB)" or "memory," and can be configured to store various types of information according to the present invention. In the present invention, the storage unit 140 may be provided in the AI ​​workload optimal execution plan generation system 100 itself. Furthermore, at least a portion of the storage unit 140 may be configured as a cloud server (or cloud storage). In other words, the storage unit 140 only needs to be a space where information necessary for the operation of the AI ​​workload optimal execution plan generation system 100 according to the present invention is stored, and it can be understood that there are no physical space constraints.

[0044] The user described above may have an account pre-registered with the AI ​​Workload Optimal Execution Plan Generation System 100 according to the present invention. In this case, the account can be generated via a page (or screen) linked to the AI ​​Workload Optimal Execution Plan Generation System 100. Alternatively, the account can also be generated in at least one other system linked to the AI ​​Workload Optimal Execution Plan Generation System 100. However, in this specification, without distinguishing the system 100 to which the user account was issued, all accounts that can utilize the various services provided by the AI ​​Workload Optimal Execution Plan Generation System 100 according to the present invention are referred to as "accounts pre-registered with the AI ​​Workload Optimal Execution Plan Generation System."

[0045] As a result, the storage unit 140 may store various information related to the user account. Here, the information related to the user account may include the user's history information.

[0046] More specifically, user history information may include information about various events that occurred in the user account. For example, events that occur in the user account may include at least one of the following: i) input of AI workload definition information, ii) input of optimization requirement details information, and iii) setting weights for the optimization requirement details information, which are necessary to execute a particular AI workload.

[0047] As a result, the user's history information may include, for example, at least one of the following: i) input records (or history or breakdown) of AI workload definition information entered by the user; ii) input records of optimization request details information entered by the user; iii) records of weight settings for the user's optimization request details information; iv) execution records of the user's AI workload; v) breakdown of cloud environment configuration information recommended (or suggested) by the user; vi) the user's cloud environment configuration information (or cloud environment information used by the user).

[0048] The users described above may include specific companies (or entities or businesses), and users who do not have an account under the present invention (non-members) can also use the various services provided under the present invention.

[0049] Furthermore, the storage unit 140 may store data and instructions necessary for the operation of the AI ​​workload optimal execution plan generation system 100 according to the present invention. For example, the storage unit 140 may store training datasets necessary for training an artificial neural network (or artificial intelligence model 152), and may also store data, software, firmware, program code (source code), and instructions that have been processed or are scheduled to be processed by the control unit 150.

[0050] Furthermore, the storage unit 140 may store different cloud environment information related to the cloud computing infrastructure provided by each of the multiple cloud service providers 21, 22, and 23. As another example, the storage unit 140 may store different network path information related to the network performance and network transmission paths required to send and receive data in different cloud environments.

[0051] On the other hand, the control unit 150, also called a "processor," can play a role in controlling the overall operation of the AI ​​workload optimal execution plan generation system 100 according to the present invention. The control unit 150 can process signals, data, information, etc. that are input or output via the above-described components, or can perform a series of data processing to provide or process appropriate information and functions for the user.

[0052] As shown in Figure 2, the control unit 150 can generate (or calculate) input data used to recommend (or suggest) a cloud service that meets the user's requirements.

[0053] First, the control unit 150 can receive input from the user terminal 10 for user AI workload definition information 210 and user optimization request detail definition information 240.

[0054] Here, the user AI workload definition information 210 may include elements that have different characteristics (or meanings) in relation to the type (or category) of the AI ​​workload, the type of artificial intelligence model, and the data characteristics.

[0055] Furthermore, the user optimization requirement details definition information 240 may include elements that have different characteristics in relation to the time, price, and resource utilization required (or used) to execute the user AI workload definition information. For example, the user optimization requirement details definition information 240 may include at least one of the following: response time, latency, mean time to detection, mean time to resolution, mean time between failures, price, utilization, compliance, and scalability. However, the elements included in the user optimization requirement details information in the present invention are not limited to these, and may include various other elements besides those described above.

[0056] Furthermore, the control unit 150 can extract necessary information from the user AI workload definition information 210 (for example, workload type, dataset characteristics, artificial intelligence model type, hyperparameters, etc.).

[0057] Next, the control unit 150 can sample the cloud environment information 220 and network path information 230 stored in the storage unit 140 and collect N data pairs (or data sets or sample group data) in which the cloud environment information 220 and network path information 230 form a pair.

[0058] Furthermore, the control unit 150 can generate N sample group data corresponding to the N data pairs, each containing the extracted user AI workload definition information, by combining each of the N data pairs, each containing the cloud environment information 220 and the network path information 230, with the information extracted from the user AI workload definition information. More specifically, for one user AI workload definition information 210, there may be N data pairs, each containing the cloud environment information 220 and the network path information 230, and N sample group data (input data) may be generated by combining these.

[0059] The control unit 150 can perform validation and preprocessing processes on N sample group data and calculate the preprocessed input data (N sample group data).

[0060] On the other hand, the control unit 150 can convert different environmental information (e.g., parameter names, resource types and combinations, etc.) collected on-premises and in heterogeneous clouds into a unified expression. More specifically, the control unit 150 can use the intermediate expression data converter 151 to convert N pre-processed sample group data into N intermediate expressions (or intermediate expressions).

[0061] Furthermore, the control unit 150 can remove noise and perform normalization on the N converted intermediate expressions. This can be understood as a pre-processing step that prevents overfitting, which degrades the generalization performance of the artificial neural network 152, and adjusts all attributes (or features) of the input data (N intermediate expressions) to the same scale.

[0062] Next, the control unit 150 can input the N intermediate expressions, which have undergone noise reduction and normalization, into the artificial neural network 152. In this case, the artificial neural network 152 may be executed for each of the N intermediate expressions in a different environment, and N calculations, corresponding to the number of N expressions, may be performed in parallel. As a result, the artificial neural network 152 can output N predicted values ​​for user optimization requirement detail information elements (such as time and price) that are expected to be necessary to execute the user AI workload definition information 210 in each of the different environments (e.g., a cloud environment, a network environment).

[0063] Next, the control unit 150 can identify the optimal prediction value using the optimal prediction computer 153 (or optimal prediction calculation). The optimal prediction computer 153 receives N prediction values, which are output data from the artificial neural network 152, as input, calculates a final score for the N optimal prediction values ​​based on the input N prediction values ​​and the user optimization requirement detail definition information 240, and sorts the N optimal prediction values ​​based on the final score. The user can select any one of the N optimal prediction values. Alternatively, the AI ​​workload optimal execution plan generation system 100 itself may identify the optimal prediction value with the highest score (Top-1) and provide it to the user without requiring the user to make a separate selection.

[0064] Once the identification of the optimal prediction values ​​is complete, the control unit 150 can use the optimal execution data generator 154 to identify the intermediate expression corresponding to the identified optimal prediction values, identify the cloud environment information and network path information corresponding to the identified intermediate expression, and then generate the final optimal execution data using the identified cloud environment and network path information.

[0065] Furthermore, the control unit 150 can identify at least one cloud environment configuration information that satisfies the user's requirements (user AI workload definition information and user optimization requirement details information) based on the optimal predicted values, and recommend the identified cloud environment configuration information to the user. For example, as shown in Figure 2, the control unit 150 can provide recommendation information to the user terminal 10 that includes cloud environment configuration information that satisfies the user's requirements (for example, "The price and response time are good!", "By choosing Goo*Cloud, you can save $100 in the cost required to run your AI workload and expect a response time reduction of about 50ms." 200).

[0066] However, the intermediate representation data converter 151, the artificial neural network 152, the optimal prediction computer 153, and the optimal execution data generator 154 are all components of the control unit 150, and for the sake of clarity, they may be referred to collectively as the control unit 150 below.

[0067] The following describes in more detail the neural network-based method for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments according to the present invention, based on the configuration of the AI ​​workload optimal execution plan generation system 100 as described above. Figure 3 is a flowchart illustrating the neural network-based method for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments, Figures 4a, 4b, 4c, 5, 6, 7, 8, 9a, 9b, and 10 are conceptual diagrams illustrating the neural network-based method for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments according to the present invention, and Figures 11, 12, and 13 are conceptual diagrams illustrating the method by which the present invention recommends cloud environment setting information to the user.

[0068] In the present invention, a process may be performed to receive user AI workload definition information and user optimization requirement details information from the user terminal (S310, see Figure 3).

[0069] The control unit 150 can provide a user environment that allows the user to input AI workload definition information and optimization request details. For example, as shown in Figures 4a and 4b, the control unit 150 can provide a page (or screen) on the user terminal 10 that is configured to allow the user to input user AI workload definition information 410 and user optimization request details 420.

[0070] As discussed above, the user AI workload definition information 410 may include elements that have different meanings in relation to the type of AI workload, the type of artificial intelligence model, and the characteristics of the dataset. For example, as shown in Figure 4a, the user AI workload definition information 410 may include at least one of the following: workload type 411 (e.g., "work type", "detailed work", "use case"), artificial intelligence model type 412, artificial intelligence model architecture 413 (e.g., "number of layers", "number of nodes per layer", "number of parameters"), and dataset information 414 (e.g., "data size", "data format").

[0071] The control unit 150 can receive user AI workload definition information 410 from the user terminal 10. For example, based on the selection of a graphic object (e.g., "Confirmation" 410a) linked to the function of receiving user AI workload definition information 410, the control unit 150 can receive user input for elements with different characteristics included in the user AI workload definition information 410 (e.g., "Work type: Deep learning model training", "Detailed work: Sentiment Analysis", "Use case: Customer Review Analysis", "Model type: Recurrent Neural Networks (RNNs)", "Number of layers: 5", "Number of nodes per layer: [128, 64, 32, 16, 8]", "Number of parameters: 1.2 million parameters", "Data size: 90GB", "Data format: CSV (Comma-Separated Values)" 411, 412, 413, 414).

[0072] Furthermore, as discussed above, the user optimization requirement details information 420 may include elements that have different characteristics from each other in relation to the time required (or usage) to execute the user AI workload definition information 410, price, and resource utilization. For example, as shown in Figure 4b, the user optimization requirement details information 420 may include at least one of the following: response time 421, latency 422, mean time to detection 423, mean time to resolution 424, mean time between failures 425, price 426, utilization 427, compliance 428, and scalability 429.

[0073] Furthermore, the control unit 150 can receive user optimization request details information 420 from the user terminal 10. For example, the control unit 150 can receive user input from the user terminal 10 regarding elements 421, 422, 423, 424, 425, 426, 427, 428, 429 (for example, "Response Time: 200ms or less", "Latency: 50ms or less", "Mean Time to Detection: 2 minutes or less", "Mean Time to Resolution: within 30 minutes", "Mean Time Between Failures: 1000 hours or more", "Price: $500 or less per month", "Utilization: 80% or less", "Compliance: GDPR compliant", "Scalability: Automatic expansion when traffic increases") which are contained in the user optimization request details information 420 and have different characteristics from each other.

[0074] On the other hand, the control unit 150 may set the weights of elements 421, 422, 423, 424, 425, 426, 427, 428, and 429, which have different characteristics and are included in the user optimization request details information 420, from the user terminal 10.

[0075] To this end, the control unit 150 can provide a user environment in which the user can set weights for elements 421, 422, 423, 424, 425, 426, 427, 428, and 429, which have different characteristics from each other. For example, as shown in Figure 4c, the control unit 150 may provide a graphic object (e.g., a slider) that is linked to a function that allows the user to set weights for each of the different elements 421, 422, 423, 424, 425, 426, 427, 428, and 429. The user can increase or decrease the weight for each element by adjusting the slider to the left or right. In this case, the sum of the weights may always be fixed at a preset value (e.g., "1").

[0076] However, the method for setting user weights in the present invention is not limited to this, and a user environment can be provided in which users can set weights using various other methods (for example, checkboxes, text boxes for direct numerical input, voice, dropdown menus, etc.).

[0077] Furthermore, the control unit 150 can receive weight setting information (or user input regarding weight settings) from the user terminal 10 for elements 421, 422, 423, 424, 425, 426, 427, 428, and 429 having different characteristics from each other. For example, the control unit 150 can receive user weight setting information, including "1. Response time: 0.3" and "6. Price: 0.7", from the user terminal 10 based on the selection of a graphic object (e.g., "Confirmation" 430) in conjunction with the reception of weights.

[0078] In contrast, the weights of elements 421, 422, 423, 424, 425, 426, 427, 428, and 429, which have different characteristics from each other, may be set by the AI ​​workload optimization execution plan generation system 100 itself. For example, the control unit 150 can set the weights of elements 421, 422, 423, 424, 425, 426, 427, 428, and 429, which have different characteristics from each other, based on information (e.g., history information) of the user account logged into the user terminal 10. More specific details will be described later.

[0079] On the other hand, the present invention may involve sampling information from different cloud environments and different network paths to generate a plurality of sample group data that include different cloud environments and different network paths (see S320, Figure 3).

[0080] As discussed above, information on different cloud environments (e.g., heterogeneous clouds) may include elements that have different characteristics in relation to the cloud computing infrastructure provided by different cloud service providers. For example, as shown in Figures 5(a) and (b), elements with different characteristics included in information on different cloud environments may include at least one of the following: cloud service providers, regions and available space, computing resources, storage resources, network resources, security and compliance, cost management, resource management and autoscaling, service and application management, and data movement and integration between clouds. In this case, the first cloud environment information 510 of the different cloud environment information 510, 520 may include elements relating to the cloud computing infrastructure provided by the first cloud service provider (e.g., "Amazon Web Services (AWS)"), and the second cloud environment information 520 may include elements relating to the cloud computing infrastructure provided by the second cloud service provider (e.g., "Google Cloud (GCP)").

[0081] Furthermore, as discussed above, network routing information may include elements that have different characteristics in relation to the network performance and network transmission paths required to send and receive data in a cloud environment (or different cloud environments or different network environments). For example, as shown in Figure 6, the different network routing information 610 and 620 may include at least one of the following: network location (or region), network bandwidth, latency, packet loss rate, jitter, availability, reliability, congestion state, path length, security level, cost, and ISP information.

[0082] However, the elements included in the different cloud environment information and different network path information in the present invention are not limited to those described above, and may include various other elements in addition to those described above.

[0083] On the other hand, the control unit 150 can sample different cloud environment information 510, 520 and different network path information 610, 620 that are pre-stored in the AI ​​workload optimization execution plan generation system 100, and collect (or generate) N (or more) sample group data where the different cloud environment information 510, 520 and different network path information 610, 620 form pairs. As another example, the control unit 150 can sample cloud environment information 510, 520 and different network path information 610, 620 that are received (or collected) from a server (e.g., a cloud service provider) linked to the AI ​​workload optimization execution plan generation system 100, and generate N sample group data where the different cloud environment information 510, 520 and different network path information 610, 620 form pairs.

[0084] The control unit 150 can generate multiple sample group data by combining each of the sampled, different cloud environment information 510, 520 and different network path information 610, 620 with the user AI workload definition information 410, based on the user AI workload definition information 410.

[0085] More specifically, if there are N sample group data sets in which each set consists of pairs of different cloud environment information sets 510, 520 and different network path information sets 610, 620, the control unit 150 can generate N sample group data sets containing user AI workload definition information 410, different cloud environment information sets 510, 520 and different network path information sets 610, 620 by duplicating one user AI workload definition information set 410 to correspond to the N sets. For example, as shown in Figure 7, the control unit 150 can generate multiple sample group data sets 701, 702, and 703 containing user AI workload definition information 410, different cloud environment information sets 510, 520 and different network path information sets 610, 620.

[0086] Here, "c" can mean user AI workload definition information, "s" can mean cloud environment information, and "t" can mean network path information. This can be understood as, if there is one user AI workload definition information (c) and N 2-tuples (pairs of cloud environment information (s) and network path information (t)) that match it, then the c is duplicated N times to generate a 3-tuple (user AI workload definition information (c), cloud environment information (s), network path information (t)).

[0087] In other words, the control unit 150 can generate multiple sample group data 701, 702, and 703 that include not only different cloud environment information and different network path information, but also user AI workload definition information, by combining each of the multiple (N) sample group data, each of which consists of collected different cloud environment information 510, 520 and different network path information 610, 620, with user AI workload definition information 410.

[0088] Furthermore, the control unit 150 can perform tests on the generated sample group data 701, 702, and 703. For example, testing on multiple sample group data can be understood as testing whether or not NULL values ​​exist in the multiple sample group data.

[0089] However, the verification process in the present invention may be performed during the process of receiving user AI workload definition information and user optimization requirement details information. As an example, the control unit 150 can verify whether an appropriate user response corresponding to each element included in the user AI workload definition information and user optimization requirement details information has been input (for example, whether information for the model has been input for the model type, and whether a data format suitable for the workload type and model type has been input for the data format).

[0090] On the other hand, in the present invention, a process may be performed in which each of the multiple sample group data is input into a neural network, and multiple predicted values ​​for the multiple sample group data are received from the neural network (S330, see Figure 3).

[0091] The control unit 150 can input each of the multiple sample group data 701, 702, and 703, which have undergone preprocessing (e.g., hypothesis testing), into the artificial neural network 152.

[0092] At this time, the control unit 150 can use the intermediate representation data converter 151 to convert the pre-processed sample group data 701, 702, and 703 into multiple intermediate representation data.

[0093] In this case, the number of intermediate representation data to be converted may be as many as the number of sample group data (N) corresponding to the number of sample group data 701, 702, 703. For example, if we assume that there are "10" sample group data, the number of intermediate representation data to be converted could be 10.

[0094] The control unit 150 can convert each of the multiple sample group data 701, 702, and 703 into multiple intermediate representation data based on a preset format. For example, as shown in Figure 7, the control unit 150 converts the first sample group data 701 (e.g., "(c, s_1, t_1)") into the first intermediate representation data 711 (e.g.,

number

number

number

[0095] In other words, as discussed above, the present invention ensures generality that can adequately represent information from heterogeneous systems (hybrid and multi-cloud) through the intermediate representation data conversion process, and ensures data representation efficiency that promotes effective learning and inference of neural network models.

[0096] Furthermore, the control unit 150 can remove noise and perform normalization on the converted intermediate representation data 711, 712, and 713. This can be understood as a preprocessing step that prevents overfitting, which degrades the generalization performance of the artificial neural network 152, and adjusts all attributes (or features) of the input data (e.g., multiple intermediate representation data) to the same scale.

[0097] As an example, as shown in Figure 8, all input data for the artificial neural network 152 must be defined as real values. Therefore, the control unit 150 can perform a type conversion process to convert any data with Boolean values ​​to real values. However, if no input data with Boolean values ​​exists, the type conversion process does not need to be performed.

[0098] Furthermore, the control unit 150 can perform normalization and outliers removal (robustness) processes on multiple intermediate representation data 711, 712, and 713. For example, the range and variability of values ​​may differ depending on the information in the input data. If the latency (LA) value range is "[0.0, 1.0]", the price (PR) value range may be "[0.0, 20,000,000]". Also, outliers present in the input data can affect the performance of the artificial neural network 152.

[0099] The techniques (or methods) used in this invention for normalization and outlier sensitivity removal can be confirmed by referring to the table and mathematical formulas regarding the normalization method shown in the figure below. [Table 1]

[0100] On the other hand, the control unit 150 can input each of the pre-processed intermediate representation data into the artificial neural network 152. For example, as shown in Figure 8, the control unit 150 can input each of the pre-processed first intermediate representation data 801, second intermediate representation data 802, and Nth intermediate representation data 803 into the artificial neural network 152.

[0101] At this point, the architecture of the artificial neural network 152 in the present invention can be confirmed by referring to the table and mathematical formulas shown in the figure below. [Table 2]

[0102] In this regard, the artificial neural network 152 can make predictions for each of the multiple intermediate representation data 801, 802, and 803, and output multiple predicted values ​​for each of the multiple intermediate representation data. In this case, the artificial neural network 152 can make predictions for each of the multiple intermediate representation data 801, 802, and 803 in parallel, and output multiple predicted values ​​for each of the multiple intermediate representation data 801, 802, and 803 simultaneously. For example, the artificial neural network 152 can make predictions for each of the first intermediate representation data 801, the second intermediate representation data 802, and the Nth intermediate representation data 803 in parallel, and output the first predicted value 811 for the first intermediate representation data 801 (for example,

number

number

number

[0103] Here,

number

number

number

number

[0104] This is due to the weights (for example,) in the optimal prediction calculation process described later.

number

number

[0105] Furthermore, the multiple prediction values ​​811, 812, and 813 output from the artificial neural network 152 may include the time and cost required to execute the user AI workload definition information 410, with different cloud environment information 510, 520 and different network path information 610, 620. However, the information included in the multiple prediction values ​​811, 812, and 813 is not limited to these and may further include resource utilization rates used to execute the user AI workload definition information.

[0106] In other words, since there are a total of N intermediate representation data 801, 802, and 803 for the user AI workload definition information 410, the control unit 150 can apply the neural network a total of N times to one user AI workload and receive (or acquire) N (or more) predicted values, including time and price, through the results.

[0107] On the other hand, in the present invention, a process may be performed to identify the optimal predicted value that satisfies the user AI workload definition information and the user optimization requirement details information using optimal prediction calculation (see S340, Figure 3).

[0108] The control unit 150 receives as input multiple predicted values ​​811, 812, and 813 output by the artificial neural network 152 for the user AI workload definition information 410, and user optimization requirement details information 420, and can calculate (or identify) at least one predicted value that satisfies the user's detailed requirements (e.g., AI workload definition information and optimization requirement details information).

[0109] As discussed above, the user optimization requirement details information 420 may include elements having different characteristics from each other, and the control unit 150 may set weights for the elements 421, 422, 423, 424, 425, 426, 427, 428, and 429 having different characteristics from each other, based on the user terminal 10.

[0110] Referring to Figure 9a,

number

number

number

number

[0111] The control unit 150 can use an optimal prediction calculation (or optimal prediction computer) to identify the optimal prediction value that satisfies the user AI workload definition information 410 and the user optimization requirement details information 420.

[0112] Specifically, as shown in 9b, the control unit 150 takes a plurality of predicted values ​​901, 902, and 903 output from the artificial neural network 152 and weights 911 (for example,) for elements having different characteristics received from the user terminal 10.

number

[0113] First, the optimal prediction computer 153 can define a score function based on multiple prediction values ​​901, 902, and 903 and weights 911 for elements having different characteristics from each other. In this case, for the sake of explanation, the present invention will assume and explain only a few elements (for example, 7) (and therefore 7 weights) from the elements described as an example of user optimization requirement details.

number

[0114] [Mathematical formula 1]

number

[0115] Next, the optimal prediction computer 153 can calculate a score function and sort multiple optimal prediction values ​​according to the calculated score. For example, multiple optimal prediction (or inference) values ​​may be calculated by calculating the score function, and the multiple optimal prediction values ​​may be sorted in descending order.

[0116] Furthermore, the control unit 150 selects an optimal prediction value 920 from among the sorted plurality of optimal prediction values ​​that satisfies the user AI workload definition information 410 and the user optimization requirement details information 420 (for example,

number

[0117] In this case, the identified optimal prediction value 920 may correspond to either the optimal prediction value identified by the AI ​​workload optimal execution plan generation system 100, or the optimal prediction value identified based on the user's selection.

[0118] In this regard, the user's priority is to identify the optimal predicted value (i.e., (r_{k,1}, ..., r_{k,9}) for a specific k that calculates the maximum score in the optimal predicted value identification step described above, and then finally identify the corresponding input data in the optimal execution data generation step. Therefore, the control unit 150 can identify the optimal predicted value (Top-1) that has the maximum score (Top-1 auto-return).

[0119] However, the user has multiple optimal prediction values ​​(

number

[0120] More specific details regarding "Top-1 Auto Return" and "User Selection" can be found by referring to the table and mathematical formulas shown in the figure below. However, this invention does not limit the method for identifying the optimal predicted value to just one of these methods. [Table 3]

[0121] In this way, the control unit 150 can use the multiple predicted values ​​901, 902, 903 and the weight 911 of the user optimization requirement details information to calculate the optimal predicted value 920 that satisfies the user AI workload definition information 410 and the user optimization requirement details information 420.

[0122] The optimal prediction calculation described above may be designed to reflect user optimization requirement details, have a "weight" mechanism, and at the same time allow the user to review the inference results and make a manual selection regardless of the weights of the details.

[0123] On the other hand, in the present invention, optimal execution data can be generated based on the optimal predicted value 920.

[0124] More specifically, in this invention, optimal execution data generation can be used to identify (specify) the cloud environment configuration information for which the optimal predicted value 920 was calculated, and then optimal execution data can be generated using the identified cloud environment configuration information.

[0125] Cloud environment configuration information includes cloud environment and network routing information (information about cloud environment settings) necessary to run a specific AI workload (e.g., a user AI workload), and may include values ​​set for optimal execution of that workload.

[0126] As shown in Figure 10, first, the control unit 150 can confirm the identified optimal prediction value 1001 using optimal prediction calculation. This means obtaining the expected optimal time and price information required to execute the user's AI workload (user AI workload definition information). Starting from this, the control unit can then identify the intermediate representation data used to calculate the optimal prediction value 1001, and then identify specific cloud environment information and specific network path information included in the sample group data corresponding to the identified intermediate representation data.

[0127] Next, the control unit 150 can identify (or specify) the intermediate representation data from which the optimal prediction value 1001 was calculated, based on the calculation result of the optimal prediction value 1001. For example, the control unit 150 can identify the intermediate representation data 1010 for the optimal execution data based on the optimal prediction value 1001. However, as an example, in the case of intermediate representation data, it is not necessarily the case that only one but multiple intermediate representation data 1010, 1020 may be identified.

[0128] Furthermore, the control unit 150 can convert the identified intermediate representation data 1010 into final optimal execution data. More specifically, the control unit 150 can combine the cloud environment information and network path information 1010a contained in the intermediate representation data 1010 to generate optimal execution data (or execution instructions). As an example, the optimal execution data may be generated in the form of program code (or source code). However, the form of the optimal execution data in this invention is not limited to this, and can be realized in various other forms.

[0129] On the other hand, the present invention can provide a user environment that recommends an optimal cloud environment that satisfies the user's requirements (e.g., user AI workload, user optimization requirement details, etc.) based on the process of the AI ​​workload optimization execution plan generation system 100 described above.

[0130] To this end, the control unit 150 can first identify at least one cloud environment configuration information that satisfies the user AI workload definition information 410 and the user optimization requirement details information 420 based on the optimal prediction value 1001. For example, by calculating the optimal prediction value 1001, the control unit 150 can analyze various cloud environments and network paths related to the user AI workload definition information 410 to calculate the expected time and expected price, and then identify cloud environment configuration information that satisfies the user AI workload definition information 410 and the user optimization requirement details information 420 based on the calculation results.

[0131] For example, cloud environment configuration information may include at least one of the following: cloud service provider, instance type and configuration (e.g., CPU, GPU, and memory specifications), storage options (e.g., SSD, HDD), network settings (e.g., network bandwidth and latency), operating system and software environment (e.g., Windows, Linux®, Python, TensorFlow, etc.), cost management information (e.g., estimated cost, spot instances, reserved instances, etc.), and performance information (e.g., estimated execution time, estimated resource utilization). However, the elements included in cloud environment configuration information are not limited to these, and may include other elements with different characteristics.

[0132] Next, the control unit 150 can generate recommendation information for the identified cloud environment configuration information.

[0133] The "recommended information" described in this invention may include the estimated time and estimated price required to execute a user-defined AI workload (user AI workload definition information) based on identified cloud environment configuration information. In this case, the recommended information may further include information on the estimated resource utilization used to execute the user-defined AI workload, in addition to the estimated time and estimated price.

[0134] Furthermore, the control unit 150 can provide the generated recommendation information to the user terminal 10. For example, as shown in Figure 11, the control unit 150 can provide recommendation information 1101 regarding cloud environment settings (for example, "The price and response time are good!", "By choosing Goo*Cloud, you can save $100 in the cost required to run your AI workload and expect a response time reduction of about 50ms.") to the user terminal (or service page 10) where user account U is logged in.

[0135] On the other hand, in the present invention, the recommendation information provided to the user terminal 10 may include a plurality of recommendation pieces having different types (or characteristics) from each other.

[0136] Specifically, the multiple recommendation pieces having different types may include a first type of recommendation piece identified based on user optimization requirement details information 420, and a second type of recommendation piece identified based on conditions pre-set in the AI ​​workload optimization execution plan generation system 100. However, the types in this invention are not necessarily limited to the first and second types.

[0137] Here, the first type of recommendation information may be recommendation information provided based on weights set for elements 421, 422, 423, 424, 425, 426, 427, 428, and 429 having different characteristics, which are included in the user optimization requirement details information 420. As an example, suppose the user terminal 10 sets weights for "Response Time" and "Price". The control unit 150 can provide recommendation information 1101 that satisfies the user AI workload definition information 410 and the weights for response time and price.

[0138] Furthermore, the second type of recommendation information may also be recommendation information identified by the AI ​​workload optimization execution plan generation system 100 itself and provided to the user (or user account U).

[0139] In this regard, the AI ​​workload optimization execution plan generation system 100 may have pre-configured conditions for providing a second type of recommendation information. For example, the AI ​​workload optimization execution plan generation system 100 may have pre-configured conditions based on i) information matching user account U (e.g., user history information), ii) specific cloud environment configuration information that was most frequently selected during a specific period (or a pre-configured period), and iii) at least one of the cloud environment configuration information of multiple users registered with the AI ​​workload optimization execution plan generation system 100. However, the criteria for setting the pre-configured conditions are not limited to these and may include various other criteria besides those mentioned above.

[0140] The control unit 150 may provide the user terminal 10, which has been provided with the first type of recommendation information 1101, with second type of recommendation information identified based on pre-set conditions. For example, as shown in Figures 11 and 12, the control unit 150 can provide the user terminal 10 with second type of recommendation information 1201, 1202, and 1203 based on the selection of a graphic object (e.g., "See more other recommendations" 1110) linked to the second type of recommendation information provision function from the user terminal 10.

[0141] For example, among the second type of recommendations 1201, 1202, and 1203, the first recommendation 1201 (e.g., "It's good in terms of price!", "Choosing Ama*Cloud will save you $150 in the cost required to run your AI workload, but you can expect response times to be about 100ms longer.") may be a recommendation provided based on the historical information of user account U. This may be provided based on the weights of elements included in the user optimization request details that the user has previously set or preferred elements.

[0142] As another example, among the second type of recommendations 1201, 1202, and 1203, recommendation 1202 (e.g., "Currently the HOT!", "Choosing N* Cloud will increase the cost required to run your AI workload by $50, but is expected to reduce response times by approximately 100ms.") may be provided based on the specific cloud environment configuration information that was most frequently selected over a particular period.

[0143] As another example, among the second type of recommendation information 1201, 1202, and 1203, the third recommendation information 1203 (for example, "Pick users similar to the first user!", "Choosing AW* Cloud can save you $100 in the costs required to run your AI workload, but response times are expected to be about 200ms longer.") may be provided based on the cloud environment configuration information of multiple users registered in the AI ​​workload optimization execution plan generation system 100.

[0144] Furthermore, the control unit 150 can sort the identified multiple recommendation pieces of information in order and provide them to the user terminal 10.

[0145] Here, "sorting and providing in order" can be understood as sorting and providing multiple identified recommendations according to their priority based on the user optimization requirements details.

[0146] For example, as shown in Figures 12 and 13, the control unit 150 can sort multiple recommendation information 1301, 1302, and 1303 according to priority and provide them to the user terminal 10 based on the selection of a graphic object 1210 (e.g., "Compare all suggestions") linked to multiple recommendation information sorting functions from the user terminal 10.

[0147] On the other hand, the control unit 150 may select at least one of several recommended pieces of information from the user terminal 10.

[0148] Specifically, the control unit 150 can receive the user's selection from the user terminal 10 for either the first type of recommendation information or the second type of recommendation information. For example, as shown in Figure 13, the control unit 150 can receive the user's selection from the user terminal 10 for the first recommendation information 1301 from among a plurality of recommendation information 1301, 1302, and 1303.

[0149] Furthermore, the control unit 150 can generate optimal execution data corresponding to the recommendation information recommended by the user. More specifically, the control unit 150 can generate optimal execution data corresponding to the recommendation information selected from the user terminal 10 and register the generated optimal execution data to the user account U. For example, as shown in Figure 13, suppose the first recommendation information 1301 corresponding to the first type of recommendation information is selected from the user terminal 10. The control unit 150 can generate optimal execution data corresponding to the first recommendation information 1301 and register it to the user account U logged into the user terminal 10.

[0150] Thus, the present invention provides a user environment that allows users to select the optimal cloud environment from various perspectives by simultaneously providing a first type of recommendation information and a second type of recommendation information having different characteristics from each other.

[0151] In other words, users can receive not only customized recommendations that meet their specific requirements, but also recommendations from various perspectives, allowing them to select the optimal cloud environment (cloud environment configuration information) that simultaneously meets the time, cost, and performance optimization requirements necessary for running their AI workloads.

Claims

1. The steps include receiving user AI workload definition information and user optimization request details information from the user terminal, A step of sampling information from different cloud environments and different network paths to generate multiple sample group data containing the different cloud environments and different network paths, The steps include inputting each of the multiple sample group data into a neural network and receiving multiple predicted values ​​for the multiple sample group data from the neural network, A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment, comprising the step of using an optimal prediction calculation to identify an optimal prediction value that satisfies the user AI workload definition information and the user optimization requirement details information.

2. The user AI workload definition information is as follows: This includes information on AI workload types, artificial intelligence model types, and dataset characteristics. The aforementioned information on different cloud environments is, This includes information on cloud service providers, cloud service locations, cloud service pricing policies, and types of cloud services. The aforementioned mutually different network routing information is, A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 1, characterized by including information regarding network performance and network transmission paths.

3. The aforementioned sample group data further includes the user AI workload definition information, In the step of generating the aforementioned multiple sample group data, A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 1, characterized in that, based on the user AI workload definition information, the sampled different cloud environment information and different network path information are combined with the user AI workload definition information to generate the plurality of sample group data.

4. A step of converting each of the multiple sample group data into multiple intermediate representation data based on a pre-set format, A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 3, further comprising the step of inputting each of the plurality of intermediate representation data into the neural network.

5. The aforementioned neural network is Predictions are made for each of the aforementioned multiple intermediate representation data, and the aforementioned multiple predicted values ​​are output for each of the aforementioned multiple intermediate representation data. The aforementioned multiple predicted values ​​are, A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 4, characterized in that it includes the time and cost required to execute the user AI workload definition information using the mutually different cloud environment information and the mutually different network path information.

6. The aforementioned multiple predicted values ​​further include resource utilization rates used to execute the user AI workload definition information, The aforementioned neural network is A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment, according to claim 5, characterized in that the predictions are performed in parallel for each of the plurality of intermediate representation data, and the plurality of predicted values ​​are output simultaneously for each of the plurality of intermediate representation data.

7. The user optimization requirement details information includes elements having different characteristics from each other. The elements having different characteristics are, The time and cost required to execute the user AI workload definition information, and the resource utilization rate used to execute the user AI workload definition information, further include: The step of receiving user optimization request details is: A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 1, further comprising the step of setting the weights of the elements having different characteristics from the user terminal.

8. The aforementioned optimal prediction calculation is: A score function is defined based on the multiple predicted values ​​and the weights of the elements having different characteristics. The score function is calculated, and for each of the calculated scores, a plurality of optimal predicted values ​​are sorted. A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 6, characterized in that, from among the sorted plurality of optimal prediction values, the optimal prediction value that satisfies the user AI workload definition information and the user optimization requirement details information.

9. A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 8, further comprising the step of generating optimal execution data based on the optimal predicted values.

10. The steps include identifying at least one cloud environment configuration information that satisfies the user AI workload definition information and the user optimization requirement details information based on the aforementioned optimal predicted value, A step of generating recommended information for the identified cloud environment configuration information, A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 1, further comprising the step of providing the generated recommendation information to the user terminal.

11. The aforementioned recommended information is, A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 10, characterized in that the cloud environment configuration information includes the estimated time and estimated price required to execute the user AI workload definition information.

12. The aforementioned recommended information is, A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 11, further comprising, in the cloud environment configuration information, an expected resource utilization rate used to execute the user AI workload definition information.

13. The aforementioned recommended information is, A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 10, characterized by including a first type of recommendation information identified based on the user optimization requirement details information and a second type of recommendation information identified based on pre-set conditions.

14. When either the first type of recommendation information or the second type of recommendation information is selected from the user terminal, the optimal execution data corresponding to the selected recommendation information is generated. A neural network-based method for generating an optimal execution plan for an AI workload in a hybrid and multi-cloud environment according to claim 13, further comprising the step of registering the aforementioned optimal execution data to a user account.

15. A communication unit that receives user AI workload definition information and user optimization request details from the user terminal, Includes a control unit that samples information from different cloud environments and different network paths to generate a plurality of sample group data that include the different cloud environments and different network paths, The control unit, Each of the above-mentioned multiple sample group data is input into a neural network, and multiple predicted values ​​for the above-mentioned multiple sample group data are received from the neural network. A neural network-based system for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments, characterized by using optimal prediction calculations to identify optimal predicted values ​​that satisfy the user AI workload definition information and the user optimization requirement details information.

16. The control unit, Based on the aforementioned optimal prediction values, at least one cloud environment configuration information that satisfies the user AI workload definition information and the user optimization requirement details information is identified. Recommended information is generated for the identified cloud environment configuration information. A neural network-based system for generating an optimal execution plan for AI workloads in hybrid and multi-cloud environments according to claim 15, characterized in that the generated recommendation information is provided to the user terminal.

17. A program executed by one or more processes on an electronic device and stored on a computer-readable recording medium, The aforementioned program, The steps include receiving user AI workload definition information and user optimization request details information from the user terminal, A step of sampling information from different cloud environments and different network paths to generate multiple sample group data containing the different cloud environments and different network paths, The steps include inputting each of the multiple intermediate representation data for the multiple sample group data into a neural network and receiving multiple predicted values ​​for the multiple sample group data from the neural network, A program stored on a computer-readable recording medium, characterized by including a command that performs the step of identifying an optimal predicted value that satisfies the user AI workload definition information and the user optimization requirement details information using an optimal prediction calculation.