Deployment of AI model based on AI model

By splitting the artificial intelligence model in a distributed system and using a second artificial intelligence model to predict the optimal configuration, the problem of efficiently processing AI models in a constrained computing environment is solved, achieving secure, resource-efficient execution and adaptive optimization of AI models.

CN121970031APending Publication Date: 2026-05-01INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2024-10-01
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

The need to efficiently process artificial intelligence models in constrained computing environments remains unmet, especially in distributed systems, where existing technologies struggle to achieve the optimal combination of resource utilization and model structure adaptability.

Method used

By splitting the first artificial intelligence model into input blocks, intermediate blocks, and output blocks, and using the second artificial intelligence model to predict the optimal split configuration, the deployment of the model in the distributed system is dynamically adjusted. By combining reinforcement learning and federated learning techniques, resource utilization and network conditions are optimized.

Benefits of technology

It enables efficient execution of artificial intelligence models on constrained computing devices, ensuring security and optimal resource utilization, adapting to changes in computing and network environments, and providing secure model splitting and efficient data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970031A_ABST
    Figure CN121970031A_ABST
Patent Text Reader

Abstract

The invention relates to a method. The method comprises the following steps: determining a current resource utilization rate condition in a distributed system; inputting the splitting configuration in use of the first artificial intelligence model and the current resource utilization rate condition into a second artificial intelligence model, wherein the second artificial intelligence model is configured to predict the splitting configuration for the first artificial intelligence model; receiving an output from the second artificial intelligence, the output indicating a current split configuration for the first artificial intelligence model; splitting the first artificial intelligence model by using the current splitting configuration; deploying the split first artificial intelligence model such that input and output blocks can be executed on one or more first computer systems of the set of first computer systems and such that intermediate blocks can be executed on a second computer system of the at least one second computer system; and executing the workload.
Need to check novelty before this filing date? Find Prior Art

Description

Deployment of AI models based on AI models Background Technology

[0001] This invention relates to the field of digital computer systems, and more particularly, to a method for performing workloads in a distributed system.

[0002] A radio access network (RAN) can provide access to mobile telecommunications system sites and coordinate cross-site resource management according to a protocol stack. The RAN can provide processing resources, which can be used, for example, for inference artificial intelligence (AI) models. However, there is a need to introduce AI models that require efficient processing on constrained computing environments. Summary of the Invention

[0003] Various embodiments provide a method, computer program product, and system for performing workloads in a distributed system, as described in the subject matter of the independent claims. Advantageous embodiments are described in the dependent claims. Embodiments of the invention may be freely combined with each other if they are not mutually exclusive.

[0004] In one aspect, the present invention relates to a method for performing a workload using a first artificial intelligence model in a distributed system, the distributed system including a set of first computer systems configured to be connected to at least one second computer system, the first artificial intelligence model being configured to receive specific inputs, process specific inputs, and provide specific outputs, the first artificial intelligence model being configured to be split into a set of one or more input blocks, intermediate blocks, and a set of one or more output blocks according to a splitting configuration, such that the set of one or more input blocks receives specific inputs and provides intermediate outputs, the intermediate blocks receive the intermediate outputs as inputs and provide another intermediate output, and the set of one or more output blocks receives the other intermediate outputs as inputs and provides the specific output; the method includes a splitting method comprising: receiving The method includes: receiving a request to execute a workload using a first artificial intelligence model, the workload including the specific input; determining the current resource utilization status in the distributed system; inputting the currently used split configuration and the current resource utilization status of the first artificial intelligence model into a second artificial intelligence model, the second artificial intelligence model being configured to predict a split configuration for the first artificial intelligence model; receiving an output from the second artificial intelligence model indicating the current split configuration for the first artificial intelligence model; the method further includes: splitting the first artificial intelligence model using the current split configuration; deploying the split first artificial intelligence model such that input blocks and output blocks can be executed on one or more first computer systems in the group of first computer systems, and that intermediate blocks can be executed on second computer systems in at least one second computer system; and executing the workload.

[0005] In one aspect, the present invention relates to a computer program product comprising a computer-readable storage medium having computer-readable program code implemented therewith, the computer-readable program code being configured to implement the methods of the embodiments described above.

[0006] In one aspect, the present invention relates to a computer system for performing a workload using a first artificial intelligence model in a distributed system, the distributed system comprising a set of first computer systems configured to be connected to at least one second computer system, the first artificial intelligence model being configured to receive specific inputs, process specific inputs, and provide specific outputs, the first artificial intelligence model being configured to be split into a set of one or more input blocks, intermediate blocks, and a set of one or more output blocks according to a splitting configuration, such that the set of one or more input blocks receives specific inputs and provides intermediate outputs, the intermediate blocks receive intermediate outputs as inputs and provide another intermediate output, and the set of one or more output blocks receives the other intermediate outputs as inputs and provides the specific output; the computer system is configured to use... The algorithm: receives a request to execute a workload using a first artificial intelligence model, the workload including the specific input; determines the current resource utilization status in the distributed system; inputs the currently used split configuration and the current resource utilization status of the first artificial intelligence model into a second artificial intelligence model, the second artificial intelligence model being configured to predict the split configuration for the first artificial intelligence model; receives an output from the second artificial intelligence model indicating the current split configuration for the first artificial intelligence model; splits the first artificial intelligence model using the current split configuration; deploys the split first artificial intelligence model such that input blocks and output blocks can be executed on one or more first computer systems in the group of first computer systems and that intermediate blocks can be executed on the second computer systems in the at least one second computer system; and executes the workload. Attached Figure Description

[0007] Embodiments of the invention will be explained in more detail below by way of example only, with reference to the accompanying drawings, wherein:

[0008] Figure 1 is a block diagram of a wireless communication system based on an example of this topic.

[0009] Figure 2 is a flowchart of a method for performing workloads in a distributed system, based on an example from this topic.

[0010] Figure 3 is a schematic diagram illustrating the deployment of a split base model based on an example of this topic.

[0011] Figure 4 is a schematic diagram illustrating a method for performing workloads in a distributed system, based on an example of this topic.

[0012] Figure 5 is a schematic diagram illustrating a method for performing workloads in a distributed system, based on an example from this topic.

[0013] Figure 6 shows a computing environment based on an example from this topic.

[0014] Figure 7 depicts a cloud computing environment according to an embodiment of the present invention.

[0015] Figure 8 depicts an abstract model layer according to an embodiment of the present invention. Detailed Implementation

[0016] The description of various embodiments of the invention is presented for illustrative purposes and is not intended to be exhaustive or to limit the embodiments to those disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or improvements relative to technologies found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0017] The first AI model can be configured to perform a task. The task can refer to the type of prediction or reasoning being performed. The task can be based on the question or inquiry being asked and the available data. For example, the task can be a classification task, a clustering task, or a prediction task. For instance, a classification task assigns data to categories, and a clustering task groups data based on similarity. The first AI model can perform the task by receiving input data, processing the input data using a set of learnable parameters, and providing an output representing the task's outcome. The first AI model can be provided as a deep neural network, a transformer, or another AI model that can be broken down into blocks as described herein.

[0018] The first AI model can be configured to receive input X and provide output Y. The first AI model can be configured to split into a set of one or more input blocks, intermediate blocks, and a set of one or more output blocks according to a splitting configuration. For example, the first AI model can be configured to split into input blocks, intermediate blocks, and output blocks according to the splitting configuration. The splitting of the first AI model can be performed into any number of blocks, provided that if the same input X is used as input to the splitting model, the same output Y can be obtained. The splitting configuration can define how the first AI model can be split. For example, the splitting configuration can include the number of blocks into which the first AI model can be split and the splitting ratio indicating the portion of the task of the first AI model that can be performed by each block.

[0019] This topic provides a precise and systematic decomposition of a first artificial intelligence model by using a second artificial intelligence model. The terms "first" or "second" are used as labels for the nouns preceding them and do not imply any type of ordering (e.g., spatial, temporal, logical) unless explicitly defined as such. The second artificial intelligence model can be a trained artificial intelligence model. The second artificial intelligence model can be trained to receive the current resource utilization status in a distributed system and a splitting configuration for decomposing the first artificial intelligence model as input. The splitting configuration for decomposing the first artificial intelligence model can be referred to as the "splitting configuration in use." The splitting configuration in use can be a splitting configuration used by a computer system (e.g., the first computer system) that performs the method to decompose the first artificial intelligence model. In response to receiving said input, the second artificial intelligence model can predict current or new splitting configurations that can be used to decompose the first artificial intelligence model. The input of the second artificial intelligence model can be referred to as the splitting input, and the output of the second artificial intelligence model can be referred to as the splitting output.

[0020] The splitting inputs can include values ​​representing the current resource utilization status and the splitting configuration in use. For example, a utilization parameter describing resource utilization can be provided, allowing the determination of the current resource utilization status to be performed by evaluating the utilization parameter. The splitting configuration in use can be represented or defined by the values ​​of the configuration parameter. Therefore, the splitting inputs of the second AI model can include the values ​​of the utilization parameter and the configuration parameter, respectively. The splitting outputs of the second AI model can include the values ​​of the configuration parameter, respectively.

[0021] The predicted current splitting configuration can be used to split the first AI model. The splitting of the first AI model can produce three blocks S1, S2, and S3, where S1 is the input block, S2 is the intermediate block, and S3 is the output block. In one example, the number of blocks can be higher, where n1 and n2 are greater than or equal to two, by splitting the input block S1 into n1 input sub-blocks S1(a), S1(b)...S1(n1) and the output block S3 into n2 output sub-blocks S3(a), S3(b)...S3(n2). This makes the number of blocks equal to n1 + n2 + 1. In the following text, and for simplicity, each sub-block in the input sub-blocks S1(a), S1(b)...S1(n1) can be referred to as an input block. Similarly, each sub-block in the output sub-blocks S1(a), S1(b)...S1(n1) can be referred to as an output block. That is, the first AI model can be split into one or more input blocks, one or more output blocks, and one intermediate block.

[0022] The breakdown of the first artificial intelligence model can produce multiple blocks, each with its own structure. For example, the structure of a block can refer to the types of inputs and outputs of the block, as well as the number and types of operations performed by the block. For instance, if the first artificial intelligence model is a deep neural network, the structure of a block can be defined by the number of layers belonging to that block, where the number of layers can represent the specific number and types of operations performed by the first artificial intelligence model.

[0023] The execution of a first artificial intelligence model can be performed according to an execution pipeline. The execution pipeline may include three or more sequential execution stages, each configured to receive input, process the input using a subset of learnable parameters, and provide an output. The input to an execution stage that is not a first execution stage may be the output of a preceding execution stage. For example, input block S1 may represent one or more first execution stages of the pipeline, output block S3 may represent one or more final execution stages of the pipeline, and intermediate block S2 may represent the remaining execution stages. For example, in the case of a deep neural network, execution stages may represent the processing of one or more layers of the deep neural network. The processing performed for a network layer may include, for example, weighting operations, convolution operations, or activation operations. The first AI model may be split at two slicing layers, and the intermediate output may include, for example, slicing layer activations. Generally, the model splitting of the present invention can be applied to various AI architectures, such as CNNs or other AI architectures, such as Transformer, ResNet, LSTM, or AI models that can be executed according to the above-described execution pipeline.

[0024] Different workloads can use a first AI model to perform tasks assigned to that model. A workload can refer to one or more software applications and the data accessed by those applications. For example, a workload may include multiple steps, one or more of which may include feeding input data to the first AI model to obtain an output associated with that input data. The output can also be used by other steps in the workflow; for example, if the first AI model is trained to predict whether a communication channel is reliable, a workflow for data scheduling can use the first AI model to find reliable channels for communication scheduling on those channels. In another example, the first AI model could be trained for facial recognition of a user, where the results of the recognition can be used to enable the user to access a service.

[0025] These workloads can be efficiently executed using a distributed system in accordance with this topic. A distributed system includes multiple first computer systems remotely connected to one or more second computer systems. The first computer systems can be local computer systems, for example, accessible to users. The second computer systems may not be part of the first computer systems. The second computer systems are located remotely from the first computer systems. The first computer systems can be configured to connect to the second computer systems via wired and / or wireless digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), radio access networks (RANs), metropolitan area networks (MANs), wide area networks (WANs), microwave access interoperability (WIMAX), wireless local area networks (WLANs), all or part of the Internet, one or more other communication systems in one or more locations, or combinations thereof.

[0026] Therefore, this topic provides an accurate method for optimal execution of a workflow using available resources. Compared to existing techniques, this topic not only searches for resources to execute the program but also adapts to the structure of the program itself to find the optimal combination of resources. This example seeks a trade-off between secure execution of a first AI model and available processing resources. Execution can be secure because the inputs and outputs of the first AI model are processed locally at one or more first computer systems. This prevents problems such as model inversion attacks by malicious parties and reverse engineering attempts on sensitive input data and / or output labels by malicious parties and / or well-intentioned but curious servers (e.g., in the cloud).

[0027] This topic provides optimal model splits, such as S=(S1, S2, S3), given the current deployment-specific compute and network conditions at an edge node, that satisfy stringent Quality of Service (QoS) or other performance metrics. The optimal model split can be adaptively learned to derive a set of intelligent split configurations that define split ratios that can be applied to unpredictable and changing compute and network conditions. Changing environments and requirements can be mapped to appropriate input variables for the splitting method.

[0028] The second AI model can be a trained model, where training can be performed to find the optimal split configuration for a given environmental state. The second AI model can be trained periodically (e.g., on a time-period basis) to provide accurate policies that reflect changing conditions in the distributed system. The second AI model can be trained using reinforcement learning or other learning techniques (such as supervised learning, self-supervised learning, transfer learning, evolutionary algorithms, semi-supervised learning, federated learning), but is not limited to these. The second AI model can be trained, for example, using reinforcement learning, and therefore can be called a reinforcement learning model.

[0029] In one example, the second AI model is a reinforcement learning model. The state of the reinforcement learning model is defined by the current resource utilization status and the splitting configuration currently in use. A distributed system can define an environment for the reinforcement learning model. The state of the reinforcement learning model can be the state of the environment. The policy of the reinforcement learning model can define actions that change the splitting configuration currently in use of the first AI model based on the state. The reinforcement learning model can include a trained agent that uses the policy to define actions that change the splitting configuration currently in use using the current state of the environment, or a trained agent that uses the policy to predict the current splitting configuration using the current state of the environment. Reinforcement learning can enable the agent to learn an optimal policy for optimizing (e.g., maximizing) the reward of splitting the first AI model. The reward can be defined, for example, using the execution time of the split first AI model or the duration of an inference execution of the split first AI model. The trained agent can be configured to receive splitting inputs and predict or define the current or new splitting configuration for splitting the first AI model.

[0030] The method of the present invention can be executed by a first computer system (e.g., an edge node) within a first set of computer systems. The input and output blocks of the resulting split first artificial intelligence model can be executed by the first computer system, and the intermediate blocks can be executed by a second computer system. The current resource utilization status can, for example, depend on the first computer system executing the method thereon. For instance, the current resource utilization status can indicate the resource utilization of the first computer system executing the method thereon and the network conditions collected by the first computer system executing the method thereon (from its perspective).

[0031] According to one example, a second artificial intelligence model is trained on multiple first computer systems using federated learning techniques. Each of the multiple first computer systems can include an instance of the second artificial intelligence model and can train that instance using its own collected training data, which includes the currently used split configuration and resource utilization status used by each first computer system. The resource utilization status may depend on the first computer system that determines it. Each first computer system may have different resource utilization statuses, depending on other workloads running on that particular computer system. Therefore, each instance of the second AI model may receive different inputs to process it. The first computer systems may not share data or current resource utilization with each other. And resource utilization can be different; for example, MEC edge nodes may have additional services. This example enables the collaborative sharing of knowledge and policies from similar distributed inference tasks to dynamically update the global policy set, thereby achieving a more informed and robust model splitting mechanism.

[0032] According to one example, training based on federated learning techniques includes: combining learnable weights of a second artificial intelligence model in a corresponding first computer system to produce a combined second artificial intelligence model, wherein the second artificial intelligence model used in the splitting method is the combined second artificial intelligence model.

[0033] For example, a control unit (e.g., in one of a series of first computer systems) may request each of the multiple first computer systems to send a copy of the learnable weights of its respective instance of the second artificial intelligence model. The control unit may aggregate the weights via a specified federated learning aggregation strategy (such as federated averaging). The control unit may also request each of the multiple first computer systems to receive a globally federated copy of the second artificial intelligence model and subsequently update the corresponding instance using the globally federated second artificial intelligence model.

[0034] Based on the example, the second artificial intelligence model is the Deep Q Neural Network (DQNN).

[0035] According to one example, the method is repeatedly executed, where the splitting method also includes: retraining an instance of the second AI model; and after successful retraining, replacing the second AI model with the retrained second AI model for further execution of the splitting method. For example, if retraining ends successfully in the nth iteration of the splitting method, the next (n+1)th and subsequent iterations can use the retrained second AI model.

[0036] The first artificial intelligence model can be deployed using a deployment configuration. This provides flexible and optimal splitting of the model's execution. Indeed, splitting based on resource utilization can indicate whether the block can be executed on a specific first computer system (i.e., the first computer system executing the method) and a specific second computer system. However, the deployment configuration may be advantageous when the input and output blocks can be executed on one or more first computer systems in addition to or instead of the specific first computer system executing the method, and when the intermediate blocks are to be deployed on another or an additional second computer system in addition to the specific second computer system. The deployment configuration can be defined by at least one of the following: the second computer system for executing the intermediate blocks, and one or more first computer systems for executing the input and output blocks of the first artificial intelligence model. For example, when the first artificial intelligence model is split into three blocks S1, S2, and S3, the one or more first computer systems can process the input and output blocks S1 and S3. When the first artificial intelligence model is split into a higher number of blocks (e.g., n1+n2+1 blocks) S1(a), S1(b)...S1(n1), S2, S3(a), S3(b)... and S3(n2), the one or more first computer systems can process the n1+n2 input blocks and output blocks S1(a), S1(b)...S1(n1), S3(a), S3(b)... and S3(n2).

[0037] This topic provides useful techniques for defining deployment configurations using current resource utilization status. The deployment configuration can also be defined based on estimates of the resources required by each processing stage or step of the first AI model. The resources required by each block of the first AI model can be referred to as the required block resources.

[0038] For example, deployment configuration can be determined using current resource utilization status. As an example, defining deployment configuration using current resource utilization status includes: performing capacity profiling on a first and a second computer system to determine whether each of the first and second computer systems can execute one or more blocks of an artificial intelligence model; and defining the deployment configuration based on the capacity profiling.

[0039] According to one example, the utilization parameter includes at least one of the following: the utilization level of the network resources of the distributed system by the first computer system executing the method; the utilization level of resources of one or more first computer systems including the first computer system executing the method; the utilization level of resources of the second computer system; or resources in each of the one or more first computer systems and the second computer system. For example, the utilization parameter may include a set of hardware specifications, such as RAM, CPU, memory, cache, information about computing conditions (e.g., hardware utilization, node occupancy), and information about network conditions (e.g., link reliability, latency, data transfer rate). According to one example, the configuration parameter may include any one of the following: the maximum number of blocks in the model, the minimum number of blocks in the model, and the maximum resource usage per block. In one example, the processing resources required by each block in the blocks of the first artificial intelligence model can be estimated. The current resource utilization status of the distributed system can be used, given the estimated processing resources, to identify the first and second computer systems that can run the blocks.

[0040] As an example, the distributed system is a wireless communication system where the first computer system is a multi-access edge computing (MEC) node, and the second computer system is a cloud system. The second computer system can be part of a public cloud, a private cloud, or a hybrid cloud. For example, multiple public clouds, private clouds, and hybrid clouds can be provided, where each cloud can provide resources for the second computer system (e.g., a cloud can provide a second computer system as a virtual machine). For example, for defining a deployment configuration, a specific cloud can be selected first, and then the second computer system provided by the selected cloud can be used for the deployment configuration. Multiple clouds can be provided by the same cloud service provider or different cloud service providers.

[0041] As an example, the first artificial intelligence model can be provided as a neural network (e.g., a deep neural network), a transformer, or any other artificial intelligence model and corresponding architecture (e.g., in terms of parallelization, distribution) that can be broken down into blocks as described in the methods described in this paper. Input blocks can represent the first network layer, intermediate blocks can represent intermediate network layers, and output blocks can represent the final network layer.

[0042] According to one example, the workload involves at least one of the following: data analytics, sensor measurement fusion from different sources, image analytics, processing data streams destined for the cloud or another internal or external workload. The internal workload can be a workload running on the first computer system performing this method, while the external workload can be a workload executed on a system outside the first computer system and used by the first computer system. Examples include any other intensive workload, such as internal workloads from other running services, external workloads from services that provide computational instance operations (e.g., network operators in MEC nodes), other ML model inference and training, etc.

[0043] The execution of the workflow involves executing a first AI model once or multiple times. Execution of the first AI model may include providing input to the first AI model and receiving its output. According to one example, execution of the first AI model includes: for every two consecutive blocks of the first AI model deployed on different systems: encoding the output of the first block in the two blocks using an encoding protocol, and sending the encoded output to either a first computer system or a second computer system to be used as input to the second block in the two blocks. In one example where the first AI model is split into more than three blocks, if input blocks S1(a) and S1(b) are deployed on two separate first computer systems, the output of input block S1(a) can be encoded using an encoding protocol, and the encoded output can be sent to the first computer system where S1(b) is deployed. The output can be decoded using an encoding protocol at the receiving first computer system and then used as input to block S1(b). In another example where the first AI model is split into three blocks, the output of input block S1 can be encoded using an encoding protocol, and the encoded output can be sent to a second computer system where S2 is deployed. The output can be decoded using an encoding protocol at the second computer system and then used as input to block S2. Similarly, the output of main block S2 can be encoded using an encoding protocol, and the encoded output can be sent to the first computer system where input block S3 is deployed. The output can be decoded at the first computer system using the encoding protocol and then used as input to block S3.

[0044] According to one example, the encoding protocol includes at least one of compression or encryption.

[0045] An encoding protocol defines methods for encoding raw data to obtain encoded data, and defines corresponding decoding methods that enable the recovery of the original data from the encoded data. Data encoding can include any of the following: encryption, compression, scrambling, formatting, or the allocation or interpretation of specific bit patterns in the data. This ensures the security of data communication. Alternatively or additionally, this ensures the efficient use of network resources, for example, because compression can reduce data size. For example, in the case of encoding by compressing the output, decoding of the compressed output is performed by decompressing the compressed output. In the case of encoding by encrypting the output, decoding of the encrypted output is performed by decrypting the encrypted output.

[0046] This topic offers the following advantages: It can introduce data-saving operations suitable for artificial intelligence models, enabling efficient operation on constrained computing devices. It can ensure optimal data transfer sizes to guarantee efficient utilization of network resources. It can adjust data-saving operations to accommodate changes in computing resource availability on constrained computing devices. It can dynamically schedule the model split ratio of artificial intelligence models to distribute inference tasks across constrained first-level computing systems (e.g., edge computing devices) and second-level computing systems (e.g., cloud server instances). It can introduce security for distributed inference using large and complex artificial intelligence models (e.g., base models), which can be efficiently handled in constrained computing environments such as edge computing and Internet of Things (IoT) devices.

[0047] According to one example, one or more output blocks can be removed from a first computer system before executing one or more input blocks. For example, a management server can distribute input blocks, intermediate blocks, and output blocks to one or more first and second computer systems respectively, while ensuring that one or more first computer systems process only one or more input blocks or one or more output blocks per time instance to ensure maximum utilization of hardware resources. For example, input block configurations can be removed after execution, encoding, and transmission to free up computational power for output blocks. The management server can, for example, be configured to connect to and control the operation of the first and second computer systems.

[0048] As an example, after executing one or more input blocks, one or more output blocks can be deployed to one or more first computer systems. For instance, one or more output blocks can be downloaded from a management server after executing one or more input blocks. This can further improve the resource utilization of the first computer systems because processing resources used to maintain the output blocks can be saved when they are not in use.

[0049] As an example, after executing one or more input blocks, the input blocks can be deleted and one or more output blocks can be deployed to the first computer system. For instance, one or more output blocks can be downloaded from a management server after executing one or more input blocks. This can further improve the resource utilization of the first computer system, since only one block type can be processed and managed by the first computer system(s) at a time.

[0050] According to one example, the execution of a first artificial intelligence model includes the execution of a series of processing steps, wherein the splitting of the first artificial intelligence model allows the input block to execute an initial number (N1) of consecutive processing steps and the output block to execute a final number (N3) of consecutive processing steps, wherein intermediate blocks are configured to execute a second number (N2) of consecutive processing steps immediately following the initial processing steps of the input block, and the sum of the first, second, and third numbers is the total number of processing steps in the artificial intelligence model, i.e., N1 + N2 + N3 is the number of processing steps in the first artificial intelligence model. The execution phase of the previously defined first AI model may include one or more processing steps of the first AI model. In the case where the input block is further split into multiple input blocks, N1 refers to the initial consecutive processing steps executed by all input blocks. Similarly, in the case where the output block is further split into multiple output blocks, N3 refers to the final consecutive processing steps executed by all output blocks.

[0051] The intermediate block can be called the main block because it can include most of the processing steps of the first artificial intelligence model.

[0052] As an example, the first AI model is a trained model. In this case, performing the first AI model using this method is the inference of the first AI model.

[0053] As an example, the first computer system has a smaller amount of processing resources than the second computer system. The second computer system can be any computer system, for example, with processing resources for executing any defined intermediate blocks of the artificial intelligence model. For instance, the second computer system can be any computer system with processing resources for executing the entire first artificial intelligence model.

[0054] According to one example, a second computer system is provided as a service within a cloud computing environment. In one example, the second computer system may be provided as a cloud instance within the cloud computing environment. A cloud instance can be server resources provided by a cloud service. In one example, the second computer system may be implemented using one or more functional abstraction layers provided by the cloud computing environment; for example, the hardware and software resources of the second computer system may be provided by the hardware and software layers of the cloud computing environment. The workload layer of the cloud computing environment may, for example, be used to implement the steps to be performed by the second computer system. The cloud computing environment can remain unaware of any data and output labels because it does not possess a complete first AI model, and external model inversion and / or reverse engineering attacks can be mitigated through secure model coding, thereby protecting the data. Therefore, inference can uniquely and securely occur at the edge device, using the cloud computing environment as a purely computational and processing instance without knowledge of the specific use cases of the edge device or the results of the inference.

[0055] As an example, the first AI model is the foundation model. The foundation model can be a large-scale AI model trained on massive amounts of data, producing models capable of adapting to a wide range of downstream tasks. Examples of foundation models include the Bidirectional Encoder Representation (BERT) from Transformers and the generation of pre-trained Transformer n-series models (GPT-n series). The initial and final foundation model (FM) layers are processed on the device, and their intermediate slice activations are securely transmitted (received) through application compression (decompression), ensuring efficient low-bandwidth transmission for communication.

[0056] In one example, the first artificial intelligence model is a deep neural network, where the input block represents the initial network layer, the intermediate block represents the intermediate network layer, and the output block represents the final network layer. Continuing with this example, each processing step of the first artificial intelligence model can be represented as a process performed by the corresponding layer of the deep neural network. That is, the input block comprises N1 initial layers of the deep neural network, the intermediate block comprises N2 layers of the deep neural network, and the output block comprises N3 final layers of the deep neural network, where the total number of layers in the deep neural network is N1 + N2 + N3.

[0057] According to one example, the first computer system is any of the following: an edge device (e.g., such as an MEC node), a user equipment (UE), or an Internet of Things (IoT) device. This example can be seamlessly integrated into a wireless or mobile communication system. The mobile communication system provides wireless connectivity to users. Users may include, for example, mobile devices, tablets, laptops, or individuals. The mobile communication system may include a radio access network (RAN) and a core network. The core network may provide Internet Protocol (IP) connectivity to the radio access network. The radio access network may use radio equipment such as base stations to manage the radio spectrum for users. The radio access network may enable packet processing according to a processing pipeline. The processing pipeline has different layers. These layers include a baseband processing layer and a radio frequency (RF) processing layer. The baseband processing layer may be defined according to a protocol stack and may be executed by a baseband unit, which is contained within an edge device.

[0058] A baseband unit can be associated with one or more base stations. For example, each of the one or more base stations can serve users located within the service geographic area or cell of the base station. The baseband unit can process the baseband signals of the users of the one or more base stations it serves. Therefore, the baseband unit is considered to be serving said users. The baseband unit can implement layers of the protocol stack, such as the Packet Data Convergence Protocol (PDCP) layer, Radio Link Control (RLC) layer, Media Access Control (MAC) layer, and Physical (PHY) layer. In one example, the baseband unit can be divided into functional entities configured to perform corresponding functions; for example, a function can implement one or more layers of the stack protocol. For example, the baseband unit can be divided into two functional entities named a Centralized Unit (CU) and a Distributed Unit (DU). The CU can provide support for higher layers of the protocol stack, such as the PDCP layer, while the DU can provide support for lower layers of the protocol stack, such as the RLC, MAC, and Physical layers.

[0059] The baseband unit can be implemented using the specific hardware and software configuration of the edge device. The software configuration of the baseband unit may include an operating system and software modules for performing the functions of the baseband unit. Additionally, the software configuration may indicate one or more vendors providing the software configuration. For example, the operating system and software modules may be provided by one or more vendors. The hardware configuration may include storage resources, data communication resources, and processing resources. Additionally, the hardware configuration may indicate one or more vendors providing the hardware configuration. Resources may be provided by one or more vendors.

[0060] Figure 1 illustrates a schematic diagram of a wireless communication system based on an example of this topic.

[0061] The wireless communication system 100 includes a core network 101 and a radio access network 102. The radio access network 102 may include, but is not limited to, remote radio components 107 equipped by base stations 109 and 111. Each base station 109 or 111 may include a remote radio unit (RRU) with an antenna and may serve UE 120 in its respective cells 121 and 122. The radio access network 102 may also include a first computer system 103. For simplicity, only three first computer systems are shown, but the system is not limited thereto. Furthermore, for the purpose of simplifying the figures, only some components of one first computer system are described.

[0062] The first computer system 103 may include, for example, a group of one or more baseband units (BBUs) 105.1-n. The baseband units can be connected to corresponding RRUs in the remote radio assembly 107 via fiber or cable 113. The first computer system 103 may be configured to connect to the core network 101 via a backhaul link 115. The first computer system 103 may include a central unit 117 configured to control the operation and deployment of the baseband units 105.1-n. For example, each first computer system in the first computer system 103 may be provided as an MEC node. MEC nodes can improve user service (e.g., with low latency). The first computer system 103 may, for example, use advanced techniques (e.g., for image analysis) to process data provided by the BBUs. The radio access network 102 may include a control unit 110 for managing workloads in the wireless communication system 100. Although shown as a separate component, in another example, the control unit 110 may be part of the first computer system 103.

[0063] The remote radio component 107 and the first computer system 103 can be configured to connect to the cloud computing environment 130. The cloud computing environment 130 may include a second computer system 131. In one example, the second computer system 131 may be provided as a cloud instance within the cloud computing environment 130.

[0064] In one example implementation, the cloud computing environment 130 may be provided, for example, as described with reference to Figures 8 and 9. For instance, the second computer system 131 may be implemented using one or more functional abstraction layers provided by the cloud computing environment 130; for example, the hardware and software resources of the second computer system 131 may be provided by the hardware and software layers of the cloud computing environment 130. The workload layer of the cloud computing environment 130 may, for example, consist of a main block for implementing AI models executed by the second computer system 131.

[0065] In one example implementation, system 100 may be provided as an Open Radio Access Network (O-RAN), wherein a first computer system 103 may be in one or more edge sites, and a remote radio component 107 may be in one or more cell sites.

[0066] The first computer system 103 and processing unit 110 may include a second artificial intelligence model. The second artificial intelligence model can be trained to receive the current resource utilization status and the currently used splitting configuration of the first artificial intelligence model, and predict the splitting configuration for the first artificial intelligence model. The first artificial intelligence model can be configured to receive specific inputs, process specific inputs, and provide specific outputs. The first artificial intelligence model is configured to split according to the splitting configuration into a set of one or more input blocks, intermediate blocks, and a set of one or more output blocks, such that the set of one or more input blocks receives specific inputs and provides intermediate outputs, the intermediate blocks receive intermediate outputs as inputs and provide another intermediate output, and the set of one or more output blocks receives another intermediate output as inputs and provides the specific output.

[0067] Figure 2 is a flowchart of a method for executing an artificial intelligence model according to an example of this subject matter. For illustrative purposes, the method of Figure 2 can be implemented in the system of Figure 1, but is not limited thereto. The method can be executed, for example, by control unit 110 or first computer system 103.

[0068] A request to perform a workload using an artificial intelligence model may be received in step 201. The workload may include the step of receiving specific inputs to be used as inputs to a first artificial intelligence model. The current resource utilization status in the distributed system may be determined in step 203.

[0069] In step 205, the current resource utilization status and the split configuration in use can be input into the second artificial intelligence model so that in step 206, the predicted current split configuration for the first artificial intelligence model can be received from the second artificial intelligence model.

[0070] The first AI model can be split using the current split configuration in step 207. The first AI model can be deployed in step 208. The workload can be executed in step 209.

[0071] This method can be performed, for example, by a first computer system 103. In step 209, the first computer system 209 can control the execution of workloads in the first computer system 103 and other systems (such as a second computer system 131). Executing the workloads may include executing input and output blocks of a deployed first AI model at the first computer system and executing intermediate blocks at the second computer system.

[0072] This method can efficiently, adaptively, and collaboratively learn and infer the optimal model split ratio. Adaptive and self-learning rules (especially regarding time-varying edge environments) can produce the optimal split ratio, for example, reduced inference processing overhead at edge nodes, reduced susceptibility to model inversion attacks if edge nodes are scheduled to process almost no layers, and reduced latency due to balanced and not too long computations.

[0073] Figure 3 is a schematic diagram illustrating the deployment of the split base model. The base model can be split into three blocks. As shown, the first computer system 401 may include an input block 404 and an output block 406, while the remote second computer system 402 includes a master block 405. Input data 403 can be received at the first computer system 401 and processed by the input block 404. The output of the input block 404 can be processed in the second computer system 402 by the master block 405. Conversely, the output of the master block 405 can be processed by the output block 406 to obtain the inference result 407 of the input data 403.

[0074] Figure 4 is a schematic diagram illustrating a method for workload orchestration in a distributed system, based on an example from this topic.

[0075] The second artificial intelligence model can be, for example, a reinforcement learning model. The first artificial intelligence model can be a base model (FM). Data about the computation, network, and operating environment in edge node 503 can be collected in step 501. The collected data can be stored in storage system 502 (e.g., a database). The collected data can be used to perform Input Mapping (IM) for reinforcement learning (RL) in step 505. The IM method maps input data and other information to states, actions, and rewards related to RL. Given the computation and network environment, general information about how to construct the model split ratio (number of layers, percentage, etc.) and evaluations of the success and failure of split inference, the IM method transforms the above input data into RL-based states (current operating conditions), actions (selected model re-split), and rewards (evaluation and scoring). The results of the IM method can be used to perform an Adaptive Split Optimization (ASO) method in step 506. The ASO method can adaptively learn the optimal split ratio based on the current environment and performance requirements via reinforcement learning. Given the original FM architecture, information about the computation and network conditions of the current edge node 503, potential QoS, performance, or security constraints, and a mapping of these to states, actions, and rewards associated with reinforcement learning (RL), this method implements a single-agent neural network to learn and derive an optimal split ratio S*=(S1*, S2*, S3*) via reward-based RL evaluation. The learned optimal split ratio can be used in step 507 to split the FM into input block S1, intermediate block S2, and output block S3. Input block S1 and output block S3 can be processed on edge device 503, and intermediate block S2 can be processed at a computer server in cloud system 530. The method in Figure 4 may also include a step 508 of performing policy federation (PF) using the splitting policy learned in step 507 and model splitting policies 511 learned from other edge nodes. Policy federation can synchronize the learned model splitting policies across similar edge nodes to obtain a better and more adaptive global policy. Given a local model splitting policy for each participating edge node and information about the expected frequency of joint synchronization among them, the method achieves federated aggregation of said policies via federated learning (FL) to obtain a more robust global policy set since each participating node has already collaboratively shared its corresponding knowledge. The global splitting policy can be sent (512) to edge node 503.

[0076] Therefore, as shown in Figure 4, the initial and final foundational model (FM) layers (S1 and S3) can be processed at the edge nodes, and their intermediate slicing activations are securely transmitted, received, and processed because the cloud environment remains unaware of any data and output labels and does not possess the complete FM. Thus, inference can occur exclusively and securely at the edge nodes, leveraging the cloud environment as a pure computation and processing instance for S2. To address the changing and potentially harmful computational and network conditions at the edge, while maintaining certain performance, service, and security standards, adaptive split optimization (ASO) based on reinforcement learning (RL) can be employed to adaptively learn the optimal model split ratio for distributed split inference. For this purpose, the current computational and network conditions (and others) may need to be transformed into appropriate RL-based metrics via an input mapping (IM). Additionally, a policy federation (PF) can be invoked to aggregate learned policies and rules from nodes in the connected network designed to perform similar splitting tasks. Thus, valuable knowledge can be shared to obtain improved global policies.

[0077] Figure 5 is a schematic diagram illustrating a method for workload orchestration in a distributed system, based on an example from this topic.

[0078] The second artificial intelligence model can be, for example, a reinforcement learning model. The first artificial intelligence model can be a base model (FM). Data about the computation, network, and operating environment in the edge node 603 can be collected in step 601. This can provide a description of the current situation. The collected data can be used to perform input mapping (IM) in step 604. IM can map the input data and other information to states, actions, and rewards related to RL. The collected data may include, for example, the following items 1) through 6): 1) a set of (MEC) edge node hardware specifications, such as RAM, CPU, memory, cache, 2) information about computing conditions (hardware utilization, node occupancy), 3) information about network conditions (link reliability, latency, data transmission speed), 4) information about the FM architecture and how to perform effective architecture-dependent splits, e.g., splits after layer X; residual units Y; attention heads Z, 5) information about the desired split architecture between edge nodes and cloud servers, e.g., a U-shaped three-way split S = (S1, S2, S3) and 6) information and capabilities for measuring the total duration of successful inference processing, e.g., given as t_total = SUM_i(t_transmission_i) + SUM_j(t_processing_j). Using one or more of items 1) through 6), the input mapping can define the state associated with RL as: s=[computation_profile, network_profile, in-use_split_configuration], the action as: a=[resplit_configuration], and the reward as: r=[-t_total], by mapping given information about the operational situation to numerical values ​​that can be used as inputs to RL-based algorithms. This can be accomplished through profiling, discretization, scaling, and thresholding techniques as described below: computation_profile and network_profile, etc., can each represent integer values ​​between 1 and N (best to worst) based on the capacity profile describing the current availability of computational and network resources. Here, N is an adjustable granularity factor if more case differentiation is required. Capacity profiling can be implemented using rule-based tables, voting systems, or specific heuristics (such as those used in task scheduling and load balancing). The `in-use_split_configuration` option specifies the FM split ratio (e.g., [0.2, 0.7, 0.1]), which is used for U-shaped splits of large neural networks.`resplit_configuration` represents a new `current_resplit_configuration` vector as part of the actions taken during the RL process, adapting the model splitting to the current operational situation. Within this context, security-related aspects such as a threshold for the minimum processing ratio can be employed to ensure robustness against model inversion attacks. `-t_total` represents a negative measure of the total inference duration and is used as part of the RL process to evaluate and reward the current policy / strategy.

[0079] The obtained states, actions, and rewards can be used for adaptive split optimization in step 606. The ASO method can adaptively learn the optimal split ratio based on the current environment and performance requirements via reinforcement learning. For example, given the RL-related states, actions, and rewards of the current operation provided by the input mapping method, and given sufficient computational resources for ASO (on an edge device or at a monitoring instance / server connected to the edge), adaptive split optimization enables a scalable RL optimization process to cope with the large surge in complexity that occurs when the number of states and actions becomes very large. Deep Q-learning can be described for this purpose as follows.

[0080] Deep Q-Neural Networks (DQNNs) can be used to process RL tasks by taking into account the available computational resources allocated to ASO. The classical Bellman equation for Q-values, which measures the total reward for actions taken within the system, can be augmented by a weight vector w, which represents the weights and parameters of the DQNN. , where alpha[0, 1] is the discount factor, as is commonly used in Q-learning, and beta[0, 1] is an additional factor that can be adopted to concentrate future reward estimates onto the measured value and can optionally be learned to obtain better performance and stability. Using the Q objective function described above, the mean squared error (MSE) metric is further defined as the loss function of the DQNN, i.e. Next, the DQNN is trained for a specified number of epochs or until convergence based on available observations (some collected dataset), ultimately producing a policy π for the optimal model split. For continued learning, ASO can replicate the learned DQNN so that one instance can be used for further training with new observations (policy updater), and another instance can be used as an independent inference engine (policy provider 607) to infer the optimal model split ratio for the current operation. This mechanism can be implemented as follows: while policy provider 607 provides fast inference results, policy updater 608 retrains the DQNN in parallel with the same input. After successful retraining of the DQNN at policy updater 608, a new policy π* is accepted, and policy updater 608 and policy provider 607 synchronize their models. In the continuous loop, policy updater 608 and policy provider 607 continue training and providing inference respectively without interfering with inference request 610 to obtain the optimal model split ratio. FM can be split in response to inference request 610 using the policy provided by policy provider 607. As shown in Figure 5, the split FM can be deployed or distributed across edge nodes and cloud systems.

[0081] For example, the method in Figure 5 may further include a step 611 of performing policy federation (PF), which uses a splitting policy learned from policy provider 607 and other model splitting policies learned from other edge nodes 612. The policy federation method can synchronize learned model splitting policies across similar edge nodes to obtain a better and more adaptive global policy. The policy federation method can use inputs 1) through 3): 1) sufficient computing power at an external FL aggregation server to handle policy federation, 2) information about participating edge nodes within a connected MEC region or distributed FMaaS network, and 3) information about the expected synchronization frequency of model splitting policies among said participating edge nodes. Given inputs 1) through 3), the policy federation method can implement the federated learning process to aggregate the learned DQNN weights without interfering with the local inference process by: requesting each participant's policy provider to send a copy of its weights to the FL server, aggregating the weights via a specified FL aggregation policy such as federated averaging, and requesting each participant's policy updater to receive a copy of the globally federated DQNN model and subsequently updating the corresponding policy provider.

[0082] This topic may include the following terms.

[0083] Clause 1. A method for performing a workload using a first artificial intelligence model in a distributed system, the distributed system including a set of first computer systems configured to connect to at least one second computer system in the distributed system, the first artificial intelligence model being configured to receive specific input, process specific input, and provide specific output, the first artificial intelligence model being configured to split into a set of one or more input blocks, intermediate blocks, and a set of one or more output blocks according to a splitting configuration, such that the set of one or more input blocks receives specific input and provides intermediate output, the intermediate blocks receive intermediate output as input and provide another intermediate output, and the set of one or more output blocks receives the other intermediate output as input and provides the specific output; the method includes a splitting method comprising: receiving input using the first artificial intelligence model... The method further includes: an AI model executing a workload request, the workload including the specific input; determining the current resource utilization status in the distributed system; inputting the currently used split configuration of a first AI model and the current resource utilization status into a second AI model, the second AI model being configured to predict a split configuration for the first AI model; receiving an output from the second AI model indicating the current split configuration for the first AI model; the method also includes: splitting the first AI model using the current split configuration; deploying the split first AI model such that input and output blocks can be executed on one or more first computer systems in the group of first computer systems, and that intermediate blocks can be executed on second computer systems in the at least one second computer system; and executing the workload.

[0084] Clause 2. According to the method described in Clause 1, the second artificial intelligence model is a reinforcement learning model, wherein the state of the reinforcement learning model is defined by the current resource utilization status and the split configuration in use, and wherein the policy of the reinforcement learning model defines the action of changing the split configuration in use of the first artificial intelligence model.

[0085] Clause 3. The second artificial intelligence model is trained on multiple first computer systems in the group of first computer systems according to the method of any one of the preceding Clauses 1 to 2, based on federated learning techniques.

[0086] Clause 4. The method described in Clause 3 includes: combining learnable weights of a second artificial intelligence model in a corresponding first computer system to produce a combined second artificial intelligence model, wherein the second artificial intelligence model used in the splitting method is the combined second artificial intelligence model.

[0087] Clause 5. The method according to any one of Clauses 1 to 4 above, wherein the current resource utilization status is defined by at least one of the following: the utilization level of the network resources of the distributed system; the utilization level of the resources of one or more first computer systems; and the utilization level of the resources of the second computer system.

[0088] Clause 6. The distributed system according to any one of Clauses 1 to 5 above is a wireless communication system, wherein the first computer system is a multi-access edge computing (MEC) node and the second computer system is a cloud system.

[0089] Clause 7. The method according to any one of the preceding Clauses 1 to 6 shall be repeated.

[0090] Clause 8. The method according to any one of the preceding Clauses 1 to 7, wherein the first artificial intelligence model is the base model.

[0091] Clause 9. The method according to any one of Clauses 1 to 8 above, wherein the first artificial intelligence model is a deep neural network or transformer, wherein the input block represents the initial network layer, the intermediate block represents the intermediate network layer, and the output block represents the final network layer.

[0092] Clause 10. The method according to any one of the preceding Clauses 1 to 9, wherein the first computer system has a smaller amount of processing resources than the second computer system.

[0093] Clause 11. The method according to any one of Clauses 1 to 10 above, wherein the workload involves at least one of: data analysis, sensor measurement fusion from different sources, image analysis, processing of data streams destined for the cloud or another internal or external workload.

[0094] Clause 12. The method according to any one of Clauses 1 to 11 above, wherein the execution of the first artificial intelligence model comprises: executing every two consecutive blocks of the first artificial intelligence model deployed on different computer systems by at least the following: encoding the output of the first block of the two blocks using an encoding protocol, and sending the encoded output to the first computer system or the second computer system for use as input to the second block of the two blocks.

[0095] Clause 13. The method described in Clause 12, wherein the encoding protocol includes at least one of compression or encryption.

[0096] Clause 14. The second artificial intelligence model according to any one of the preceding Clauses 1 to 13 is a DeepQ Neural Network (DQNN).

[0097] Clause 15. The method according to any one of Clauses 1 to 14 above is repeated, wherein the splitting method further comprises: retraining an instance of the second artificial intelligence model; and, after successful retraining, replacing the second artificial intelligence model with the retrained second artificial intelligence model for further execution of the splitting method.

[0098] Clause 16. The method according to any one of Clauses 1 to 15 above, deployment includes: defining a deployment configuration using current resource utilization status, the deployment configuration indicating a second computer system to execute intermediate blocks, and one or more first computer systems to execute input blocks and output blocks; and using the deployment configuration for deployment.

[0099] Clause 17. Defining the deployment configuration according to the method described in Clause 16 includes: performing a capacity profile of the first computer system and the second computer system to determine the capability of each of the first computer system and the second computer system to execute one or more blocks of the first artificial intelligence model; and defining the deployment configuration based on the capacity profile.

[0100] The computing environment 800 includes examples of an environment for executing and performing at least some of the computer code involved in the inventive methods, such as code 900 for workload orchestration. In addition to block 900, the computing environment 800 also includes, for example, a computer 801, a wide area network (WAN) 802, an end-user equipment (EUD) 803, a remote server 804, a public cloud 805, and a private cloud 806. In this embodiment, the computer 801 includes a processor set 810 (including processing circuitry 820 and a cache 821), a communication structure 811, volatile memory 812, persistent storage device 813 (including an operating system 822 and block 900, as identified above), a peripheral device set 814 (including a user interface (UI) device set 823, a storage device 824, and an Internet of Things (IoT) sensor set 825), and a network module 815. The remote server 804 includes a remote database 830. The public cloud 805 includes a gateway 840, a cloud orchestration module 841, a host physical machine set 842, a virtual machine set 843, and a container set 844.

[0101] Computer 801 may take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or to be developed in the future capable of running programs, accessing networks, or querying databases (such as remote database 830). As is well known to those skilled in the art and depending on the technology, the performance of the computer-implemented method may be distributed among multiple computers and / or across multiple locations. On the other hand, in the description of this computing environment 800, the detailed discussion focuses on a single computer, specifically computer 801, to keep the description as simple as possible. Although not shown in the cloud in Figure 6, computer 801 may be located in the cloud. On the other hand, unless explicitly stated otherwise, computer 801 is not required to be located in the cloud.

[0102] Processor set 810 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 820 may be distributed across multiple packages, for example, multiple coordinated integrated circuit chips. Processing circuitry 820 may implement multiple processor threads and / or multiple processor cores. Cache 821 is memory located within the processor chip package(s) and is typically used for data or code that should be quickly accessible to the threads or cores running on processor set 810. Cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuitry. Alternatively, some or all of the caches in the processor set may be located “off-chip”. In some computing environments, processor set 810 may be designed to work with qubits and perform quantum computing.

[0103] Computer-readable program instructions are typically loaded onto computer 801 to cause a series of operational steps to be executed by processor set 810 of computer 801, thereby implementing a computer-implemented method such that the instructions executed instantiate the method specified in the flowcharts and / or descriptive passages of the computer-implemented method included herein (collectively, the “inventive method”). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 821 and other storage media discussed below. Processor set 810 accesses the program instructions and associated data to control and direct the execution of the inventive method. In computing environment 800, at least some of the instructions for performing the inventive method may be stored in block 900 of persistent memory 813.

[0104] The communication structure 811 is a signal transmission path that allows the various components of the computer 801 to communicate with each other. Typically, this structure consists of switches and electrically conductive paths, such as those forming buses, bridges, physical input / output ports, etc. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.

[0105] Volatile memory 812 is any type of volatile memory now known or to be developed in the future. Examples include dynamically typed random access memory (RAM) or statically typed RAM. Typically, volatile memory 812 is characterized by random access, but this is not required unless explicitly stated. In computer 801, volatile memory 812 is located in a single package and is internal to computer 801; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally relative to computer 801.

[0106] The persistent storage device 813 is any form of non-volatile memory for a computer, now known or to be developed in the future. The non-volatility of this storage device means that the stored data is retained regardless of whether the computer 801 is powered or not, or whether the persistent storage device 813 is directly powered. The persistent memory 813 may be a read-only memory (ROM), but typically at least a portion of the persistent memory allows data to be written, deleted, and rewritten. Some familiar forms of persistent memory include hard disks and solid-state storage devices. The operating system 822 may take many forms, such as various known proprietary operating systems or open-source portable operating system interfaces with a kernel. The code included in block 900 typically includes at least some computer code relating to performing the inventive methods.

[0107] Peripheral device set 814 includes a collection of peripheral devices for computer 801. Data communication connections between peripheral devices and other components of computer 801 can be implemented in various ways, such as Bluetooth connections, near field communication (NFC) connections, connections made via cables (such as Universal Serial Bus (USB) type cables), plug-in connections (e.g., secure digital (SD) cards), connections made via local area communication networks, and even connections made via wide area networks (such as the Internet). In various embodiments, UI device set 823 may include components such as displays, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage device 824 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. Storage device 824 can be persistent and / or volatile. In some embodiments, storage device 824 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments requiring computer 801 to have a large amount of storage (e.g., where computer 801 locally stores and manages a large database), this storage can be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 825 consists of sensors that can be used in IoT applications. For example, one sensor could be a thermometer, and another sensor could be a motion detector.

[0108] Network module 815 is a collection of computer software, hardware, and firmware that allows computer 801 to communicate with other computers via WAN 802. Network module 815 may include hardware (such as a modem or Wi-Fi transceiver), software for packetizing and / or de-packetizing data for transmission over the communication network, and web browser software for transmitting data over the Internet. In some embodiments, the network control and network forwarding functions of network module 815 are performed on the same physical hardware device. In other embodiments (e.g., embodiments using software-defined networking (SDN), the control and forwarding functions of network module 815 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the invention can typically be downloaded to computer 801 from an external computer or external storage device via a network adapter card or network interface included in network module 815.

[0109] A WAN 802 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any technology now known or to be developed in the future for transmitting computer data. In some embodiments, a WAN 802 may be replaced by and / or supplemented by a local area network (LAN) (such as a Wi-Fi network) designed to transmit data between devices located in a local area. WANs and / or LANs typically include computer hardware such as copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0110] End User Equipment (EUD) 803 is any computer system used and controlled by an end user (e.g., an enterprise customer operating computer 801) and can take any form as discussed above regarding computer 801. EUD 803 typically receives helpful and useful data from the operation of computer 801. For example, in the hypothetical case where computer 801 is designed to provide recommendations to the end user, these recommendations would typically communicate from computer 801's network module 815 to EUD 803 via WAN 802. In this way, EUD 803 can display or otherwise present recommendations to the end user. In some embodiments, EUD 803 can be a client device, such as a thin client, a thick client, a mainframe computer, a desktop computer, etc.

[0111] Remote server 804 is any computer system that provides at least some data and / or functionality to computer 801. Remote server 804 can be controlled and used by the same entity operating computer 801. Remote server 804 represents a machine that collects and stores helpful and useful data for use by other computers, such as computer 801. For example, in the hypothetical scenario where computer 801 is designed and programmed to provide recommendations based on historical data, that historical data could be provided to computer 801 from a remote database 830 of remote server 804.

[0112] Public cloud 805 is any computer system available to multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities (especially data storage (cloud storage) and computing power) without direct, active management by the user. Cloud computing typically leverages resource sharing to achieve consistency and economies of scale. Direct and active management of the computing resources of public cloud 805 is performed by the computer hardware and / or software of cloud orchestration module 841. The computing resources provided by public cloud 805 are typically implemented by virtual computing environments running on individual computers constituting host physical set 842, which is the entire set of physical computers in and / or available to public cloud 805. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 843 and / or containers from container set 844. It is to be understood that these VCEs can be stored as images and can be transferred between and between individual physical machine hosts, either as images or after instantiation of the VCE. The cloud orchestration module 841 manages the transfer and storage of images, deploys new VCE instantiations, and manages the instantiation of active VCE deployments. The gateway 840 is a collection of computer software, hardware, and firmware that allows the public cloud 805 to communicate via WAN 802.

[0113] Now, some further explanation of Virtualized Computing Environments (VCEs) will be provided. A VCE can be stored as an "image." New active instances of a VCE can be instantiated from an image. Two familiar types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows the existence of multiple isolated user-space instances (called containers). These isolated user-space instances typically behave like a real computer from the viewpoint of the programs running within them. Computer programs running on a regular operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running inside a container can only use the contents of the container and the devices allocated to the container; this is a feature known as containerization.

[0114] Private cloud 806 is similar to public cloud 805, except that computing resources are only available for use by a single enterprise. While private cloud 806 is depicted as communicating with WAN 802, in other embodiments, private cloud can be completely disconnected from the internet and accessed only via a local / private network. Hybrid cloud is a combination of multiple clouds of different types (e.g., private cloud, community cloud, or public cloud types), typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple component clouds. In this embodiment, both public cloud 805 and private cloud 806 are part of a larger hybrid cloud.

[0115] It should be understood that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings recorded herein is not limited to a cloud computing environment. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed hereafter.

[0116] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.

[0117] The features are as follows:

[0118] On-demand self-service: Cloud consumers can unilaterally and automatically supply computing power, such as server time and network storage, on demand, without requiring manual interaction with the service provider.

[0119] Extensive network access: Capabilities can be accessed through network availability and through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms, such as mobile phones, laptops, and PDAs.

[0120] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. There exists a perception of location independence, meaning that consumers typically do not have control or knowledge of the exact location of the resources they are provided with, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0121] Rapid elasticity: Capacity can be configured quickly and flexibly, and in some cases automatically, to expand rapidly and contract rapidly. To consumers, the capacity available for supply often appears unlimited and can be purchased in any quantity at any time.

[0122] Measurable services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at several levels of abstraction appropriate to service types, such as storage, processing, bandwidth, and active user accounts. Resource usage can be monitored, controlled, and reported, providing transparency to both the providers and consumers of the services used.

[0123] The service model is as follows:

[0124] Software as a Service (SaaS): This provides consumers with the ability to use the provider's applications running on cloud infrastructure. These applications can be accessed from various client devices through thin client interfaces, such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, storage devices, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0125] Platform as a Service (PaaS): This provides consumers with the ability to deploy consumer-created or acquired applications, built using programming languages ​​and tools supported by the provider, onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage devices, but they have control over the configuration of deployed applications and potential application hosting environments.

[0126] Infrastructure as a Service (IaaS): This provides consumers with the capability to supply processing, storage, networking, and other basic computing resources, where consumers can deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but have control over the operating system, storage devices, deployed applications, and possibly limited control over selected network components (e.g., host firewalls).

[0127] The deployment model is as follows:

[0128] Private cloud: The cloud infrastructure is operated solely by the organization. It can be managed by the organization or a third party, and can exist either on-premises or off-premises.

[0129] Community cloud: Cloud infrastructure shared by multiple organizations and supporting specific communities with shared concerns (such as mission, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can exist inside or outside the deployment.

[0130] Public cloud: Cloud infrastructure available to the public or large industry groups and owned by organizations that sell cloud services.

[0131] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain separate entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0132] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure of a network of interconnected nodes.

[0133] Referring now to Figure 7, an illustrative cloud computing environment 1050 is depicted. As shown, the cloud computing environment 1050 includes one or more cloud computing nodes 1010, with local computing devices that can communicate with these nodes and are used by cloud consumers, such as, for example, a personal digital assistant (PDA) or cellular phone 1054A, a desktop computer 1054B, a laptop computer 1054C, and / or an automotive computer system 54N. The nodes 1010 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as a private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof, as described above. This allows the cloud computing environment 1050 to provision infrastructure, platform, and / or software-as-a-service for which cloud consumers do not need to maintain resources on their local computing devices. It should be understood that the types of computing devices 1054A-N shown in Figure 7 are intended to be illustrative only, and that computing node 1010 and cloud computing environment 1050 can communicate with any type of computerized device across any type of network and / or network-addressable connectivity (e.g., using a web browser).

[0134] Referring now to Figure 8, a set of functional abstraction layers provided by the cloud computing environment 1050 (Figure 7) is shown. It should be understood beforehand that the components, layers, and functions shown in Figure 8 are intended to be illustrative only, and embodiments of the invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0135] The hardware and software layer 1060 includes hardware and software components. Examples of hardware components include: a mainframe 1061; a server 1062 based on a RISC (Reduced Instruction Set Computer) architecture; a server 1063; a blade server 1064; a storage device 1065; and network and networking components 1066. In some embodiments, software components include network application server software 1067 and database software 1068.

[0136] The virtualization layer 1070 provides an abstraction layer from which examples of the following virtual entities can be provided: virtual server 1071; virtual storage device 1072; virtual network 1073, including virtual private network; virtual application and operating system 1074; and virtual client 1075.

[0137] In one example, management layer 1080 can provide the functionality described below. Resource provisioning 1081 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 1082 provides cost tracking as resources are used within the cloud computing environment, and bills or invoices for consuming these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 1083 provides consumers and system administrators with access to the cloud computing environment. Service level management 1084 provides cloud resource allocation and management to ensure that required service levels are met. Service level agreement (SLA) planning and fulfillment 1085 provides pre-scheduling and procurement of cloud resources for anticipated future needs in accordance with the SLA.

[0138] Workload layer 1090 provides examples of functionalities that can leverage a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: map creation and navigation 1091; software development and lifecycle management 1092; virtual classroom education delivery 1093; data analysis and processing 1094; transaction processing 1095; and an AI model inference engine (AIIE) 1096 that executes the main block of the first artificial intelligence model according to this topic.

[0139] Various aspects of this disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). With respect to any flowchart, depending on the technology involved, operations may be performed in a different order than shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, concurrently, or in a manner that at least partially overlaps in time.

[0140] Computer Program Product Embodiment (“CPP Embodiment” or “CPP”) is a term used in this disclosure to describe a group of one or more storage media (also referred to as “media”) collectively included in a group of one or more storage devices, which collectively include machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A “storage device” is any tangible device capable of holding and storing instructions used by a computer processor. Without limitation, a computer-readable storage medium can be an electrical storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices including these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), optical disc read-only memory (CD-ROM), digital versatile optical disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punched cards or pits / planes formed on the main surface of the disk), or any suitable combination of the foregoing. As used in this disclosure, computer-readable storage media are not to be construed as storing transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses transmitted through fiber optic cables, electrical signals transmitted through lines, and / or other transmission media. As those skilled in the art will understand, during normal operation of a storage device (such as during access, defragmentation, or garbage collection), data typically moves at some random points in time, but this does not make the storage device transient, because the data is not transient when it is stored.

Claims

1. A method for performing a workload using a first artificial intelligence model in a distributed system, the distributed system comprising a set of first computer systems configured to connect to at least one second computer system of the distributed system, the first artificial intelligence model configured to receive specific input, process the specific input, and provide specific output, the first artificial intelligence model configured to be split into a set of one or more input blocks, intermediate blocks, and a set of one or more output blocks according to a splitting configuration, such that the set of one or more input blocks receives the specific input and provides intermediate output, the intermediate blocks receive the intermediate output as input and provide another intermediate output, and the set of one or more output blocks receives the other intermediate output as input and provides the specific output; The method includes a splitting method, the splitting method comprising: receiving a request to execute a workload using a first artificial intelligence model, the workload including the specific input; determining the current resource utilization status in the distributed system; inputting the currently used splitting configuration of the first artificial intelligence model and the current resource utilization status into a second artificial intelligence model, the second artificial intelligence model being configured to predict a splitting configuration for the first artificial intelligence model; receiving an output from the second artificial intelligence model, the output indicating the current splitting configuration for the first artificial intelligence model; the method further comprising: splitting the first artificial intelligence model using the current splitting configuration; deploying the split first artificial intelligence model such that input blocks and output blocks can be executed on one or more first computer systems in the set of first computer systems, and such that intermediate blocks can be executed on second computer systems in at least one second computer system; and executing the workload.

2. The method according to claim 1, wherein the second artificial intelligence model is a reinforcement learning model, wherein the state of the reinforcement learning model is defined by the current resource utilization status and the split configuration in use, and wherein the policy of the reinforcement learning model defines the action of changing the split configuration in use of the first artificial intelligence model.

3. The method according to any one of the preceding claims, wherein the second artificial intelligence model is trained on a plurality of first computer systems in the set of first computer systems using federated learning techniques.

4. The method according to claim 3, wherein the training comprises: The learnable weights of the second artificial intelligence model in the corresponding first computer system are combined to produce a combined second artificial intelligence model, wherein the second artificial intelligence model used in the splitting method is the combined second artificial intelligence model.

5. The method according to any one of the preceding claims, wherein the current resource utilization status is defined by at least one of the following: the utilization level of network resources of the distributed system; the utilization level of resources of one or more first computer systems in the first computer system; and the utilization level of resources of the second computer system.

6. The method according to any one of the preceding claims, wherein the distributed system is a wireless communication system, wherein the first computer system is a multi-access edge computing (MEC) node, and the second computer system is a cloud system.

7. The method according to any one of the preceding claims, wherein the method is performed repeatedly.

8. The method according to any one of the preceding claims, wherein the first artificial intelligence model is a base model.

9. The method according to any one of the preceding claims, wherein the first artificial intelligence model is a deep neural network or a transformer, wherein the input block represents the initial network layer, the intermediate block represents the intermediate network layer, and the output block represents the final network layer.

10. The method according to any one of the preceding claims, wherein the first computer system has a smaller amount of processing resources than the second computer system.

11. The method according to any one of the preceding claims, wherein the workload involves at least one of: data analysis, sensor measurement fusion from different sources, image analysis, and processing of data streams destined for the cloud or another internal or external workload.

12. The method according to any one of the preceding claims, wherein execution of the first artificial intelligence model comprises: Each pair of consecutive blocks of a first artificial intelligence model deployed on different computer systems is executed by at least the following: encoding the output of the first block of the two blocks using an encoding protocol, and sending the encoded output to either the first or the second computer system so as to be used as input to the second block of the two blocks.

13. The method of claim 12, wherein the encoding protocol includes at least one of compression or encryption.

14. The method according to any one of the preceding claims, wherein the second artificial intelligence model is a deep Q-neural network (DQNN).

15. The method according to any one of the preceding claims, wherein the method is repeatedly performed, and wherein the splitting method further comprises: An example of retraining a second artificial intelligence model; And after successful retraining, the retrained second artificial intelligence model is used to replace the second artificial intelligence model for further execution of the splitting method.

16. The method according to any one of the preceding claims, wherein the deployment comprises: The deployment configuration is defined using the current resource utilization status, and the deployment configuration indicates the second computer system to execute the intermediate block and one or more first computer systems to execute the input and output blocks; And use the deployment configuration for the deployment.

17. The method of claim 16, wherein defining the deployment configuration includes: Perform a capacity profile of the first computer system and the second computer system to determine the ability of each of the first computer system and the second computer system to execute one or more blocks of the first artificial intelligence model; and define the deployment configuration based on the capacity profile.

18. A computer program product comprising a computer-readable storage medium having computer-readable program code implemented therewith, the computer-readable program code being configured to implement the method according to any one of the preceding claims.

19. A computer system for performing a workload using a first artificial intelligence model in a distributed system, the distributed system comprising a set of first computer systems configured to connect to at least one second computer system of the distributed system, the first artificial intelligence model configured to receive specific input, process the specific input, and provide a specific output, the first artificial intelligence model configured to be split into a set of one or more input blocks, intermediate blocks, and a set of one or more output blocks according to a splitting configuration, such that the set of one or more input blocks receives the specific input and provides an intermediate output, the intermediate blocks receive the intermediate output as input and provide another intermediate output, and the set of one or more output blocks receives the other intermediate output as input and provides the specific output; the computer system is configured to: The process includes: receiving a request to execute a workload using a first artificial intelligence model, the workload including the specific input; determining the current resource utilization status in the distributed system; inputting the currently used split configuration and the current resource utilization status of the first artificial intelligence model into a second artificial intelligence model, the second artificial intelligence model being configured to predict a split configuration for the first artificial intelligence model; receiving an output from the second artificial intelligence model, the output indicating the current split configuration for the first artificial intelligence model; splitting the first artificial intelligence model using the current split configuration; deploying the split first artificial intelligence model such that input blocks and output blocks can be executed on one or more of the first computer systems in the set of first computer systems, and that intermediate blocks can be executed on the second computer systems in at least one of the second computer systems; and executing the workload.

20. The computer system according to claim 19, wherein the computer system is the first computer system in the group of first computer systems.