A category-based 6g network multi-dimensional resource ai model dynamic deployment optimization method
By employing a joint optimization method based on category theory and the Lyapunov drift penalty framework, the cross-layer consistency and multi-dimensional resource coupling problems of AI model deployment in 6G networks are solved, achieving low latency and stability optimization in dynamic environments. This method is suitable for large-scale end-edge-cloud collaborative inference and IoT intelligent services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-02-28
- Publication Date
- 2026-05-29
AI Technical Summary
The static or semi-static deployment schemes of AI models in existing 6G networks are difficult to adapt to the dynamic changes in task arrival and network status, resulting in resource imbalance and increased queuing latency. Dynamic deployment schemes, on the other hand, suffer from insufficient cross-layer consistency description, difficulty in handling multi-dimensional resource coupling, and online optimization challenges in dynamic environments, and cannot meet long-term service quality requirements.
A unified modeling framework based on category theory is constructed. Cross-layer consistency dependencies are formally characterized by functors. By combining the Lyapunov drift penalty framework and genetic algorithm, joint optimization of task scheduling and model deployment is achieved, reducing long-term average end-to-end latency and ensuring queue stability.
It achieves unified modeling and joint optimization of task scheduling and dynamic deployment of AI models under multi-dimensional resource constraints, reduces long-term average end-to-end latency, avoids node resource overruns and load imbalance, and ensures system stability and performance.
Smart Images

Figure CN122120342A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent collaborative optimization technology, specifically involving a dynamic deployment optimization method for AI models of multi-dimensional resources in 6G networks based on category theory. Background Technology
[0002] As 6G networks evolve towards "integrated sensing, intelligence, computing, and storage," network service models for intelligent businesses are gradually expanding from traditional communication connections to edge-cloud collaborative intelligent inference and decision-making services. A large number of sensing, interaction, and generative tasks generated by intelligent terminals are characterized by strong randomness in arrival, large data scale, strict latency constraints, and differentiated inference accuracy requirements. To meet these business needs, the industry generally adopts a collaborative architecture that deploys computing resources and AI models at the edge, provides elastic computing power and model libraries in the cloud, and handles lightweight inference and data preprocessing at the edge. This low-latency delivery of inference services is achieved through task offloading and model deployment.
[0003] In existing technologies, there are generally two typical solutions for edge-cloud collaborative reasoning:
[0004] (1) Static or semi-static model deployment + task scheduling scheme: The model is pre-deployed on some edge or cloud nodes, and only task routing or unloading decisions are made at runtime. This type of scheme is simple to implement, but it is difficult to adapt to the dynamic changes in task arrival and network status. It is easy to have resource imbalance problems such as some nodes being congested while other nodes are idle, which leads to increased queuing latency and decreased service quality.
[0005] (2) Dynamic Deployment + Online Scheduling Scheme: The model is dynamically deployed, scaled up, shrunk, or migrated according to the task load and network status, while simultaneously performing task scheduling and resource allocation to improve resource utilization and service performance. However, this type of problem usually involves multi-variable joint optimization such as task-model selection, model-node placement, and multi-dimensional resource allocation. It has the characteristics of discrete decision-making coupled with continuous resources, strong cross-layer dependencies, multiple constraint dimensions, and complex and non-convex feasible regions, making it difficult to directly obtain the global optimal solution and run stably in a real-time system.
[0006] In 6G AIaaS inference service scenarios, in addition to meeting computational resource constraints, the system also needs to meet multi-dimensional resource constraints such as storage capacity, memory capacity, and bandwidth usage. Furthermore, due to the randomness of task arrival and link / computing power status, the system typically experiences dynamic queue evolution caused by task queuing. Existing methods utilize queuing theory, Lyapunov online optimization, and reinforcement learning to achieve online decision-making, but they still generally suffer from the following shortcomings:
[0007] Insufficient description of cross-layer consistency: Existing methods mostly describe task scheduling and model deployment separately using empirical rules or hierarchical modeling, lacking a unified formal characterization of the cross-layer mapping relationship of "task-model-physical node", which easily leads to inconsistencies or duplicate constraints during joint decision-making, increasing the complexity of modeling and solving.
[0008] Multidimensional resource coupling is difficult to handle: When multiple constraints such as computing, memory, storage, and bandwidth exist at the same time, simple greedy or local patching strategies are prone to frequent infeasible solutions or performance fluctuations, making it difficult to balance end-to-end latency and long-term stability.
[0009] Online optimization in dynamic environments is difficult to implement: As task arrival and network state change over time, offline solutions or long-term iterative global optimizations cannot meet real-time requirements; and existing online algorithms, if lacking effective structured decomposition, often have excessively high computational complexity or difficulty in guaranteeing convergence and stability.
[0010] Queuing latency and service stability are not adequately balanced: Optimizing only instantaneous transmission or computation latency may lead to continuous queue growth, resulting in a significant increase in queuing latency or even system instability, which cannot meet long-term service quality requirements.
[0011] Therefore, there is an urgent need for a dynamic deployment optimization method for AI models that can perform unified modeling and joint optimization of task scheduling and dynamic deployment of AI models under multi-dimensional resource constraints, effectively reduce long-term average end-to-end latency while ensuring queue stability, and has an engineering-featured solution process and controllable complexity. Summary of the Invention
[0012] The purpose of this invention is to provide a dynamic deployment optimization method for AI models of multi-dimensional resources in 6G networks based on category theory, so as to solve the problems mentioned in the background art.
[0013] The objective of this invention is achieved as follows: a dynamic deployment optimization method for multi-dimensional resources in 6G networks based on category theory using AI models, characterized by the following steps:
[0014] Step S1: Construct a categorized unified modeling framework to formally characterize the cross-layer consistency dependencies of end-to-end AI inference services through functors;
[0015] Step S2: Under the conditions of dynamic task arrival and time-varying network state, establish a joint optimization model with the goal of minimizing long-term average end-to-end latency, while being constrained by multi-dimensional resources and quality of service.
[0016] Step S3: Construct a two-stage optimization algorithm based on the Lyapunov drift penalty framework to transform the long-term stochastic optimization problem into a time-slot-by-time online decision problem;
[0017] Step S4: Based on two-stage alternating optimization and genetic algorithm, alternately solve the task-model matching and model-node deployment sub-problems to obtain the dynamic deployment and task scheduling results of AI model.
[0018] Preferably, in step S1, a categorized unified modeling framework is constructed to formally characterize the cross-layer consistency dependencies of end-to-end AI inference services using functors, specifically as follows:
[0019] Step S1-1: Construct a 6G network architecture that integrates end-edge-cloud collaboration. Based on this 6G network architecture, category theory modeling is introduced, specifically as follows:
[0020] The 6G network architecture includes a physical layer, a model layer, and a task layer. The physical layer contains heterogeneous computing nodes from terminals, edge, and cloud. The model layer contains a collection of AI models. The task layer contains intelligent inference task requests generated by service terminals.
[0021] Abstracting physical computing nodes into the physical layer category Abstracting AI models into model categories Abstracting intelligent tasks into task categories ,in, The path categories are generated by the network topology. and It belongs to the discrete category;
[0022] Physical layer scope Abstracting all physical computing nodes into the physical layer category A collection of objects, expressed as:
[0023] ;
[0024] in, Indicates the first One physical computing node;
[0025] Model category Abstracting AI models into model categories A collection of objects, expressed as:
[0026] ;
[0027] in, Indicates the first An AI model;
[0028] Task scope Abstracting intelligent tasks into task categories A collection of objects, expressed as:
[0029] ;
[0030] in, express One intelligent reasoning task instance / request;
[0031] Step S1-2: Construct the matching functor from task to model and the deployment functor from model to physical node, and form an end-to-end consistent mapping of task-model-node through the composite functor;
[0032] Step S1-3: Introduce discrete-time indexes to describe the dynamic characteristics of task scheduling and resource allocation in the time dimension.
[0033] Preferably, in steps S1-2, the matching functor from task to model and the deployment functor from model to physical node are constructed, and an end-to-end consistent mapping of task-model-node is formed through composite functors, specifically as follows:
[0034] Define functor Hanzi Mapping task objects to model objects; under the influence of the functor, for each task object in the task scope... Hanzi Map it to a model object ,Right now ;
[0035] To ensure the effectiveness of the mapping, task-model matching must satisfy the following constraints:
[0036] ;
[0037] set up Indicates the model's inference accuracy. If we represent the minimum precision requirement for the task, then:
[0038] ;
[0039] In dynamic network environments, the mapping from task to model is re-determined at different time slots depending on the system state, that is, the corresponding mapping functor is solved in each discrete time slot. To adapt to task arrival, resource changes and node load fluctuations;
[0040] Define deployment functor Mapping model objects to physical node objects: For any model ,Model Deployed and run on physical nodes The expression is:
[0041] ;
[0042] Set physical nodes At any time / time slot The available computing, memory, storage, and bandwidth resources are respectively , , and The model deployment must satisfy the multi-dimensional resource capacity constraints of the nodes:
[0043] ;
[0044] ;
[0045] ;
[0046] in, , and Representing the model respectively At any moment The computational, memory, and storage resources consumed during operation; to avoid duplicate deployment of models, each model is further constrained to be deployed on only one physical node in any given time slot as follows:
[0047] ;
[0048] To uniformly characterize the mapping relationship between tasks, models, and physical nodes, a composite functor is introduced:
[0049] ;
[0050] Task After model selection and deployment decisions, the data is ultimately allocated to physical nodes. Execution, the expression is:
[0051] .
[0052] Preferably, in steps S1-3, a discrete-time index is introduced to describe the dynamic characteristics of task scheduling and resource allocation in the time dimension, specifically as follows:
[0053] Introducing the category of discrete-time index The object is discrete time slots ;
[0054] In the time slot Using scheduling indicator variables This indicates the temporal evolution relationship between time slots;
[0055] when The condition indicates that the ternary relation is true; otherwise, it is not true.
[0056] Define task At the node The calculated load function is This indicates that the node is in the time slot. The equivalent computational workload provided by the task, according to the principle of load conservation, can be formalized as follows:
[0057] ;
[0058] in, This is the logical computational complexity factor when the task is processed by the model. Describe the task at time The amount of input data;
[0059] Based on the computational load, further characterize the memory and storage resource consumption of the task on the node:
[0060] ;
[0061] in, For model parameter memory usage, The activation memory requirements generated by the task during model inference;
[0062] ;
[0063] in, Storage size for the model, The required data storage size for the task.
[0064] Preferably, in step S2, under the conditions of dynamic task arrival and time-varying network state, a joint optimization model is established with the objective of minimizing long-term average end-to-end latency, while being constrained by multi-dimensional resources and quality of service. Specifically:
[0065] Step S2-1: Determine the arrival of dynamic tasks and network status:
[0066] Task In the time slot Transmit from source node to target node The transmission delay is defined as:
[0067] ;
[0068] in, For a moment bandwidth rate, The distance from the source node. For link transmission speed; Describe the task at time The amount of input data;
[0069] When the rate at which tasks arrive exceeds the instantaneous processing capacity of a node, queuing will occur, introducing queuing delay. Let... Represents a node In the time slot Given the queue length, the evolution of the queue state between adjacent time slots is represented as:
[0070] ;
[0071] in, The time slot length;
[0072] Under stable system operation conditions, the queue length satisfies:
[0073] ;
[0074] According to Little's Law, in a steady state, the task... At the node The average queuing delay is approximately expressed as:
[0075] ;
[0076] Combining transmission, computation, and queuing delays, the end-to-end delay of a task is expressed as:
[0077] ;
[0078] Since different tasks have different latency requirements, the end-to-end latency constraint for a task is as follows:
[0079] ;
[0080] Let the time window be The set of tasks arriving within this window is denoted as The total number of tasks is The average end-to-end delay of the system within the time window is defined as:
[0081] ;
[0082] in, Indicates task In its corresponding physical node End-to-end latency during execution;
[0083] Step S2-2: Establish a joint optimization problem with the objective of minimizing the long-term average end-to-end delay.
[0084] Preferably, step S2-2 establishes a joint optimization problem with the objective of minimizing the long-term average end-to-end delay, specifically as follows:
[0085] ;
[0086] ;
[0087] in, Indicates a time range The upper limit of the average end-to-end delay as it approaches infinity.
[0088] Preferably, in step S3, a two-stage optimization algorithm based on the Lyapunov drift penalty framework is constructed to transform the long-term stochastic optimization problem into a time-slot-by-time online decision problem, specifically as follows:
[0089] set up Represents physical nodes In the time slot The virtual queue length is used to characterize the computational load backlog on the node side. A quadratic Lyapunov function is defined as follows:
[0090] ;
[0091] In the time slot Inside, node Its service capacity is determined by the computing resources allocated to it. Let its equivalent service volume be:
[0092] ;
[0093] in, Represents a node In the time slot For the model Equivalent computing power allocated; Indicates the duration of the time slot;
[0094] The node queue update equation is expressed as:
[0095] ;
[0096] in, For time slots The computational load of newly arriving nodes; Represents physical nodes In the time slot The amount of virtual queue backlog;
[0097] According to Lyapunov drift analysis, the upper bound of the drift is obtained as follows:
[0098] ;
[0099] in, Represents the conditional expectation operator; Describes a quadratic Lyapunov function; Indicates task After being assigned to a node End-to-end latency during execution; Represents physical nodes In the time slot The amount of virtual queue backlog; This represents a constant related to the system size;
[0100] To optimize system performance while ensuring queue stability, a drift penalty term is introduced, given control parameters. Now, consider minimizing the following single-slot objective function:
[0101] ;
[0102] The original joint optimization problem is equivalently transformed into a time-slot-by-time static optimization subproblem. :
[0103] ;
[0104] ;
[0105] ;
[0106] in, This ensures that each task uniquely matches a single AI model, and This ensures that the inference accuracy of the selected model meets the task-level QoAIS requirements; in a dynamic system environment, functors Re-optimize within each time slot to adapt to time-varying task arrival processes and resource states.
[0107] set up , , and They represent time slots respectively Internal physical nodes Available computing, memory, storage, and bandwidth resources; model deployment must meet the following resource capacity constraints:
[0108] ;
[0109] ;
[0110] ;
[0111] Furthermore, each model can be deployed on at most one physical node per time slot:
[0112] ;
[0113] in, This indicates that resource feasibility is guaranteed. This indicates that the deployment must be unique.
[0114] To meet the QoS requirements at the task level, the following latency constraints are introduced:
[0115] .
[0116] Preferably, in step S4, based on two-stage alternating optimization and genetic algorithm, the task-model matching and model-node deployment sub-problems are solved alternately to obtain the dynamic deployment and task scheduling results of the AI model, specifically as follows:
[0117] Step S4-1: Use a two-stage alternating optimization approach to perform time-slot static optimization subproblems. It can be broken down into two alternating subproblems, as follows:
[0118] In the time slot Using scheduling indicator variables To represent the ternary mapping relationship, and to facilitate the two-stage alternating solution, two types of discrete decision variables are introduced:
[0119] ;
[0120] Indicates task Select Model and satisfy ;
[0121] Representation Model Whether to deploy on a node and satisfy ;
[0122] The time-slot static optimization subproblem Break it down into two alternating subproblems:
[0123] Subproblem P2-F is the task matching phase: given a deployment decision Optimization This mainly affects transmission delay, accuracy constraints, and some queuing estimations;
[0124] Subproblem P2-G is the model deployment phase: given task matching Optimization It mainly affects computation latency and the feasibility of multi-dimensional resource constraints on nodes;
[0125] Step S4-2: Use GA to efficiently search for approximate optimal solutions in the discrete solution space.
[0126] Preferably, in step S4, the GA algorithm is used to efficiently search for an approximate optimal solution in the discrete solution space, specifically as follows:
[0127] The GA algorithm consists of two iterative layers: an outer layer for traffic optimization and an inner layer for genetic algorithm iteration.
[0128] Traffic optimization iteration: alternating fixed routes optimization Then fix optimization Continue until the maximum number of iterations is reached or the objective convergence is achieved;
[0129] Genetic Algorithm Iteration: For the discrete combination search of each subproblem, a genetic algorithm is used to approximate the solution, so as to reduce the search complexity of the high-dimensional integer space;
[0130] Task matching phase: Task matching initialization :
[0131] The initial task matching rules are as follows:
[0132] ;
[0133] in, ;
[0134] Model deployment initialization The initialization model deployment rules are as follows:
[0135] ;
[0136] in, Representation Model If deployed to nodes At that time, there are penalties for exceeding the limits of computing, memory, and storage resources;
[0137] Task matching subproblem:
[0138] Chromosomes are used in sets of length based on the number of tasks. The integer encoding is ;
[0139] in Indicates task Select Model ,Right now ;
[0140] The fitness function employs an exterior penalty function, which consists of a slot-by-slot end-to-end delay estimate and a queue penalty. Penalties are imposed on individuals that violate the delay constraints, as shown in the following formula:
[0141] ;
[0142] in, Indicates that in a given Under the given conditions, the estimated end-to-end latency of the task; V is the weight parameter of the Lyapunov optimization; It is the precision default penalty coefficient. It is the time-based penalty coefficient for breach of contract;
[0143] Model deployment subproblem: Chromosomes are used with a length equal to the number of models. The integer encoding is ;
[0144] in, Representation Model Deployed on nodes ,Right now ;
[0145] The fitness function primarily calculates the relevant latency and resource overrun penalties, employing an exterior penalty function. The fitness formula is as follows:
[0146] ;
[0147] in, For the current and The collective resource consumption of nodes is determined jointly; It is the resource default penalty coefficient;
[0148] Obtain the initial solution Then, proceed to the outer AO iteration, the first... The next iteration is executed:
[0149] fixed The task matching subproblem was solved using GA to obtain... ;
[0150] fixed The model deployment subproblem was solved using GA to obtain... ;
[0151] If the objective function decreases by less than the threshold or the maximum number of iterations is reached, then stop.
[0152] The computational scale of the GA algorithm depends on the number of iterations of the outer AO. And the number of generations in the inner genetic algorithm. With population size The overall time complexity is expressed as:
[0153] ;
[0154] The main complexity of the task matching phase is: The main complexity of the model deployment phase is .
[0155] Compared with the prior art, the present invention has the following improvements and advantages:
[0156] 1. By using category theory-based unified modeling and functor composite mechanism, we can achieve cross-layer consistent description of task scheduling and model deployment, reduce inconsistencies and redundant constraints caused by hierarchical modeling, and improve the structure and interpretability of joint decision-making. Under the constraints of multi-dimensional resources such as computing, memory, storage and bandwidth, we can achieve dynamic adaptive optimization, effectively avoid node resource over-limit and load imbalance, and reduce congestion and queuing latency.
[0157] 2. By combining the Lyapunov drift penalty online optimization framework, queue stability can be guaranteed in a stochastic environment with fluctuating task arrival and network resources, and the long-term average end-to-end latency can be continuously reduced. The discrete combinatorial decision is solved with low complexity through a two-stage alternating genetic optimization algorithm, which takes into account both real-time performance and performance. It is suitable for the application needs of large-scale edge-cloud collaborative inference, Internet of Things (IoT) smart services and 6G smart wireless networks. Attached Figure Description
[0158] Figure 1 This is a schematic diagram of the scenario structure of the present invention.
[0159] Figure 2 The flowchart shows the online optimization and two-stage alternating genetic optimization algorithms.
[0160] Figure 3 This is a schematic diagram showing the average queue backlog change results of the AO-GA algorithm over multiple time slots.
[0161] Figure 4 This is a schematic diagram showing the comparison of the average latency of different algorithms as the number of iterations changes within a single time slot.
[0162] Figure 5 This is a schematic diagram showing the comparison of the average latency of different algorithms as the amount of data changes within a single time slot.
[0163] Figure 6 This diagram illustrates the comparison of the average time delay of different algorithms as they change over multiple consecutive time slots. Detailed Implementation
[0164] The invention will be further summarized below with reference to the accompanying drawings.
[0165] A dynamic deployment optimization method for multi-dimensional resources in 6G networks based on category theory using AI models, comprising the following steps:
[0166] Constructing a 6G network architecture that integrates edge, cloud, and endpoint collaboration, and introducing category theory modeling based on this 6G network architecture:
[0167] This invention comprises three main categories: the task category, the AI model category, and the physical layer category; the physical layer category: let the physical layer category be... Its object collection is Each object This refers to a physical computing node in the system, including terminal nodes, edge nodes, and cloud nodes. (In the scope...) In this context, morphisms are used to characterize the reachability and communication relationships between physical nodes. For any two objects... If a morphism exists , then it represents a node With nodes There exists a reachable data transmission path; morphic This describes the transmission direction of data flow from the source node to the destination node and its communicability. (Physical layer scope) Modeled as a path category generated by network topology, its basic morphisms are generated by physical links, and its general morphisms are composed of a finite composite of multiple links, corresponding to multi-hop communication paths; identity morphisms Represents a node This includes local computation or scenarios where cross-node transfer is not required. The aforementioned set of morphisms collectively characterizes the feasible path space for task and model migration and communication between physical nodes.
[0168] AI Model Layer Category: Let the AI model category be... Its object collection is Each object This represents a deployable AI model entity, corresponding to a model instance or model type with a defined structure, parameter size, and inference performance characteristics. It defines the model scope. For the discrete category, for any two distinct model objects There is no source arrive Non-trivial morphisms; each model object contains only its identity morphisms. .
[0169] Task scope: Let the task scope be... Its object collection is Each object This represents a single instance of an intelligent inference task or a task request, corresponding to a computational task with a defined service type, input size, and QoS requirements. The task scope is also defined. For the discrete category, any two distinct task objects There are no nontrivial morphisms between them; each task object contains only its identity morphism. .
[0170] Functor mapping from task to model:
[0171] Define functor It operates at the object level of two categories. In the functor... Under its influence, for each task object within the task scope Hanzi Map it to a model object ,Right now ; indicates a task Assigned to model Execute; because and For the discrete category, its morphisms only include identity morphisms, and functors The mapping in the morphological layer naturally maintains identity, and the main focus is on the task characterized by its object mapping—model matching decision.
[0172] To ensure the effectiveness of the mapping, the task matching process must meet the following constraints:
[0173] (1)
[0174] (2)
[0175] in, This ensures that each task uniquely matches a single AI model, and This ensures that the inference accuracy of the selected model meets the task-level QoAIS requirements; in a dynamic system environment, functors Re-optimize within each time slot to adapt to time-varying task arrival processes and resource states.
[0176] In dynamic network environments, the mapping from task to model is re-determined at different time slots depending on the system state; that is, a corresponding mapping functor is solved in each time slot. This adapts to task arrivals, resource changes, and node load fluctuations. Through this object-layer functor mapping mechanism, the system achieves structured alignment between the task layer and the model layer.
[0177] Model-to-Physical Node Functor Mapping: After completing the task-to-model mapping, to further realize the computational execution and resource optimization allocation of the model at the physical layer, this study introduces a model-to-physical node mapping functor. This paper defines the functor... ,in Indicates the category of the model. This represents the scope of physical nodes. This functor is used to map logical model entities in the model layer to computational nodes in the physical layer, thereby characterizing the deployment location of the model in the actual network environment.
[0178] For each object in the model scope Hanzi Map it to an object in the physical node scope ,Right now , representing the model Deployed and running on physical nodes Due to the scope of the model Modeled as a discrete category, its morphisms only include identity morphisms, functors Mappings in the morphological layer naturally maintain identity, with a primary focus on the model deployment decisions described by the object mappings.
[0179] Set physical nodes In any time slot The available computing, memory, storage, and bandwidth resources are as follows: , , and The deployment of the model on nodes must meet the following resource capacity constraints:
[0180] ;
[0181] ;
[0182] ;
[0183] in, , and Representing the model respectively At any moment The computational, memory, and storage resources consumed during operation; to avoid duplicate deployment of models, each model is further constrained to be deployed on only one physical node in any given time slot as follows:
[0184] ;
[0185] To uniformly characterize the mapping relationship between tasks, models, and physical nodes, a composite functor is introduced:
[0186] ;
[0187] Task After model selection and deployment decisions, the data is ultimately allocated to physical nodes. Execution, the expression is:
[0188] .
[0189] Within the framework of category theory, a discrete-time index category is introduced to describe the dynamic characteristics of task scheduling and resource allocation processes in the time dimension. Its object is discrete time slots A morphism is used to represent the temporal evolution relationship between time slots. This time domain mainly serves as an index structure for system states and scheduling decisions, used to characterize the evolution of task arrival, load changes, and resource allocation across different time slots.
[0190] Load distribution mapping: Given the deployment results of the model, it is necessary to characterize the available resource space of physical nodes across various resource dimensions. To this end, a resource labeling mapping is introduced as { ,in Within the scope of physical nodes, This represents the resource description space, used to characterize the resource capacity of a node in dimensions such as computing, memory, storage, and bandwidth. For any physical node object... , mapping This corresponds to a resource vector ,in , , and Representing nodes respectively The capacity of compute, memory, storage, and bandwidth resources available in the current time slot.
[0191] Calculate the load function: in the time slot Within, each task object Through composite functors Mapped to node objects The corresponding task-model-node ternary relationship can be written as: To accurately describe the computational load of a task, a task is defined. At the node The calculated load function is This indicates that the node is in the time slot. Internal task The equivalent computational workload provided can be formalized according to the load conservation principle as follows:
[0192] ;
[0193] in, It is a scheduling indicator variable, representing the task. Is it in a time slot? Through the model And deployed on nodes implement; It is a task Through the model The complexity factor of logical computation during processing; Describe the task At any moment The amount of input data.
[0194] Resource load function: Based on the definition of the computation load function, it further characterizes the memory and storage resource usage of tasks on nodes.
[0195] Task With model At the node The memory load function is defined as follows:
[0196] ;
[0197] in, For the model The parameters of memory usage; For the task In the model The activation memory requirements generated during the inference process.
[0198] ;
[0199] in, For the model Storage size; For the task Required data storage capacity.
[0200] After completing the mapping decision between tasks, models, and physical nodes, each task is unloaded and transferred to the target computing node for execution according to its corresponding deployment result. The total latency of a task in the system mainly consists of transmission latency, queuing latency, and computation latency.
[0201] Transmission delay model: task In the time slot Transmit from the task source node to the target computing node The transmission delay experienced is defined as:
[0202] ;
[0203] in, Input the amount of data for the task. For at any time bandwidth rate, This represents the distance from the task source to the node. This refers to the link transmission speed.
[0204] Computational latency model: Task At physical nodes The computation latency on a node is determined by both its computational load and the node's allocable computational capacity, and is defined as:
[0205] ;
[0206] in, For nodes Assigned to model Computational power;
[0207] Queuing delay model: at physical nodes In this context, data from different tasks can be viewed as job requests entering the same computation queue. When the rate at which tasks arrive exceeds the instantaneous processing capacity of a node, queuing will occur, introducing queuing latency.
[0208] set up Represents a node In the time slot Given the queue length, the evolution of the queue state between adjacent time slots can be expressed by the following difference equation:
[0209] ;
[0210] in, It is a scheduling indicator variable that indicates the effective mapping between tasks, models, and nodes; The time slot length; This represents the increased effective computational workload.
[0211] Under stable system operation conditions, the queue length satisfies:
[0212] ;
[0213] According to Little's Law, in a steady state, the task... At the node The average queuing delay is approximately expressed as:
[0214] ;
[0215] End-to-end latency model: combining transmission latency, computation latency, and queuing latency, task... At the node The end-to-end delay can be expressed as:
[0216] ;
[0217] Since different tasks have different latency requirements, the end-to-end latency constraint for a task is as follows:
[0218] .
[0219] To achieve end-to-end low latency assurance and maintain long-term stable operation of the system in the edge-cloud collaborative inference scenario of 6G AIaaS, traditional heuristic algorithm decision-making is difficult to achieve: On the one hand, task arrival is sudden and time-varying, and link bandwidth, node computing power and queuing backlog fluctuate over time, which makes it easy for fixed model placement and static task unloading to cause resource congestion and amplified queuing latency in dynamic environments; on the other hand, AI inference services consume multi-dimensional resources such as computing, memory, storage and bandwidth during execution, and there is a strong coupling relationship between QoAIS constraints such as task accuracy and maximum latency and model capabilities and node capacity. If optimization is only performed at a single level, there will often be infeasible solutions that meet latency requirements but exceed resource limits, or end-to-end performance degradation caused by cross-layer fragmentation. Therefore, in order to coordinate the trade-off between task offloading and model deployment location in a dynamic network state, and to obtain feasible and high-performance decisions under the premise that multidimensional resource capacity and service quality constraints are simultaneously met, we need to construct a joint optimization problem of task scheduling and model deployment under multidimensional resource constraints, with minimizing the long-term average end-to-end latency of the system as the core objective, so as to achieve robust adaptation to instantaneous fluctuations and overall optimal long-term performance.
[0220] Let the time window be The set of tasks arriving within this time window is denoted as The total number of tasks is The average end-to-end delay of the system within the time window is defined as:
[0221] ;
[0222] in, Indicates task In its corresponding physical node End-to-end latency during execution;
[0223] A joint optimization problem is established with the objective of minimizing the long-term average end-to-end delay, specifically:
[0224] ;
[0225] ;
[0226] in, Indicates a time range The upper limit of the average end-to-end delay as it approaches infinity.
[0227] To address the objectives of minimizing long-term average end-to-end delay and queue stability constraints, the Lyapunov drift penalty method is employed to transform the long-term stochastic optimization problem into a time-slot-by-time deterministic optimization problem.
[0228] set up Represents physical nodes In the time slot The virtual queue length is used to characterize the computational load backlog on the node side. A quadratic Lyapunov function is defined as follows:
[0229] ;
[0230] In the time slot Inside, node Its service capacity is determined by the computing resources allocated to it. Let its equivalent service volume be:
[0231] ;
[0232] in, Represents a node In the time slot For the model Equivalent computing power allocated; Indicates the duration of the time slot;
[0233] The node queue update equation is expressed as:
[0234] ;
[0235] in, For time slots The computational load of newly arriving nodes; Represents physical nodes In the time slot The amount of virtual queue backlog;
[0236] According to Lyapunov drift analysis, the upper bound of the drift is obtained as follows:
[0237] ;
[0238] in, Represents the conditional expectation operator; Describes a quadratic Lyapunov function; Indicates task After being assigned to a node End-to-end latency during execution; Represents physical nodes In the time slot The amount of virtual queue backlog; This represents a constant related to the system size;
[0239] To optimize system performance while ensuring queue stability, a drift penalty term is introduced, given control parameters. Now, consider minimizing the following single-slot objective function:
[0240] ;
[0241] The original joint optimization problem is equivalently transformed into a time-slot-by-time static optimization subproblem. :
[0242] ;
[0243] ;
[0244] ;
[0245] This ensures that each task uniquely matches a single AI model, and This ensures that the inference accuracy of the selected model meets the task-level QoAIS requirements; in a dynamic system environment, functors Re-optimize within each time slot to adapt to time-varying task arrival processes and resource states.
[0246] set up , , and They represent time slots respectively Internal physical nodes Available computing, memory, storage, and bandwidth resources; model deployment must meet the following resource capacity constraints:
[0247] ;
[0248] ;
[0249] ;
[0250] Furthermore, each model can be deployed on at most one physical node per time slot:
[0251] ;
[0252] in, This indicates that resource feasibility is guaranteed. Forced deployment uniqueness;
[0253] To meet the QoS requirements at the task level, the following latency constraints are introduced:
[0254] .
[0255] Based on a two-stage alternating optimization and genetic algorithm, the task-model matching and model-node deployment sub-problems are solved alternately to obtain the dynamic deployment and task scheduling results of the AI model, specifically:
[0256] Step S4-1: Use a two-stage alternating optimization approach to perform time-slot static optimization subproblems. It can be broken down into two alternating subproblems, as follows:
[0257] In the time slot Using scheduling indicator variables To represent the ternary mapping relationship, and to facilitate the two-stage alternating solution, two types of discrete decision variables are introduced:
[0258] ;
[0259] Indicates task Select Model and satisfy ;
[0260] Representation Model Whether to deploy on a node and satisfy ;
[0261] The time-slot static optimization subproblem Break it down into two alternating subproblems:
[0262] Subproblem P2-F is the task matching phase: given a deployment decision Optimization This mainly affects transmission delay, accuracy constraints, and some queuing estimations;
[0263] Subproblem P2-G is the model deployment phase: given task matching Optimization It mainly affects computation latency and the feasibility of multi-dimensional resource constraints on nodes;
[0264] The GA algorithm consists of two iterative layers: an outer layer for traffic optimization and an inner layer for genetic algorithm iteration.
[0265] Traffic optimization iteration: alternating fixed routes optimization Then fix optimization Continue until the maximum number of iterations is reached or the objective convergence is achieved;
[0266] Genetic Algorithm Iteration: For the discrete combination search of each subproblem, a genetic algorithm is used to approximate the solution, so as to reduce the search complexity of the high-dimensional integer space;
[0267] Task matching phase: Task matching initialization :
[0268] The initial task matching rules are as follows:
[0269] ;
[0270] in, ;
[0271] Model deployment initialization The initialization model deployment rules are as follows:
[0272] ;
[0273] in, Representation Model If deployed to nodes At that time, there are penalties for exceeding the limits of computing, memory, and storage resources;
[0274] Task matching subproblem:
[0275] Chromosomes are used in sets of length based on the number of tasks. The integer encoding is ;
[0276] in Indicates task Select Model ,Right now ;
[0277] The fitness function employs an exterior penalty function, which consists of a slot-by-slot end-to-end delay estimate and a queue penalty. Penalties are imposed on individuals that violate the delay constraints, as shown in the following formula:
[0278] ;
[0279] in, Indicates that in a given Under the given conditions, the estimated end-to-end latency of the task; V is the weight parameter of the Lyapunov optimization; It is the precision default penalty coefficient. It is the time-based penalty coefficient for breach of contract;
[0280] Model deployment subproblem: Chromosomes are used with a length equal to the number of models. The integer encoding is ;
[0281] in, Representation Model Deployed on nodes ,Right now ;
[0282] The fitness function primarily calculates the relevant latency and resource overrun penalties, employing an exterior penalty function. The fitness formula is as follows:
[0283] ;
[0284] in, For the current and The collective resource consumption of nodes is determined jointly; It is the resource default penalty coefficient;
[0285] Obtain the initial solution Then, proceed to the outer AO iteration, the first... The next iteration is executed:
[0286] fixed The task matching subproblem was solved using GA to obtain... ;
[0287] fixed The model deployment subproblem was solved using GA to obtain... ;
[0288] If the objective function decreases by less than the threshold or the maximum number of iterations is reached, then stop.
[0289] The computational scale of the GA algorithm depends on the number of iterations of the outer AO. And the number of generations in the inner genetic algorithm. With population size The overall time complexity is expressed as:
[0290] ;
[0291] The main complexity of the task matching phase is: The main complexity of the model deployment phase is .
[0292] Algorithm computational overhead and system size It exhibits approximately linear growth and has better scalability in large-scale online scenarios compared to exponential search across the entire space.
[0293] To verify the effectiveness and robustness of this invention in the 6G AIaaS scenario, a systematic evaluation of the algorithm's end-to-end latency performance under different parameter conditions was conducted based on an edge-cloud collaborative scheduling simulation platform, and compared with typical algorithms. All comparison algorithms were run under the same network topology, task arrival distribution, and multi-dimensional resource constraints to ensure the fairness of the experimental results. This paper uses the system's average end-to-end latency as the main performance indicator. This indicator comprehensively reflects the transmission latency, computation latency, and queuing latency experienced by tasks in the system, and can comprehensively characterize the impact of task scheduling and model deployment decisions on service performance.
[0294] To verify the effectiveness of the proposed two-stage alternating genetic optimization algorithm in a 6G AIaaS scenario, this invention uses PyCharm Community Edition 2024 on an AMD Ryzen 9 7945HX @ 2.50 GHz as the programming software for evaluating the algorithm. An edge-cloud collaborative scheduling simulation platform was built using Python 3.12, and the simulation environment was constructed using the Abilene backbone network topology. The AO-GA algorithm was compared with several typical scheduling algorithms. This topology contains 12 geographically distributed network nodes, corresponding to major US cities. The computing power, memory, storage, and access bandwidth of each node are differentiated according to its role to characterize a heterogeneous computing network environment. In the experiment, tasks are generated using a dynamic arrival mechanism. The system generates new arriving tasks according to a standard Poisson distribution within each time slot. Each task is obtained by randomly perturbing a task template generated based on the Abilene real traffic matrix. The total number of task templates is set to 500 to ensure diversity in task type and scale. The system uses a discrete time slot approach for scheduling decisions, with each time slot lasting 1 second. At the start of each time slot, the system performs a joint optimization scheduling based on the current active task set and network state. The end-to-end completion delay of a task consists of transmission delay, computation delay, and queuing delay, where transmission delay considers both access bandwidth constraints and the propagation delay due to geographical distance between nodes. In the performance analysis, we compare the proposed AO-GA algorithm with three comparative algorithms: Genetic Algorithm (GA), Particle Swarm Optimization (PSO), and DA-MAB algorithm based on a multi-armed gambling machine. All algorithms are run under the same network topology, task arrival distribution, and resource constraints to ensure the fairness of the comparison results. Detailed simulation parameter settings and algorithm configurations are shown in Table 1.
[0295] Table 1 shows the simulation parameters.
[0296]
[0297] like Figure 3 As shown, the queue backlog initially rises rapidly, then the growth gradually slows down, converging to a stable plateau after approximately 40 time slots, with only slight fluctuations thereafter. This result indicates that AO-GA can maintain the virtual queue at a bounded level, demonstrating that the algorithm, under the Lyapunov drift penalty mechanism, can achieve stable control of long-term constraints and gradually reach steady-state operation.
[0298] like Figure 4As shown, the DA-MAB algorithm does not have an iterative process, so it is not considered in this evaluation. The AO-GA algorithm achieves a rapid reduction in latency in the early stages of iteration and converges to a stable state within a relatively small number of iterations. Although GA and PSO can also reduce latency in the early stages, their overall convergence speed is slower, and they still exhibit some fluctuations in the mid-to-late stages, with their final stable latency being significantly higher than that of AO-GA. The above phenomena indicate that AO-GA has significant advantages in both solution space search efficiency and convergence stability. The main reason for this is that AO-GA, through an alternating optimization mechanism, decomposes the originally highly coupled joint optimization problem of task matching-model deployment-resource allocation into two sub-problems, significantly reducing the decision complexity of a single search; within each sub-problem, a genetic algorithm is introduced to search the discrete solution space, enabling the algorithm to approach high-quality solutions faster while ensuring feasibility, thus exhibiting a faster convergence speed and lower final latency in the overall iteration process.
[0299] like Figure 5 As shown, the average latency of all algorithms increases with the increase in task data volume. This is because the increased scale of task input data directly increases transmission and computation latency, and under limited computing resources, it also exacerbates queuing backlog on the node side, further amplifying queuing latency and computation waiting time. Across all data volume ranges, the AO-GA algorithm consistently maintains the lowest average latency, and its latency growth slope is significantly lower than GA, PSO, and DA-MAB. This indicates that AO-GA can more effectively alleviate node resource contention and queue backlog problems by jointly optimizing task matching and model deployment decisions when facing the load pressure introduced by data volume growth. In contrast, DA-MAB performs relatively stably across all data volume ranges, but its performance is significantly worse than AO-GA under low and high data volume conditions. This is because DA-MAB's model placement decisions are usually updated based on a longer time scale, and its placement strategy struggles to adapt to the current optimal state when the load changes rapidly. Although GA's performance is relatively stable under high data volume conditions, its latency is still higher than DA-MAB and AO-GA. PSO exhibits more pronounced performance degradation when dealing with large amounts of data, indicating its limited adaptability to high-load scenarios under strong constraints of multidimensional resources.
[0300] like Figure 6As shown, the latency performance of each algorithm varies significantly with the dynamic evolution of the system state across time slots. AO-GA maintains the lowest latency level in most time slots with small overall fluctuations, demonstrating good dynamic adaptability and stability. DA-MAB experiences a periodic increase in latency in some time slots. This is because the model deployment strategy is updated relatively infrequently. When the task load or network state changes significantly in a short period, the current deployment scheme may not meet the instantaneous optimal requirements, leading to an increase in queuing latency. GA and PSO exhibit more pronounced volatility during time slot changes, indicating that their ability to jointly satisfy cross-layer consistency and multi-dimensional resource constraints is relatively insufficient in highly time-varying environments.
[0301] This invention targets the 6G AIaaS scenario, constructing an edge-cloud collaborative scheduling simulation platform. Under a unified task-model-physical layer architecture, it achieves joint optimization of dynamic AI model placement and task unloading. First, it introduces category theory modeling methods, proposing a cross-layer unified modeling framework to characterize the structured dependencies and collaborative relationships between AI inference tasks and multi-dimensional network resources. Through task-model functors, model-node functors, and their composite mappings, it achieves a precise description of the cross-layer consistency of end-to-end inference services. Based on this, it constructs a joint optimization problem under multi-dimensional resource constraints, aiming to minimize the long-term average end-to-end latency. The Lyapunov drift penalty method is used to transform the long-term stochastic optimization problem into a time-slot-by-time online decision problem. Furthermore, an AO-GA algorithm is designed to achieve a low-complexity near-optimal solution while ensuring system stability. Extensive simulation results demonstrate that the proposed method outperforms benchmark algorithms such as GA, PSO, and DA-MAB in terms of convergence speed, average end-to-end latency, and dynamic environment adaptability, effectively improving multidimensional resource utilization efficiency and AI inference service performance. Future work will further develop a prototype testing platform and validate the effectiveness of the proposed method in a real network environment.
[0302] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A method for dynamic deployment optimization of AI models for multi-dimensional resources in 6G networks based on category theory, characterized in that: The method includes the following steps: Step S1: Construct a categorized unified modeling framework to formally characterize the cross-layer consistency dependencies of end-to-end AI inference services through functors; Step S2: Under the conditions of dynamic task arrival and time-varying network state, establish a joint optimization model with the goal of minimizing long-term average end-to-end latency, while being constrained by multi-dimensional resources and quality of service. Step S3: Construct a two-stage optimization algorithm based on the Lyapunov drift penalty framework to transform the long-term stochastic optimization problem into a time-slot-by-time online decision problem; Step S4: Based on two-stage alternating optimization and genetic algorithm, alternately solve the task-model matching and model-node deployment sub-problems to obtain the dynamic deployment and task scheduling results of AI model.
2. The method for dynamic deployment optimization of AI models for multi-dimensional resources in 6G networks based on category theory, as described in claim 1, is characterized in that: In step S1, a categorized unified modeling framework is constructed, which formally characterizes the cross-layer consistency dependencies of end-to-end AI inference services using functors. Specifically: Step S1-1: Construct a 6G network architecture that integrates end-edge-cloud collaboration. Based on this 6G network architecture, category theory modeling is introduced, specifically as follows: The 6G network architecture includes a physical layer, a model layer, and a task layer. The physical layer contains heterogeneous computing nodes from terminals, edge, and cloud. The model layer contains a collection of AI models. The task layer contains intelligent inference task requests generated by service terminals. Abstracting physical computing nodes into the physical layer category Abstracting AI models into model categories Abstracting intelligent tasks into task categories ,in, The path categories are generated by the network topology. and It belongs to the discrete category; Physical layer scope Abstracting all physical computing nodes into the physical layer category A collection of objects, expressed as: ; in, Indicates the first One physical computing node; Model category Abstracting AI models into model categories A collection of objects, expressed as: ; in, Indicates the first An AI model; Task scope Abstracting intelligent tasks into task categories A collection of objects, expressed as: ; in, express One intelligent reasoning task instance / request; Step S1-2: Construct the matching functor from task to model and the deployment functor from model to physical node, and form an end-to-end consistent mapping of task-model-node through the composite functor; Step S1-3: Introduce discrete-time indexes to describe the dynamic characteristics of task scheduling and resource allocation in the time dimension.
3. The method for dynamic deployment optimization of AI models for multi-dimensional resources in 6G networks based on category theory, as described in claim 2, is characterized in that: In steps S1-2, a matching functor from task to model and a deployment functor from model to physical node are constructed, and an end-to-end consistent mapping of task-model-node is formed through a composite functor. Specifically: Define functor Hanzi Mapping task objects to model objects; under the influence of the functor, for each task object in the task scope... Hanzi Map it to a model object ,Right now ; To ensure the effectiveness of the mapping, task-model matching must satisfy the following constraints: ; set up Indicates the model's inference accuracy. If we represent the minimum precision requirement for the task, then: ; In dynamic network environments, the mapping from task to model is re-determined at different time slots depending on the system state, that is, the corresponding mapping functor is solved in each discrete time slot. To adapt to task arrival, resource changes and node load fluctuations; Define deployment functor Mapping model objects to physical node objects: For any model ,Model Deployed and run on physical nodes The expression is: ; Set physical nodes In any time slot The available computing, memory, storage, and bandwidth resources are respectively , , and The model deployment must satisfy the multi-dimensional resource capacity constraints of the nodes: ; ; ; in, , and Representing the model respectively At any moment The computational, memory, and storage resources consumed during operation; to avoid duplicate deployment of models, each model is further constrained to be deployed on only one physical node in any given time slot as follows: ; To uniformly characterize the mapping relationship between tasks, models, and physical nodes, a composite functor is introduced: ; Task After model selection and deployment decisions, the data is ultimately allocated to physical nodes. Execution, the expression is: 。 4. The method for dynamic deployment optimization of AI models for multi-dimensional resources in 6G networks based on category theory, as described in claim 2, is characterized in that: Steps S1-3 introduce a discrete-time index to describe the dynamic characteristics of task scheduling and resource allocation in the time dimension, specifically: Introducing the category of discrete-time index The object is discrete time slots ; In the time slot Using scheduling indicator variables This indicates the temporal evolution relationship between time slots; when The condition indicates that the ternary relation is true; otherwise, it is not true. Define task At the node The calculated load function is This indicates that the node is in the time slot. The equivalent computational workload provided by the task, according to the principle of load conservation, can be formalized as follows: ; in, This is the logical computational complexity factor when the task is processed by the model. Describe the task at time The amount of input data; Based on the computational load, further characterize the memory and storage resource consumption of the task on the node: ; in, For model parameter memory usage, The activation memory requirements generated by the task during model inference; ; in, For model storage size, The required data storage size for the task.
5. The method for dynamic deployment optimization of AI models for multi-dimensional resources in 6G networks based on category theory, as described in claim 1, is characterized in that: In step S2, under the conditions of dynamic task arrival and time-varying network state, a joint optimization model is established with the objective of minimizing long-term average end-to-end latency, while being constrained by multi-dimensional resources and quality of service. Specifically: Step S2-1: Determine the arrival of dynamic tasks and network status: Task In the time slot Transmit from source node to target node The transmission delay is defined as: ; in, For a moment bandwidth rate, The distance from the source node. For link transmission speed; Describe the task at time The amount of input data; When the rate at which tasks arrive exceeds the instantaneous processing capacity of a node, queuing will occur, introducing queuing delay. Let... Represents a node In the time slot Given the queue length, the evolution of the queue state between adjacent time slots is represented as: ; in, The time slot length; Under stable system operation conditions, the queue length satisfies: ; According to Little's Law, in a steady state, the task... At the node The average queuing delay is approximately expressed as: ; Combining transmission, computation, and queuing delays, the end-to-end delay of a task is expressed as: ; Since different tasks have different latency requirements, the end-to-end latency constraint for a task is as follows: ; Let the time window be The set of tasks arriving within this window is denoted as The total number of tasks is The average end-to-end delay of the system within the time window is defined as: ; in, Indicates task In its corresponding physical node End-to-end latency during execution; Step S2-2: Establish a joint optimization problem with the objective of minimizing the long-term average end-to-end delay.
6. The method for dynamic deployment optimization of AI models for multi-dimensional resources in 6G networks based on category theory, as described in claim 5, is characterized in that: Step S2-2 establishes a joint optimization problem with the objective of minimizing the long-term average end-to-end delay, specifically as follows: ; ; in, Indicates a time range The upper limit of the average end-to-end delay as it approaches infinity.
7. The method for dynamic deployment optimization of AI models for multi-dimensional resources in 6G networks based on category theory, as described in claim 1, is characterized in that: In step S3, a two-stage optimization algorithm based on the Lyapunov drift penalty framework is constructed to transform the long-term stochastic optimization problem into a time-slot-by-time online decision problem, specifically as follows: set up Represents physical nodes In the time slot The virtual queue length is used to characterize the computational load backlog on the node side. A quadratic Lyapunov function is defined as follows: ; In the time slot Inside, node Its service capacity is determined by the computing resources allocated to it. Let its equivalent service volume be: ; in, Represents a node In the time slot For the model Equivalent computing power allocated; Indicates the duration of the time slot; The node queue update equation is expressed as: ; in, For time slots The computational load of newly arriving nodes; Represents physical nodes In the time slot The amount of virtual queue backlog; According to Lyapunov drift analysis, the upper bound of the drift is obtained as follows: ; in, Represents the conditional expectation operator; Describes a quadratic Lyapunov function; Indicates task After being assigned to a node End-to-end latency during execution; Represents physical nodes In the time slot The amount of virtual queue backlog; This represents a constant related to the system size; To optimize system performance while ensuring queue stability, a drift penalty term is introduced, given control parameters. Now, consider minimizing the following single-slot objective function: ; The original joint optimization problem is equivalently transformed into a time-slot-by-time static optimization subproblem. : ; ; ; in, This ensures that each task uniquely matches a single AI model, and This ensures that the inference accuracy of the selected model meets the task-level QoAIS requirements; in a dynamic system environment, functors Re-optimize within each time slot to adapt to time-varying task arrival processes and resource states; set up , , and They represent time slots respectively Internal physical nodes Available computing, memory, storage, and bandwidth resources; model deployment must meet the following resource capacity constraints: ; ; ; Furthermore, each model can be deployed on at most one physical node per time slot: ; in, This indicates that resource feasibility is guaranteed. This indicates that the deployment must be unique. To meet the QoS requirements at the task level, the following latency constraints are introduced: 。 8. The method for dynamic deployment optimization of AI models for multi-dimensional resources in 6G networks based on category theory, as described in claim 1, is characterized in that: In step S4, based on two-stage alternating optimization and genetic algorithms, the task-model matching and model-node deployment sub-problems are solved alternately to obtain the dynamic deployment and task scheduling results of the AI model. Specifically: Step S4-1: Use a two-stage alternating optimization approach to perform time-slot static optimization subproblems. It can be broken down into two alternating subproblems, as follows: In the time slot Using scheduling indicator variables To represent the ternary mapping relationship, and to facilitate the two-stage alternating solution, two types of discrete decision variables are introduced: ; Indicates task Select Model and satisfy ; Representation Model Whether to deploy on a node and satisfy ; The time-slot static optimization subproblem Break it down into two alternating subproblems: Subproblem P2-F is the task matching phase: given a deployment decision Optimization This mainly affects transmission delay, accuracy constraints, and some queuing estimations; Subproblem P2-G is the model deployment phase: given task matching Optimization It mainly affects computation latency and the feasibility of multi-dimensional resource constraints on nodes; Step S4-2: Use GA to efficiently search for approximate optimal solutions in the discrete solution space.
9. The method for dynamic deployment optimization of AI models for multi-dimensional resources in 6G networks based on category theory, as described in claim 8, is characterized in that: In step S4, the GA algorithm is used to efficiently search for an approximate optimal solution in the discrete solution space, specifically as follows: The GA algorithm consists of two iterative layers: an outer layer for traffic optimization and an inner layer for genetic algorithm iteration. Traffic optimization iteration: alternating fixed routes optimization Then fix optimization Continue until the maximum number of iterations is reached or the objective convergence is achieved; Genetic Algorithm Iteration: For the discrete combination search of each subproblem, a genetic algorithm is used to approximate the solution, so as to reduce the search complexity of the high-dimensional integer space; Task matching phase: Task matching initialization : The initial task matching rules are as follows: ; in, ; Model deployment initialization The initialization model deployment rules are as follows: ; in, Representation Model If deployed to nodes At that time, there are penalties for exceeding the limits of computing, memory, and storage resources; Task matching subproblem: Chromosomes are used in sets of length based on the number of tasks. The integer encoding is ; in Indicates task Select Model ,Right now ; The fitness function employs an exterior penalty function, which consists of a slot-by-slot end-to-end delay estimate and a queue penalty. Penalties are imposed on individuals that violate the delay constraints, as shown in the following formula: ; in, Indicates that in a given Under the given conditions, the estimated end-to-end latency of the task; V is the weight parameter of the Lyapunov optimization; It is the precision default penalty coefficient. It is the time-based penalty coefficient for breach of contract; Model deployment subproblem: Chromosomes are used with a length equal to the number of models. The integer encoding is ; in, Representation Model Deployed on nodes ,Right now ; The fitness function primarily calculates the relevant latency and resource overrun penalties, employing an exterior penalty function. The fitness formula is as follows: ; in, For the current and The collective resource consumption of nodes is determined jointly; It is the resource default penalty coefficient; Obtain the initial solution Then, proceed to the outer AO iteration, the first... The next iteration is executed: fixed The task matching subproblem was solved using GA to obtain... ; fixed The model deployment subproblem was solved using GA to obtain... ; If the objective function decreases by less than the threshold or the maximum number of iterations is reached, then stop. The computational scale of the GA algorithm depends on the number of iterations of the outer AO. And the number of generations in the inner genetic algorithm. With population size The overall time complexity is expressed as: ; The main complexity of the task matching phase is: The main complexity of the model deployment phase is .