Control method, intelligent agent and system for dnn inference in metropolitan optical networks

CN117278564BActive Publication Date: 2026-09-29BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210654601.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2026-09-29
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

然而,基于深度学习等自适应方案,往往面临适应性和扩展性问题

Benefits of technology

[0036]本申请提供的城域光网络中DNN推理的控制方法,若城域光网络中当前接收DNN推理请求的源服务器过载,则基于深度Q网络及迁移学习算法,对所述城域光网络的城域范围内的通信资源和算力资源进行联合调度,以在所述城域光网络中选定目标服务器;输出所述目标服务器的标识,以基于该目标服务器的标识生成DNN推理卸载策略,并根据该DNN推理卸载策略控制所述目标服务器与所述源服务器共同完成所述DNN推理请求对应的推理任务;能够实现城域网络全局范围内的针对DNN分布式推理场景的算力资源和通信资源的联合调度,能够有效降低城域光网络中DNN分布式推理场景中任务的处理时间,进而能够有效保证城域光网络处理DNN推理业务的性能和服务质量,即在获得高性能的前提下,进一步提高智能代理器执行DNN分布式推理的控制过程的适应性和泛化性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117278564B_ABST
    Figure CN117278564B_ABST
Patent Text Reader

Abstract

The application provides a control method, an intelligent agent and a system for DNN inference in a metropolitan optical network. The method comprises: if a source server currently receiving a DNN inference request is overloaded in the metropolitan optical network, jointly scheduling communication resources and computing resources within the metropolitan range of the metropolitan optical network based on a deep Q network and a transfer learning algorithm to select a target server in the metropolitan optical network; outputting an identifier of the target server to generate a DNN inference offloading strategy, and controlling the target server and the source server to jointly complete an inference task corresponding to the DNN inference request according to the DNN inference offloading strategy. The application can realize joint scheduling of computing resources and communication resources in the global range of the metropolitan network for the DNN distributed inference scene, effectively guarantee the performance and service quality of the metropolitan optical network in processing DNN inference services, and improve the adaptability and generalization of the control process of DNN distributed inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of DNN inference technology, and in particular to control methods, intelligent agents and systems for DNN inference in metropolitan optical networks. Background Technology

[0002] Driven by the development of machine learning (ML), the amount of data and computational demands of DNN inference tasks uploaded to edge servers via networks are growing exponentially. Continuing to use classic centralized DNN inference would place significant pressure on communication links and edge servers in metropolitan area networks (MANs). To overcome this problem, distributed DNN inference has been proposed, which utilizes the computing resources of multiple edge nodes distributed within a MAN to provide computational services for a DNN inference task.

[0003] Currently, existing distributed inference processes using deep neural networks (DNNs) typically employ adaptive methods (such as deep reinforcement learning) to achieve optimal performance. However, these adaptive schemes often face challenges related to adaptability and scalability. When the environment in which the intelligent agent operates changes (e.g., the addition or removal of links or servers in the network, or changes in network service characteristics), the agent often requires extensive training to reconverge in the new environment. However, training the intelligent agent is extremely time-consuming and costly, which is unacceptable for optical network control algorithms.

[0004] Therefore, there is an urgent need to design a method that can ensure the performance of DNN distributed inference while also improving the adaptability and generalization of the control process of intelligent agents executing DNN distributed inference. Summary of the Invention

[0005] In view of this, embodiments of this application provide a control method, intelligent agent, and system for DNN inference in metropolitan area optical networks, in order to eliminate or improve one or more defects existing in the prior art.

[0006] One aspect of this application provides a method for controlling DNN inference in a metropolitan area optical network, comprising:

[0007] If the source server currently receiving DNN inference requests in the metropolitan optical network is overloaded, then based on the deep Q network and transfer learning algorithm, the communication resources and computing resources within the metropolitan area of ​​the metropolitan optical network are jointly scheduled to select the target server in the metropolitan optical network.

[0008] The identifier of the target server is output, and a DNN inference offloading strategy is generated based on the identifier of the target server. The target server and the source server are then controlled to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

[0009] In some embodiments of this application, after selecting the target server, the method further includes:

[0010] Based on a greedy strategy, the splitting point latency simulation is performed on each subtask in the inference task corresponding to the DNN inference request, and the splitting point with the lowest latency is selected as the target splitting point based on the corresponding latency simulation results.

[0011] Correspondingly, the output of the target server's identifier includes:

[0012] Output the description of the target server and the target split point, generate a DNN inference offloading strategy based on the identifier of the target server and the target split point, and control the source server to process the subtasks before the target split point and control the target server to process the subtasks after the target split point according to the DNN inference offloading strategy.

[0013] In some embodiments of this application, the joint scheduling of communication and computing resources within the metropolitan area of ​​the metropolitan optical network based on deep Q-networks and transfer learning algorithms to select a target server within the metropolitan optical network includes:

[0014] Based on the current status data of the metropolitan optical network, a deep Q network is used to select a target server in the metropolitan optical network.

[0015] The deep Q-network is pre-learned from the environment in another metropolitan optical network based on the transfer learning algorithm, using the knowledge learned from one metropolitan optical network.

[0016] In some embodiments of this application, a deep Q-network is applied to select a target server in the metropolitan area optical network based on the current state data of the metropolitan area optical network, including:

[0017] The current communication resource status data of the metropolitan optical network, the computing resource status data of each server, and the service feature data of the DNN inference request are obtained to obtain the MDP status of the deep Q network, and an MDP action space corresponding one-to-one between each server and each action is generated.

[0018] The state data is input into the Q-Net of the deep Q-network, so that the Q-Net outputs the simulated long-term discount rewards obtained by taking different actions at the corresponding state, and selects one of the servers corresponding to each long-term discount reward as the target server based on a greedy strategy.

[0019] In some embodiments of this application, the Q-Net includes: a Q-matrix prediction module and a feature extraction module;

[0020] The feature extraction module is used for reuse in different metropolitan area optical networks;

[0021] If the metropolitan optical network where the intelligent agent is located changes or migrates from one metropolitan optical network to another, only the Q matrix prediction module will be trained for migration.

[0022] Another aspect of this application provides a smart agent, comprising:

[0023] The decision module is used to jointly schedule the communication and computing resources within the metropolitan area of ​​the metropolitan area optical network based on deep Q-network and transfer learning algorithm if the source server currently receiving DNN inference requests in the metropolitan area optical network is overloaded, so as to select a target server in the metropolitan area optical network.

[0024] The output module is used to output the identifier of the target server, generate a DNN inference offloading strategy based on the identifier of the target server, and control the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

[0025] In some embodiments of this application, each server in the metropolitan area optical network is configured with an intelligent agent, which is only used to jointly schedule the communication resources and computing resources within the metropolitan area of ​​the metropolitan area optical network, so as to select the target server in the metropolitan area optical network and output the DNN inference offloading strategy.

[0026] Each of the intelligent agents shares the parameters of the feature extraction module and maintains its own unique Q-matrix prediction module.

[0027] Another aspect of this application provides a control system for DNN inference in a metropolitan optical network, comprising: a data processing module and an intelligent agent disposed in the control plane of the software-defined network;

[0028] The data processing module is used to collect data plane resource information and DNN inference requests from the data plane of the software-defined network through the SDN southbound interface, construct a state matrix based on the data plane resource information, and input the corresponding state into the Q-Net of the intelligent agent;

[0029] The intelligent agent is used to execute the control method for DNN inference in the metropolitan optical network to send the identifier of the target server to the data processing module;

[0030] The data processing module is further configured to generate a DNN inference offloading strategy based on the identifier of the target server, and send the DNN inference offloading strategy to the data plane, so that the data plane controls the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

[0031] In some embodiments of this application, the control plane further includes: an experience pool and a reward calculation module;

[0032] The reward calculation module is used to receive the reward for the action based on the DNN inference time fed back by the data plane after the data plane controls the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

[0033] The reward calculation module is also used to generate a corresponding reward based on the DNN inference time and send the reward to the experience pool;

[0034] The experience pool is used to store the corresponding state, action, reward, and next state based on the received reward.

[0035] In some embodiments of this application, the experience pool is also used to provide training data required for offline training of the deep Q-network.

[0036] The control method for DNN inference in a metropolitan optical network provided in this application, if the source server currently receiving DNN inference requests in the metropolitan optical network is overloaded, jointly schedules the communication and computing resources within the metropolitan area of ​​the metropolitan optical network based on deep Q-networks and transfer learning algorithms to select a target server in the metropolitan optical network; outputs the identifier of the target server, generates a DNN inference offloading strategy based on the identifier of the target server, and controls the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy; it can realize the joint scheduling of computing and communication resources for DNN distributed inference scenarios in the metropolitan optical network, effectively reduce the processing time of tasks in DNN distributed inference scenarios in the metropolitan optical network, and thus effectively guarantee the performance and service quality of the metropolitan optical network in processing DNN inference services, that is, while achieving high performance, it further improves the adaptability and generalization of the control process of intelligent agents executing DNN distributed inference.

[0037] Additional advantages, objectives, and features of this application will be set forth in part in the description which follows, and will in part become apparent to those skilled in the art upon review of the following description, or may be learned by practice of the application. The objectives and other advantages of this application can be realized and obtained by means of the structures specifically pointed out in the specification and drawings.

[0038] Those skilled in the art will understand that the purposes and advantages that can be achieved with this application are not limited to those specifically described above, and that the above and other purposes that this application can achieve will be more clearly understood from the following detailed description. Attached Figure Description

[0039] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, do not constitute a limitation thereof. The components in the drawings are not drawn to scale but are merely for illustrating the principles of this application. For ease of illustration and description of certain parts of this application, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to this application. In the drawings:

[0040] Figure 1 This is a schematic diagram illustrating an example of DNN model splitting provided in this application.

[0041] Figure 2 This is a schematic diagram of the DQN architecture provided in this application.

[0042] Figure 3 This is a schematic diagram of the overall flow of the control method for DNN inference in a metropolitan optical network according to an embodiment of this application.

[0043] Figure 4 This is a schematic diagram of a specific process for controlling DNN inference in a metropolitan optical network according to an embodiment of this application.

[0044] Figure 5 This is a schematic diagram of the structure of the intelligent agent in another embodiment.

[0045] Figure 6 A schematic diagram of the control system for DNN inference in a metropolitan optical network provided as an application example of this application.

[0046] Figure 7 A schematic diagram illustrating the structure of a modular Q-Net provided as an application example in this application.

[0047] Figure 8 This diagram illustrates an example of collaboration between multiple smart agents, providing a practical application example for this application.

[0048] Figure 9A flowchart illustrating a distributed inference offloading scheme for DNN inference requests based on DQN and transfer learning, provided as an application example of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain this application, but are not intended to limit it.

[0050] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.

[0051] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0052] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.

[0053] In the following description, embodiments of the present application will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.

[0054] In distributed machine learning, data collected from the edge doesn't need to be transmitted over long distances across networks as in cloud computing, thus improving training and inference speed. As an emerging paradigm, distributed machine learning is applicable to numerous fields such as image recognition, natural language processing, and semantic recognition. In such systems, the trend of introducing distributed inference from distributed neural networks (DNNs) is growing, as it can achieve higher inference accuracy and reduce communication overhead and latency compared to traditional alternatives such as decision trees. For example, based on sensor data collected from field Industrial Internet of Things (IIoT) devices, DNN distributed inference can be used to monitor conditions and predict impending failures, significantly improving the production efficiency of automated production lines.

[0055] In recent years, most existing research related to DNN distributed inference has focused on device-to-device communication via wireless networks, where user equipment is typically connected to cellular networks or wireless local area networks (WLANs). In this way, DNN distributed inference on wireless edge networks is well-suited for various ML services requiring low latency, such as autonomous vehicles, augmented reality, and the Internet of Things (IIoT). However, for specific scenarios requiring inference services to transmit massive amounts of data over long distances and involve significant computation, such as smart cities, wireless communication may face several technical challenges, including the instability of the wireless environment (e.g., dynamic channels and interference) and limited wireless resources (e.g., transmit power and radio spectrum), which can significantly impact the performance of distributed inference. Given the significant breakthroughs of fifth-generation fixed networks (F5G), this application naturally considers utilizing metropolitan area optical networks (MANs) to provide communication support for DNN distributed inference. MANs supporting edge computing can provide greater computing power and a wider communication range for DNN inference services, enabling MANs to mobilize global computing and communication resources to support DNN inference services with specific bandwidth and latency requirements.

[0056] When making distributed offloading decisions for DNN inference services, many factors need to be considered, including network communication resources, the computing resources of each server, and the characteristics of the DNN inference service (computing requirements, scalable nodes, and the number of parameters to be transmitted in the network). This is an NP-hard problem, therefore, in existing solutions, control schemes using adaptive methods (such as deep reinforcement learning) often have better performance.

[0057] However, adaptive solutions, especially those based on deep learning, often face challenges in adaptability and scalability. When the environment in which the intelligent agent operates changes (e.g., links or servers are added or removed from the network, or the characteristics of network services change), the agent often requires extensive training to reconverge in the new environment. Training the intelligent agent is extremely time-consuming and costly, which is unacceptable for optical network control algorithms. Therefore, while ensuring the performance of adaptive solutions, their adaptability must also be considered.

[0058] Examples of existing solutions are as follows:

[0059] Solution 1: Energy-Aware Inference Offloading for DNN Driven Applications in Mobile EdgeClouds (using wireless communication to achieve DNN distributed inference between mobile devices and multiple edge servers): An exact solution to the problem is proposed using Integer Linear Programming (ILP). Then, based on the relaxation of the ILP solution, an approximation algorithm is designed using random rounding techniques.

[0060] Option 2: Deploy DNN distributed inference between edge servers and core cloud servers in the optical network (Deep Reinforcement Learning Based DNN Model Partition in Edge Computing-enabled Metro Optical Network): The A3C deep reinforcement learning method is used to implement the DNN model partitioning and deployment algorithm between the edge nodes of the optical network and the cloud.

[0061] However, regardless of the existing method, at least one of the following problems exists:

[0062] 1) Poor performance: Current research focuses on DNN distributed inference between wireless mobile devices and edge servers, and DNN distributed inference between edge servers and cloud servers. However, the former uses wireless communication, which has low data transmission speed and is susceptible to interference, limiting it to servers near the mobile device and preventing the scheduling of computing and communication resources across the entire metropolitan optical network. The latter involves cooperation between edge servers and cloud services, but the distance between them is often vast (across provinces), resulting in high communication latency and failing to achieve joint scheduling of computing resources across the entire metropolitan optical network.

[0063] 2) Poor generalization and adaptability: The parameters of most intelligent agents are tightly coupled with their environment. This means that when the environment changes (e.g., links or servers are added or removed from the network, or the characteristics of network services change), the intelligent agent often requires extensive training to reconverge in the new environment. Agent training is extremely time-consuming and costly, which is unacceptable for optical network control algorithms. Because the intelligent agent cannot correctly guide service scheduling in the network until it reconverges, it leads to chaotic service scheduling in the optical network, affecting network performance and quality of service.

[0064] Therefore, in order to further improve the adaptability and generalization of the control process of the intelligent agent executing DNN distributed inference while achieving high performance, this application can realize the joint scheduling of computing and communication resources in the global scope of the metropolitan area network, and use deep reinforcement learning and transfer learning to implement a distributed offloading scheme for DNN inference, so as to ensure the high adaptability and generalization of the control process of the intelligent agent executing DNN distributed inference while achieving high performance.

[0065] The various services within a metropolitan area network (MAN) are characterized by uneven distribution and time-varying nature. Taking DNN inference services as an example, the number of services is higher in areas where smart factories are located; higher numbers are generated by smart city facilities near highways; and during the day, user equipment primarily generates services in commercial and office areas, while at night it is mostly located in residential areas. If computing resources can be allocated within the metropolitan area optical network, transferring services from overloaded edge servers to lightly loaded servers, the impact of uneven service distribution on service quality can be mitigated. This would allow the optical network to accommodate more DNN inference services, thereby achieving joint allocation of computing and communication resources across the entire metropolitan area network.

[0066] In one or more embodiments of this application, DNN refers to Deep Neural Networks, which are the foundation of deep learning. For example, algorithms in fields such as computer vision and natural language processing are implemented based on DNNs.

[0067] In one or more embodiments of this application, inference refers to the process of inputting data into a DNN, which then performs layer-by-layer operations on the data and finally outputs the results.

[0068] In one or more embodiments of this application, distributed inference refers to the fact that DNN inference requires a large amount of computation, so multiple servers can be used to cooperate to complete the inference task of a DNN, thereby reducing inference time.

[0069] In one or more embodiments of this application, the intelligent control scheme or control method refers to a control algorithm designed using artificial intelligence, which can intelligently determine the distributed inference scheme based on the characteristics of the communication resource status, computing resource status, and DNN inference requests in the optical network.

[0070] In one or more embodiments of this application, generalization and adaptability refer to the control method's ability to quickly adapt to changes in metropolitan area optical network topology, server distribution, and DNN inference service characteristics. That is, when the environment changes, the control method can quickly converge in the new environment.

[0071] Specifically, the basic principles of the control method, intelligent agent, and control system for DNN inference in metropolitan optical networks provided in this application are as follows:

[0072] A distributed offloading scheme for DNN inference tasks in metropolitan optical networks was implemented using the DQN (Deep Q-Network) method. This scheme can comprehensively consider the optical communication resources in the network, the computing resources and load of multiple edge servers, and the characteristics of DNN inference services (required computing power, maximum latency, scalable nodes, computational and transmission volume of each subtask), and provide a distributed offloading scheme for DNN.

[0073] By adding transfer learning to the above DQN scheme, we obtain the Transfer DQN scheme. The Transfer DQN scheme can achieve fast convergence in new network environments and has high adaptability.

[0074] By reusing some Q-Net parameters among the various intelligent agents, the generalization of the model is improved, and the convergence time of the intelligent agents is further reduced.

[0075] In one or more embodiments of this application, DNN decomposition refers to the fact that a DNN model can be divided into multiple parts, and different sub-tasks can be performed for inference on different devices. See [link to relevant documentation]. Figure 1 A DNN inference task can be represented as n i = <M i,1 M i,2 ,…,M i,j |T max,i >, each M i,j = <p i,j |c i,j > represents a subtask, where p i,j c represents the number of input parameters required for this subtask. i,j This indicates the computational cost required for this subtask. T max,i It is a reasoning request nn i The longest acceptable inference time.

[0076] In one or more embodiments of this application, the inference time of a DNN refers to the fact that, in a distributed scenario, the inference time of a DNN includes two parts: computation time and communication time.

[0077] For request nn i = <M i,1 M i,2 ,…,M i,j |T max,i The computation time can be expressed as formula (1), where each lc i,jThe computation time of a subtask is represented by formula (2). Provide subtasks to the m-th server lc i,j The computing power.

[0078] LC i =lc i,1 +lc i,2 +…+lc i,j (1)

[0079]

[0080] The transmission delay can be expressed as formula (3), where the transmission delay of each subtask is shown in formula (4). If nn i The split point j was used to perform distributed partitioning pf on the DNN. i,j =1, otherwise pf i,j =0. B is the bandwidth of each spectrum slot in the optical network, f j Is the optical network assigned to nn? i The number of spectral slots, ml j It refers to the modulation level, which is related to the communication distance. The shorter the distance, the higher the modulation level.

[0081] LT i =lt i,1 +lt i,2 +…+lt i,j (3)

[0082]

[0083] In one or more embodiments of this application, the principle of the deep Q-network is as follows:

[0084] Deep Q-Networks (DQNs) are a classic reinforcement learning method that uses neural networks to estimate state-action values ​​to guide agents in making decisions. Figure 2 As shown, it uses a target net and a replay buffer to ensure the accuracy of long-term return estimates.

[0085] 1) Markov Property of the Environment: A prerequisite for the convergence of a reinforcement learning agent is that the environment conforms to a Markov decision process (MDP). An MDP consists of tuples {S} M A M ,P M ,R M,γ} indicates that S M It is a state space, A M It is the action space. The state transition function P M :SM ×A M →S M This indicates that the next state depends on the current state and the action. (P) M It can be deterministic or random. R M γ is the immediate reward when the agent reaches a new state. γ is the discount factor for the reward. MDP requires the state transition function to conform to the Markov property, as shown in Equation (5), that is, the state at the next moment depends on the state at the current moment and the action taken at the current moment, and is independent of the state at past moments.

[0086] P{S t+1 =s(t+1)|S t =s(t),…,S1=s(1)}=P{S t+1 =s(t+1)|S t =s(t)} (5)

[0087] 2) DQN Decision Process: The DQN agent faces each state s in the MDP environment. t It will be from the action space A = {a1, a2, ..., a n Choose an action a from} t .

[0088] First, the proxy will s t When input into Q-Net, Q-Net will calculate the Q-matrix Q = {Q(s)} t ,a i ;w)|a i ∈A}, where each Q(s) t ,a i ;w) represents in s t Take action a i Potential long-term discount return U t U t The calculation is shown in formula (6). The DQN agent uses a greedy strategy to select action a. t After that, the environmental state becomes s. t+1 The agent receives a reward r t .

[0089]

[0090] 3) Training the DQN agent: DQN uses temporal-difference learning to update the Q-Net. Assume the agent obtains an experience {s} from the replay buffer. t ,a t ,r t ,st+1 The agent will first calculate TD-target y. t The loss function is shown in equations (7) and (8). Where Q... target It is an estimate made by target-net, which replicates the parameters of Q-Net at a certain frequency.

[0091] y t =r t +γ*max a Q target (s t ,a;w) (7)

[0092]

[0093] The DQN proxy updates the parameters w of Q-net using a gradient descent strategy, as shown in Equation (9), where α is the learning rate. After extensive training, Q-Net can accurately estimate U. t This guides the agent's decision-making.

[0094]

[0095] In one or more embodiments of this application, transfer learning in reinforcement learning refers to the use of knowledge from one or more related but different MDPs to improve the performance of a target domain MDP. In this application, the agent learns about the environment in another network based on knowledge learned from one metropolitan optical network, thereby reducing the amount of training required for the model to converge in the new network.

[0096] Based on this, embodiments of this application provide a control method for DNN inference in a metropolitan area optical network, see [link to relevant documentation]. Figure 3 The control method for DNN inference in the metropolitan area optical network specifically includes the following:

[0097] Step 100: If the source server currently receiving DNN inference requests in the metropolitan optical network is overloaded, then based on the deep Q network and transfer learning algorithm, the communication resources and computing resources within the metropolitan area of ​​the metropolitan optical network are jointly scheduled to select a target server in the metropolitan optical network.

[0098] In step 100, this application constructs the distributed offloading process of DNN inference services in the metropolitan area optical network as an MDP model so that the DQN method can be applied. In order to simplify the optimization problem and reduce the algorithm complexity, this application assumes that when a server (source server) is overloaded, it can only select another server (target server) to jointly complete the inference of a DNN request.

[0099] Step 200: Output the identifier of the target server, generate a DNN inference offloading strategy based on the identifier of the target server, and control the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

[0100] It is understood that the execution entity of the DNN inference control method in the metropolitan optical network can be an intelligent agent. The intelligent agent can send the identifier of the target server to a data processing module, so that the data processing module generates a DNN inference offloading strategy based on the identifier of the target server, and sends the DNN inference offloading strategy to the data plane, so that the data plane controls the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

[0101] As can be seen from the above description, the control method for DNN inference in a metropolitan optical network provided in this application embodiment can realize the joint scheduling of computing and communication resources for DNN distributed inference scenarios across the entire metropolitan optical network. This can effectively reduce the processing time of tasks and thus effectively guarantee the performance and quality of service of the metropolitan optical network in processing DNN inference services. In other words, while achieving high performance, it further improves the adaptability and generalization of the control process for the intelligent agent to execute DNN distributed inference.

[0102] In the control method for DNN inference in a metropolitan area optical network provided in this application embodiment, the following content is further included between steps 100 and 200 of the control method for DNN inference in a metropolitan area optical network:

[0103] Step 010: Based on a greedy strategy, perform split point latency simulation on each subtask in the inference task corresponding to the DNN inference request, and select the split point with the lowest latency as the target split point based on the corresponding latency simulation results.

[0104] Correspondingly, see Figure 4 Step 200 further includes the following:

[0105] Step 210: Output the description of the target server and the target split point, generate a DNN inference offloading strategy based on the identifier of the target server and the target split point, and control the source server to process the sub-tasks before the target split point according to the DNN inference offloading strategy, and control the target server to process the sub-tasks after the target split point.

[0106] As described above, this application uses a greedy strategy to guide the partitioning of the DNN, that is, to determine which subtasks are still computed on the source server and which subtasks are transferred to the target server for computation. For nn i The greedy strategy simulates the latency of all split points and selects the split point with the lowest simulated latency. Subtasks before the split point are computed by the source server, and subtasks after the split point are computed by the target server.

[0107] In the control method for DNN inference in a metropolitan optical network provided in the embodiments of this application, see [link to relevant documentation]. Figure 4 Step 100 in the control method for DNN inference in the metropolitan area optical network specifically includes the following:

[0108] Step 110: Based on the current state data of the metropolitan optical network, a deep Q-network is applied to select a target server in the metropolitan optical network; wherein, the deep Q-network is pre-learned based on the transfer learning algorithm, using knowledge learned from one metropolitan optical network to learn the environment in another metropolitan optical network.

[0109] Understandably, for distributed offloading schemes using deep reinforcement learning methods to implement DNN inference, there are many other methods besides DQN in the same MDP environment, such as the Actor-Critic method, AsynchronousAdvantage Actor-Critic (A3C), and Deep Deterministic Policy Gradient (DDPG). However, these methods are all very similar; they all learn the parameters of the neural network from experience to better predict the state-action value function or state-value function, thereby guiding the agent's decision-making.

[0110] In the control method for DNN inference in a metropolitan area optical network provided in this application embodiment, step 110 of the control method for DNN inference in the metropolitan area optical network specifically includes the following:

[0111] Step 111: Obtain the current communication resource status data of the metropolitan optical network, the computing resource status data of each server, and the service feature data of the DNN inference request to obtain the MDP status of the deep Q network, and generate the MDP action space corresponding to each server and each action.

[0112] Step 112: Input the state data into the Q-Net of the deep Q network, so that the Q-Net outputs the simulated long-term discount rewards obtained by taking different actions at the corresponding state, and selects one of the servers corresponding to each long-term discount reward as the target server based on a greedy strategy.

[0113] As can be seen from the above description, the embodiments of this application transmit the DNN inference service in the congested server to other idle servers in the metropolitan area for processing through the optical network, thereby realizing the joint scheduling of communication resources and computing resources in the metropolitan area and improving the capacity of the metropolitan optical network for DNN inference services.

[0114] In a control method for DNN inference in a metropolitan area optical network provided in this application embodiment, the Q-Net in the control method for DNN inference in the metropolitan area optical network includes: a Q matrix prediction module and a feature extraction module;

[0115] The feature extraction module is used for reuse in different metropolitan area optical networks;

[0116] If the metropolitan optical network where the intelligent agent is located changes or migrates from one metropolitan optical network to another, only the Q matrix prediction module will be trained for migration.

[0117] Specifically, Q-Net's role is to extract features from transmission resources, computational resources, and DNN inference requests, and predict the Q-matrix by analyzing these features. This application splits Q-Net into a feature extraction module (feature extractor) and a Q-matrix prediction module (predictor). The feature extraction module (e.g., routing feature extractor, server feature extractor, request feature extractor, etc.) can be reused in different optical networks. When the agent's network changes or it migrates from one network to another, only the predictor needs to be transferred for training, without changing the parameters of the feature extractor. This design further reduces the training load required for the agent to reconverge in a new environment.

[0118] This application also provides an intelligent agent for performing all or part of the control method for DNN inference in the metropolitan area optical network, see [link to relevant documentation]. Figure 5 The intelligent agent specifically includes the following:

[0119] The decision module 10 is used to jointly schedule the communication resources and computing resources within the metropolitan area of ​​the metropolitan area optical network based on deep Q-network and transfer learning algorithm if the source server currently receiving DNN inference requests in the metropolitan area optical network is overloaded, so as to select a target server in the metropolitan area optical network.

[0120] The output module 20 is used to output the identifier of the target server, generate a DNN inference offloading strategy based on the identifier of the target server, and control the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

[0121] The embodiments of the intelligent agent provided in this application can be used to execute the processing flow of the embodiment of the control method for DNN inference in the metropolitan optical network described above. Its functions will not be repeated here, but can be referred to the detailed description of the embodiment of the control method for DNN inference in the metropolitan optical network described above.

[0122] As can be seen from the above description, the intelligent agent provided in this application embodiment can realize the joint scheduling of computing and communication resources for DNN distributed inference scenarios in the metropolitan area network, which can effectively improve the service processing speed in DNN distributed inference scenarios in the metropolitan area optical network, and thus effectively guarantee the performance and service quality of the metropolitan area optical network. That is, under the premise of obtaining high performance, the adaptability and generalization of the intelligent agent are further improved.

[0123] In an embodiment of this application, each server in the metropolitan optical network is configured with an intelligent agent. The intelligent agent is only used to jointly schedule the communication resources and computing resources within the metropolitan area of ​​the metropolitan optical network, so as to select the target server in the metropolitan optical network and output the DNN inference offloading strategy.

[0124] In this system, each intelligent agent shares the parameters of the feature extraction module and maintains its own unique Q-matrix prediction module. This ensures the generalization ability of the feature extractor while allowing the predictor to better adapt to specific environments.

[0125] This application also provides a control system for DNN inference in a metropolitan optical network that includes an intelligent agent, see [link to relevant documentation]. Figure 6 The control system for DNN inference in the metropolitan area optical network specifically includes the following:

[0126] A data processing module and an intelligent agent are configured in the control plane of a software-defined network. The data processing module collects data plane resource information and DNN inference requests from the data plane of the software-defined network via the SDN southbound interface, constructs a state matrix based on the data plane resource information, and inputs the corresponding state into the Q-Net of the intelligent agent. The intelligent agent executes the control method for DNN inference in the metropolitan optical network to send the identifier of the target server to the data processing module. In this application, the intelligent agent can be written as T-DQNAgent, the Q matrix prediction module can be written as Q-matrix, the action can be written as Action, and the state can be written as State.

[0127] The data processing module is further configured to generate a DNN inference offloading strategy based on the identifier of the target server, and send the DNN inference offloading strategy to the data plane, so that the data plane controls the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

[0128] The control plane also includes: an experience pool and a reward calculation module;

[0129] The reward calculation module is used to receive the DNN inference time fed back by the data plane after the target server and the source server jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

[0130] The reward calculation module is also used to generate a corresponding reward based on the DNN inference time and send the reward to the experience pool;

[0131] The experience pool is used to store the corresponding state, action, reward, and next state based on the received reward. The experience pool is also used to provide training data for offline training of the deep Q-network.

[0132] To further illustrate this solution, this application also provides a specific application example of a control method for DNN inference in a metropolitan area optical network, specifically a distributed inference offloading scheme for DNN inference requests based on DQN and transfer learning, which includes the following:

[0133] This application constructs the distributed offloading process of DNN inference services in metropolitan optical networks as an MDP model so that the DQN method can be applied. To simplify the optimization problem and reduce algorithm complexity, this application assumes that when one server (source server) is overloaded, it can only choose another server (target server) to jointly complete the inference of a DNN request.

[0134] 1) State design in the MDP model: In order to realize the decision of DNN inference offloading, the DQN agent needs to know the communication resource status of the network, the computing resource status of the server, and the characteristics of inference requests from the environment. Therefore, this application designs the state as a one-dimensional matrix containing the above information, as shown in formulas (10) to (13).

[0135] Among them, s e (t) represents the server-related part of the state, ld i (t) represents the load of the i-th server, cp iThis refers to the computing power of the i-th server. To achieve the reuse of Q-Net in different networks, this application must ensure that Q-Net has the same input and output dimensions in different networks. The inference time of a DNN is related not only to the source and target servers but also to the distance between the two servers. A greater distance leads to lower modulation of communication, resulting in longer communication time. Simultaneously, the more links data passes through, the more communication resources are consumed, as spectrum resources need to be allocated to all links it traverses. Due to the limitations of spectrum consistency, continuity, and uniqueness, data transmission through more links also means a higher probability of resource allocation failure. Therefore, this application does not consider offloading services to servers that are too far away, but only considers the N servers closest to the source server. The biggest advantage of this is that the input and output dimensions of Q-Net are fixed, allowing Q-Net to be reused in different network environments. r (t) is the part related to communication resources, representing the resource utilization rate on the route from the source server to N target servers. nn (t) represents the characteristics of the service, including the computational load required, the size of the parameters to be transmitted, and the maximum latency tolerance.

[0136] s(t)=[s e (t),s r (t),s nn (t)] T (10)

[0137]

[0138] s r (t)={Rs1(t),Rs2(t),…,Rs N (t)} (12)

[0139] s nn (t)={p t,1 ,c t ,T max,t} (13)

[0140] 2) Action Space: The Transfer DQN agent decides which surrounding server to use as the target server for unloading DNN inference requests. Therefore, the action space can be represented by formula (14), where each action represents selecting a specific server as the target server.

[0141] A = [es1,es2,…,es] N ] T (14)

[0142] 3) Greedy Strategy for DNN Partitioning: This application uses a greedy strategy to guide the partitioning of the DNN, that is, to determine which subtasks are still computed on the source server and which subtasks are transferred to the target server for computation. For nn i The greedy strategy simulates the latency of all split points and selects the split point with the lowest simulated latency. Subtasks before the split point are computed by the source server, and subtasks after the split point are computed by the target server.

[0143] 4) Design of the reward function: The value of the reward function directly affects the discounted reward U. t This, in turn, affects the prediction values ​​of Q-Net. If the reward function is strongly correlated with the business or network features, the amount of training required for Q-Net to reconverge when the environment changes will be relatively high. Therefore, this application uses a loosely coupled reward function, as shown in formula (15), where the reward function is -1 if the business times out, and 1 otherwise.

[0144]

[0145] 5) Modular Q-Net Design: The role of Q-Net is to extract features from transmission resources, computational resources, and DNN inference requests, and to predict the Q-matrix by analyzing these features. This application decomposes Q-Net into a feature extraction module (feature extractor) and a Q-matrix prediction module (predictor), such as... Figure 7 As shown, feature extraction modules (such as route feature extractors, server feature extractors, request feature extractors, etc.) can be reused in different optical networks. When the network where the agent is located changes or migrates from one network to another, only the predictor needs to be transferred for training, without changing the parameters of the feature extractor. This design can further reduce the amount of training required for the agent to reconverge in a new environment.

[0146] 6) Collaboration among multiple agents: This application configures an agent for each server in the network. This agent is only responsible for the distributed computation offloading of DNN inference tasks uploaded to that server, such as... Figure 8 As shown. Doing so will bring two benefits.

[0147] First, besides the spectrum resource utilization of the pathways between servers, the server status, and the characteristics of DNN inference requests, other factors also affect the latency and quality of service of DNN inference. These include the degree of spectrum fragmentation on the pathways between servers and the subtask characteristics of DNN inference requests across different servers. These factors have different characteristics in different areas of the network, making it difficult to incorporate them into the MDP state. In other words, there are factors affecting U... tThe environmental factors θ. For different source servers, the agent needs to fit different θ, as shown in Equation (16). Using a multi-agent design, each agent can use a predictor to fit the environmental parameters θ around the server it is responsible for, thereby improving the prediction accuracy of the Q-matrix and thus improving the quality of agent decision-making.

[0148] On the other hand, if we view the entire network as an MDP environment with a proxy, P M The uncertainty is high, although it conforms to the Markov property, which means that s(t+1) depends on s(t) and a t The dependency is small. Because transfer DQN only considers the N servers closest to the source server, the state s(t+1) and nn are very similar. i+1 The correlation is higher than that of s(t) and a. t , and nn i+1 With s(t) and a t It is irrelevant. Therefore, the agent cannot correctly estimate U when making sequential decisions. t The only option is to decrease γ to obtain a higher immediate reward. This does not meet the global optimum expected by this application to achieve DNN inference offloading.

[0149]

[0150] Each agent shares the parameters of the feature extractor while maintaining its own independent predictor. This ensures that the feature extractor does not overfit to a specific environment, giving it better generalization ability across different servers and metropolitan area optical networks. Simultaneously, the predictor can better fit to a specific environment, guaranteeing its prediction accuracy for the Q-matrix within that environment.

[0151] 7) Module Design: Software Defined Networking (SDN) is defined as a control framework that supports the programmability of network functions and protocols by decoupling the data plane and control plane. It is currently integrated into most network devices. The SDN controller can collect information from the entire network and control the provision of services across the entire network, solving the information collection and signaling transmission problems of intelligent control methods. Simultaneously, the centralized architecture of SDN can easily provide computing power support for intelligent control methods. Therefore, this application designs a Transfer DQN deployment scheme based on the control plane. By deploying a Transfer DQN agent on the SDN control plane, this application can realize the control of DNN inference offloading in the optical network, such as... Figure 6 As shown.

[0152] See Figure 9The data processing module collects data plane resource information and DNN inference request information through the SDN southbound interface. Based on the information from the data plane, the data processing module constructs the state 's' of the MDP and inputs 's' into the Q-Net. The Q-Net generates a Q-matrix based on 's'. Then, the decision module selects an action 'a' based on the Q-matrix. The data processing module generates a DNN inference offloading policy based on 'a'. The data plane uses the policy to offload DNN inference requests and provides feedback on the DNN's inference time. The data processing module receives the policy feedback and generates a reward 'r'. The experience pool stores the state, action, corresponding reward, and next state as an experience.

[0153] Transfer DQN (T-DQN) uses data from an experience pool for offline training to ensure convergence within the network. During offline training, Transfer DQN replicates the Q-Net and trains on the replica. After training is complete and it is ensured that the Q-Net replica is not overfitting, Transfer DQN uses the replica to overlay the Q-Net for online decision-making. Therefore, offline training does not affect the real-time decision-making of Transfer DQN in metropolitan optical networks.

[0154] In summary, this application proposes a control method for DNN inference services in metropolitan optical networks, using DQN to solve the complex DNN inference offloading problem. Based on the control structure of the optical network, a deployment scheme (modular design) is designed for its use in the optical network control plane. Transfer learning is incorporated into the DQN scheme, including fixed Q-Net dimensions, a generalizable reward function, modular Q-Net design, and parameter sharing among multi-agents.

[0155] The specific advantages are as follows:

[0156] Advantage 1: By using a metropolitan area optical network to carry DNN inference tasks, joint resource scheduling within the metropolitan area is achieved. Due to the division of commercial, residential, and industrial areas within a city, as well as residents' lifestyles, services in a metropolitan area optical network often exhibit tidal characteristics. This means that a large amount of traffic tends to concentrate in a particular area at a certain time, causing local network resource strain and resulting in a decline in service quality. Due to the high bandwidth, low latency, and low interference characteristics of optical communication, this design can use the optical network to transmit DNN inference services from congested servers to other idle servers within the metropolitan area for processing. This achieves joint scheduling of communication and computing resources within the metropolitan area, thereby increasing the capacity of the metropolitan area optical network for DNN inference services.

[0157] Advantage 2: While increasing the capacity of metropolitan optical networks for DNN inference services, the design ensures the adaptability and generalization of the control scheme. A key feature of this design is its ability to quickly adapt to environmental changes. When the agent is migrated to a new optical network, it can converge rapidly with minimal transfer training. This is thanks to the modular Q-Net, the generalizable reward function, and the multi-agent parameter sharing strategy.

[0158] This application also provides a computer device (i.e., an electronic device), which may include a processor, a memory, a receiver, and a transmitter. The processor is used to execute the control method for DNN inference in a metropolitan area optical network mentioned in the above embodiments. The processor and the memory can be connected via a bus or other means, taking a bus connection as an example. The receiver can be connected to the processor and the memory via wired or wireless means.

[0159] The processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0160] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the DNN inference control method in the metropolitan area optical network in the embodiments of this application. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the DNN inference control method in the metropolitan area optical network in the above method embodiments.

[0161] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0162] The one or more modules are stored in the memory, and when executed by the processor, they perform the control method for DNN inference in the metropolitan optical network in the embodiment.

[0163] In some embodiments of this application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, memory, receiver, and transmitter may be connected via a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.

[0164] As one implementation method, the functions of the receiver and transmitter in this application can be implemented by transceiver circuits or dedicated transceiver chips, and the processor can be implemented by dedicated processing chips, processing circuits or general-purpose chips.

[0165] As another implementation approach, the server provided in this application embodiment can be implemented using a general-purpose computer. That is, the program code implementing the processor, receiver, and transmitter functions is stored in memory, and the general-purpose processor implements the processor, receiver, and transmitter functions by executing the code in memory.

[0166] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned control method for DNN inference in a metropolitan area optical network. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0167] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave.

[0168] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0169] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0170] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to the embodiments of this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A control method for DNN inference in a metropolitan area optical network, characterized in that, include: If the source server currently receiving DNN inference requests in the metropolitan optical network is overloaded, then based on the deep Q network and transfer learning algorithm, the communication resources and computing resources within the metropolitan area of ​​the metropolitan optical network are jointly scheduled to select the target server in the metropolitan optical network. Based on a greedy strategy, the splitting point latency simulation is performed on each subtask in the inference task corresponding to the DNN inference request, and the splitting point with the lowest latency is selected as the target splitting point based on the corresponding latency simulation results. The identifier of the target server and the target split point are output to generate a DNN inference offloading strategy based on the identifier of the target server and the target split point. The source server is controlled to process the sub-tasks before the target split point and the target server is controlled to process the sub-tasks after the target split point according to the DNN inference offloading strategy, so as to jointly complete the inference task corresponding to the DNN inference request.

2. The control method for DNN inference in a metropolitan area optical network according to claim 1, characterized in that, The method based on deep Q-networks and transfer learning algorithms jointly schedules communication and computing resources within the metropolitan area of ​​the metropolitan optical network to select target servers within the metropolitan optical network, including: Based on the current status data of the metropolitan optical network, a deep Q network is used to select a target server in the metropolitan optical network. The deep Q-network is pre-learned from the environment in another metropolitan optical network based on the transfer learning algorithm, using the knowledge learned from one metropolitan optical network.

3. The control method for DNN inference in a metropolitan area optical network according to claim 2, characterized in that, Based on the current state data of the metropolitan area optical network, a deep Q network is applied to select a target server within the metropolitan area optical network, including: The current communication resource status data of the metropolitan optical network, the computing resource status data of each server, and the service feature data of the DNN inference request are obtained to obtain the MDP status of the deep Q network, and an MDP action space corresponding one-to-one between each server and each action is generated. The state data is input into the Q-Net of the deep Q-network, so that the Q-Net outputs the simulated long-term discount rewards obtained by taking different actions at the corresponding state, and selects one of the servers corresponding to each long-term discount reward as the target server based on a greedy strategy.

4. The control method for DNN inference in a metropolitan area optical network according to claim 3, characterized in that, The Q-Net includes: a Q-matrix prediction module and a feature extraction module; The feature extraction module is used for reuse in different metropolitan area optical networks; If the metropolitan optical network where the intelligent agent is located changes or migrates from one metropolitan optical network to another, only the Q matrix prediction module will be trained for migration.

5. An intelligent agent, characterized in that, include: The decision module is used to jointly schedule the communication and computing resources within the metropolitan area of ​​the metropolitan area optical network based on deep Q-network and transfer learning algorithm if the source server currently receiving DNN inference requests in the metropolitan area optical network is overloaded, so as to select a target server in the metropolitan area optical network. Based on a greedy strategy, the splitting point latency simulation is performed on each subtask in the inference task corresponding to the DNN inference request, and the splitting point with the lowest latency is selected as the target splitting point based on the corresponding latency simulation results. The output module is used to output the identifier of the target server and the target split point, so as to generate a DNN inference offloading strategy based on the identifier of the target server, and control the source server to process the sub-tasks before the target split point according to the DNN inference offloading strategy and the target split point, and control the target server to process the sub-tasks after the target split point, so as to jointly complete the inference task corresponding to the DNN inference request.

6. The intelligent agent according to claim 5, characterized in that, Each server in the metropolitan area optical network is configured with an intelligent agent. The intelligent agent is only used to jointly schedule the communication resources and computing resources within the metropolitan area of ​​the metropolitan area optical network, so as to select the target server in the metropolitan area optical network and output the DNN inference offloading strategy. Each of the intelligent agents shares the parameters of the feature extraction module and maintains its own unique Q-matrix prediction module.

7. A control system for DNN inference in a metropolitan area optical network, characterized in that, include: Data processing modules and intelligent agents are set up in the control plane of a software-defined network; The data processing module is used to collect data plane resource information and DNN inference requests from the data plane of the software-defined network through the SDN southbound interface, construct a state matrix based on the data plane resource information, and input the corresponding state into the Q-Net of the intelligent agent; The intelligent agent is used to execute the control method for DNN inference in a metropolitan optical network as described in any one of claims 1 to 4, so as to send the identifier of the target server to the data processing module; The data processing module is further configured to generate a DNN inference offloading strategy based on the identifier of the target server, and send the DNN inference offloading strategy to the data plane, so that the data plane controls the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy.

8. The control system for DNN inference in a metropolitan area optical network according to claim 7, characterized in that, The control plane also includes: an experience pool and a reward calculation module; The reward calculation module is used to receive the reward for the action based on the DNN inference time fed back by the data plane after the data plane controls the target server and the source server to jointly complete the inference task corresponding to the DNN inference request according to the DNN inference offloading strategy. The reward calculation module is also used to generate a corresponding reward based on the DNN inference time and send the reward to the experience pool; The experience pool is used to store the corresponding state, action, reward, and next state based on the received reward.

9. The control system for DNN inference in a metropolitan area optical network according to claim 8, characterized in that, The experience pool is also used to provide the training data required for offline training of the deep Q network.

Citation Information

Patent Citations

  • Mobile edge computing task unloading method and device based on transfer learning

    CN113504987A