Large-scale optical network resource allocation method based on iterative learning
Through iterative learning and reinforcement learning methods, combined with node and link feature coding, and using EGAT model, the problems of high computing complexity and frequent topological changes in large-scale optical networks are solved, efficient resource allocation and cross-topology adaptation are achieved, and scheduling performance and resource utilization are improved.
Patent Information
- Application Number
- CN202510379782.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-28
AI Technical Summary
There are problems in large-scale optical networks with high computational complexity, frequent topology changes, and insufficient cross-topology generalization capabilities, resulting in low resource scheduling efficiency and difficulty in achieving efficient resource allocation in dynamic environments.
Using an iterative learning and reinforcement learning method, through node feature encoding and link feature encoding, the edge feature enhancement graph attention network (EGAT) model is used, the scheduling strategy is dynamically adjusted, and the optimal resource allocation scheme is independently learned to adapt to different network topology structures.
It significantly improves the scheduling performance and efficiency of large-scale optical networks, enhances cross-topology generalization capabilities, reduces computing complexity, and improves resource utilization and network adaptability.
Smart Images

Figure CN120264174A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technologies, and in particular, relates to a large-scale optical network resource allocation method based on iterative learning. Background Art
[0002] With the gradual expansion of optical networks from backbone networks to metropolitan area networks, the rapid growth of the scale of optical network nodes has become the main development trend of the next-generation optical networks; the introduction of digital twin technology has also provided a solid foundation for the deployment of large-scale optical networks.
[0003] However, the growth of network scale has led to an exponential increase in computational complexity, posing a huge challenge to fast and reliable network configuration; especially the demand for models that can effectively generalize in diverse and dynamic environments has further exacerbated these challenges; therefore, it is crucial to enhance the cross-topology generalization ability of service provision in large-scale optical networks.
[0004] In the prior art, there are many technical problems in the resource scheduling aspect of large-scale optical networks, such as:
[0005] (1) Huge network scale: Large-scale optical networks involve hundreds to thousands of nodes and links, with complex and frequently changing topological structures. The huge network state space makes the solution of scheduling problems extremely complex.
[0006] As the network scale continues to expand, traditional scheduling algorithms are difficult to meet its computational complexity and real-time requirements;
[0007] (2) Dynamic topology and multi-objective optimization: The topology and traffic in large-scale optical networks change frequently, which requires scheduling algorithms to have efficient real-time adaptation capabilities.
[0008] At the same time, multiple objectives such as bandwidth requirements and network congestion need to be considered during the scheduling process. How to balance these objectives and ensure efficient resource allocation in complex environments is a major challenge;
[0009] (3) Balance between global information and local decision-making: In large-scale networks, how to reasonably utilize global network information for optimization without adding excessive computational overhead is still an urgent problem to be solved.
[0010] Although local decision-making has high computational efficiency, it may lead to waste of global resources. How to find an effective balance between the two is the key to algorithm design.
[0011] To overcome the above problems, those skilled in the prior art have designed many optical network resource allocation methods. Although they can operate effectively in certain specific scenarios, when facing large-scale optical networks, there are still many deficiencies, such as:
[0012] (1) Dependence on predefined paths and fixed topologies: Existing methods usually rely on a predefined set of paths or a static topology structure.
[0013] In large-scale optical networks, due to the sharp increase in the number of nodes, the operation and maintenance difficulty of optical networks also rises rapidly. Nodes or links are dynamically added or deleted due to faults or adjustments, etc., resulting in continuous changes in the optical network topology. The above methods are difficult to quickly adapt to these changes, leading to low resource scheduling efficiency. Especially when dealing with dynamic demands, it is often impossible to quickly find the optimal path.
[0014] (2) High computational complexity: Existing resource scheduling methods, such as those based on the shortest path, heuristic search, etc., often need to enumerate and evaluate all possible paths in the network.
[0015] When the network scale expands, the computational complexity of this method increases exponentially; for large-scale optical networks, such methods are difficult to give an effective scheduling scheme within a reasonable time, especially in scenarios with high real-time requirements.
[0016] (3) Lack of cross-topology generalization ability: Many traditional scheduling algorithms are designed for specific network topologies or structures and lack good cross-topology generalization ability.
[0017] This means that when the network topology changes, traditional methods need to be readjusted or retrained, which increases the computational and storage overhead and may lead to unstable performance. Summary of the Invention
[0018] To solve the above technical problems, the present invention provides a large-scale optical network resource allocation method based on iterative learning, which gradually constructs a service allocation strategy through iterative learning and uses reinforcement learning for training. This method can not only dynamically adjust the scheduling strategy according to the real-time network state, autonomously learn the optimal end-to-end resource allocation scheme, but also effectively cope with large-scale network topology changes, significantly improving the global performance and efficiency of scheduling; in addition, the present invention has strong cross-topology generalization ability, can operate stably in optical networks of different scales, and achieve higher resource utilization and network adaptability.
[0019] In the first aspect, a large-scale optical network resource allocation method based on iterative learning includes:
[0020] Step 1: Initialize the optical network topology, background service distribution, and current service requests;
[0021] As an example, the network topology includes: the number of nodes, the number of links, the topological connection relationship, and the number of available wavelengths or frequency slots on the links; the background traffic distribution refers to filling a certain number of random background traffic in the optical network, and the number of background traffic is determined by the network load rate; the current traffic request includes: the source node, the destination node, and the bandwidth requirement.
[0022] Step 2: According to the current optical network state, calculate the node feature codes of all nodes and the link feature codes of all links {v i}, {e j}, and the mask vector M, which together constitute the input state S of the decision-making Agent I ;
[0023] As an example, the process of calculating the node feature codes of all nodes and the link feature codes of all links {v i}, {e j}, and the mask vector M is as follows: According to the topological structure relationship, the background traffic distribution, and the current traffic request status, dynamically update the node feature code v i of each node and the link feature code e j of each link; According to the topological situation, calculate a mask vector N is the current number of topological nodes; The Agent refers to the intelligent agent of reinforcement learning.
[0024] As an example, the node feature code v i of each node means that each node has a K-dimensional node feature code N is the current number of topological nodes, where: each dimension corresponds to a state feature of a node, and each feature needs to reflect the key information required for Agent decision-making.
[0025] The node feature code includes: the number of services passing through the current node, the total number of time slots occupied by the services passing through the current node, the spectral connectivity, the spectral connection block, the shortest path length from the current service end point, and the node label.
[0026] As an example, the link feature code e j of each link means that each link has an L-dimensional link feature code E is the number of edges in the entire topology; Among them: each dimension corresponds to a state feature of a link, which is customized according to the network state and the service distribution requirement.
[0027] The link feature code includes: the total number of available wavelengths or frequency slots, the starting position of the first large available spectrum block, the size of the largest continuous available spectrum block, the link fragmentation entropy, and whether there is a spectrum block that meets the conditions.
[0028] As an example, the dynamic update means: according to the dynamic change of the Agent node selection, the v is updated in real time i and e j Whenever a new node joins the partial solution, the network state is adjusted accordingly, and the spectrum resource usage of each link is updated according to the current service request; by calculating the common available wavelength or frequency gap of the links on the selected path, the v is updated i and e j to ensure that the input state S of the Agent I reflects the latest network resource situation;
[0029] The partial solution refers to: a node sequence composed of nodes that have been selected by the decision-making Agent according to the current service request, but has not yet formed a complete route from the source node to the destination node.
[0030] Step 3. According to the input state S I the Agent filters invalid adjacent nodes through the action mask mechanism and selects one of the legal adjacent nodes as the action selection result, that is, the next-hop route, and calculates the reward function value;
[0031] According to the input state S of the Agent I the Agent selects one of the adjacent nodes of the current node as the action selection result and adds it to the partial solution; by means of an incremental iteration of selecting one hop at a time, a feasible solution that meets the optical path constraints is gradually constructed; at the same time, according to the wavelength or frequency gap usage of the link between the next-hop node and the current node and the existing common wavelength or frequency gap usage, the common available wavelength or frequency gap is calculated again to determine whether the common available wavelength or frequency gap meets the frequency gap requirements;
[0032] As an example, the model in the Agent adopts an edge feature enhanced graph neural network structure, and uses a multi-layer EGAT architecture to extract the complex coupling relationship between the link feature encoding and the node feature encoding, and capture the relationship and change law among the input state S I , the action selection result and the reward function value.
[0033] As an example, in the multi-layer edge feature enhanced graph attention network EGAT, the link feature encoding e j and the node feature encoding v i are respectively sent into the neural network as link features and node features for update; through the multi-layer edge feature enhanced graph attention network EGAT, the node features and link features are iteratively updated, and the information of adjacent nodes and links is combined to optimize the path selection.
[0034] As an example, the calculation formula of the reward function value R is designed as follows:
[0035]
[0036] Where: R is the reward value, which is the return in the current state; R terminal is the termination reward, which is a fixed reward value returned when the task is successfully completed or an exception occurs; D is the network diameter, which represents the maximum length of the shortest path between any two nodes in the network; P is the penalty value, which is a penalty value calculated according to the partial spectrum usage; S is the number of steps, which represents the number of steps that the current task has proceeded; C is the condition value, which calculates the relative progress of the current node from the target node, and a positive value indicates approaching the target, while a negative value indicates moving away from the target;
[0037] The calculation formula of the said condition value C is designed as follows:
[0038] C = μ · (γ · len(P last,dst ) - len(P current,dst ))
[0039] Where: μ: is the proportionality factor or scaling factor, which is used to adjust the influence degree of the potential energy function; it controls the contribution degree of the path length difference to the reward;
[0040] γ: is a constant, which is used to adjust the weight of the path length difference and is used as the discount factor for controlling the long-term return in reinforcement learning; in this formula, it is used to weight and calculate the shortest path length from the previous node to the target node; d curr,dst : the shortest path length from the current node to the target node.
[0041] Step 4: Repeat Step 2 and Step 3 until the termination condition of the current service is met, and judge whether the current service can be successfully routed according to the termination condition. If successful, select a frequency slot according to the First Fit strategy;
[0042] As an example, the said termination condition means: reaching the end point, having no selectable next-hop node; there is no common available wavelength or frequency slot that can meet the service bandwidth requirement.
[0043] As an example, the said First Fit strategy means: arranging the wavelengths or frequency slots in a certain order. Each time a wavelength or frequency slot is allocated, traverse all wavelengths or frequency slots in order and select the first available wavelength or frequency slot block that meets the service bandwidth requirement;
[0044] Step 5: Record the input state S I generated in each iteration, the action selection result, and the reward function value during the above process for Agent training;
[0045] As an illustration, through multiple rounds of training, the agent continuously adjusts its decisions; after each training, the cross-entropy loss function is used to calculate the difference between the predicted action and the actual action, and the weights of the model are updated through the backpropagation algorithm; during the training process, a learning rate scheduler is adopted to optimize the learning rate and improve the convergence speed of the model.
[0046] In a second aspect, the present application discloses an electronic device, which includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the method described in any of the above aspects.
[0047] In a third aspect, the present application discloses a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method described in any of the above aspects.
[0048] In a fourth aspect, the present application discloses a computer program product, when the instructions in the computer program product are executed by a processor of an electronic device, enabling the electronic device to execute the method described in any of the above aspects.
[0049] Advantages of the present invention:
[0050] The present invention overcomes the problem that the existing optical network resource scheduling method cannot balance topological generalization and global optimization in a large-scale heterogeneous network, and dynamically generates service routing and spectrum allocation strategies adapted to different network topologies through deep reinforcement learning.
[0051] The present invention first proposes a method for gradually constructing a solution space based on iterative learning to autonomously learn the optimal scheduling strategy of elastic optical networks and trains through reinforcement learning.
[0052] Innovatively, a separated node feature encoding and link feature encoding method is proposed, and the above features are further extracted by using an Edge-Enhanced Graph Attention Network, enhancing the model's perception ability of optical path constraints in the optical network, so as to more precisely consider the topological structure and spectrum resource constraints of the optical network during the resource scheduling process. Compared with the traditional path-predefined scheduling method, this method can effectively avoid path dependence and improve the accuracy and flexibility of scheduling.
[0053] It can more effectively handle the model retraining problem caused by topological changes in a large-scale optical network and enhance the actual deployment ability of the algorithm; compared with traditional heuristic algorithms, the present invention can better adapt to a large-scale dynamic network environment, reduce computational complexity and optimize the scheduling effect; BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a schematic flowchart of a large-scale optical network resource allocation method based on iterative learning according to the present invention.
[0055] Figure 2 This is a schematic structural diagram of an electronic device for a large-scale optical network resource allocation method based on iterative learning according to the present invention.
[0056] Figure 3 This is a schematic structural diagram of another electronic device for a large-scale optical network resource allocation method based on iterative learning according to the present invention. Detailed implementation manners
[0057] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application. Refer to Figures 1 to 3 as shown.
[0058] In a first aspect, a large-scale optical network resource allocation method based on iterative learning includes:
[0059] Step 1: Initialize the optical network topology, background traffic distribution, and current traffic requests;
[0060] As an example, the network topology includes: the number of nodes, the number of links, the topological connection relationship, and the number of available wavelengths or frequency slots on the links;
[0061] As an example, the background traffic distribution means: filling a certain number of random background traffic in the optical network, and the number of background traffic can be determined by the network load rate;
[0062] As an example, the current traffic requests include: source node, destination node, and bandwidth requirement.
[0063] Step 2: According to the current optical network state, calculate the node feature codes of all nodes and the link feature codes {v i}, {e j} of all links, and the mask vector M, which together constitute the input state S of the decision-making agent I ;
[0064] According to the topological structure relationship, the background traffic distribution situation, and the current traffic request status, dynamically update the node feature code v i of each node and the link feature code e j of each link; according to the topological situation, calculate a mask vector N is the current number of topological nodes;
[0065] As an example, the Agent refers to the agent of reinforcement learning.
[0066] As an example, the node feature encoding v of each node i means that each node has a K-dimensional node feature encoding N is the current number of topological nodes, where each dimension corresponds to a feature of a node;
[0067] As an example of application, the node feature encoding includes but is not limited to the following information: the number of services passing through the current node, the total number of time slots occupied by the services passing through the current node, the spectrum connectivity, the spectrum connection block, the shortest path length to the end point of the current service, and the node label.
[0068] As an example, the link feature encoding e of each link j means that each link has an L-dimensional link feature encoding E is the number of edges in the entire topology; where each dimension corresponds to a state feature of a link, which is customized according to the network state and service distribution requirements.
[0069] As an example of application, the link feature encoding includes but is not limited to the following information: the total number of available wavelengths or frequency slots, the starting position of the first large available spectrum block, the size of the largest continuous available spectrum block, the link fragmentation entropy, and whether there is a spectrum block that meets the conditions.
[0070] As an example, the dynamic update means that according to the dynamic changes of the Agent node selection, the v i , e j are updated in real time. Whenever a new node joins the partial solution, the network state is adjusted accordingly, and the usage of the spectrum resources of each link is updated according to the current service request; by calculating the common available wavelengths or frequency slots on the selected path, the v i , e j are updated to ensure that the input state S of the Agent I reflects the latest network resource situation;
[0071] As an example, the partial solution refers to a node sequence composed of the nodes that have been selected by the decision-making Agent according to the current service request, but has not yet formed a complete route from the source node to the destination node;
[0072] The specific scheme of the dynamic update operation is as follows:
[0073] Represent the source node and destination node of the current service request using node labels, where: in the node feature encoding of the source node and destination node, the node label is set to 1, and for other nodes, it is set to 0;
[0074] For the nodes in the partially obtained solutions that have been selected, set the node label in the node feature encoding of each node to 1, and the rest to 0;
[0075] According to the usage of spectrum resources on each link in the optical network, update the node feature encoding and link feature encoding, which includes: based on the partial paths obtained from the selected partial solutions, obtain their common available wavelengths or frequency gaps, and then take the intersection of the usage of this common available wavelength or frequency gap with the available wavelength or frequency gap usage on any link in the whole network, that is, only retain the wavelengths or frequency gaps that are available both in the common available wavelength or frequency gap and on the current link; use the intersection result as the remaining available wavelength or frequency gap on each link, and let it be the calculation basis for the link feature encoding and node feature encoding.
[0076] Calculate a mask vector according to the topology situation N is the number of current topology nodes; in the mask vector M, the positions of the nodes adjacent to the current node are set to 1, and the rest of the non - adjacent nodes are set to 0; all nodes located in the partial solutions are set to 0.
[0077] Step 3. According to the input state S I , the Agent filters out invalid adjacent nodes through the action mask mechanism and selects one of the legal adjacent nodes as the action selection result, that is, the next - hop route, and calculates the reward function value;
[0078] According to the input state S I of the Agent, the Agent selects one of the adjacent nodes of the current node as the action selection result and adds it to the partial solution; by means of an incremental iteration of selecting one hop at a time, gradually construct a feasible solution that meets the optical path constraints; at the same time, according to the usage of the link wavelength or frequency gap between the next - hop node and the current node and the existing common wavelength or frequency gap usage, calculate the common available wavelength or frequency gap again, and judge whether the common available wavelength or frequency gap meets the bandwidth requirements;
[0079] As an example, the model in the Agent adopts an edge - feature - enhanced graph neural network structure, uses a multi - layer EGAT architecture to extract the complex coupling relationship between the link feature encoding and the node feature encoding, and captures the relationship and change law among the input state S I , the action selection result, and the reward function value.
[0080] As an example, the EGAT is: Edge-Enhanced Graph Attention Network, that is, EGAT.
[0081] As an example, in the EGAT model, the link feature encoding e j and the node feature encoding v i are respectively fed into a neural network as link features and node features for updating; through the EGAT model, the node features and link features are iteratively updated, and the information of adjacent nodes and links is combined to optimize the path selection.
[0082] As an example, the EGAT model adopts a two-layer structure. After the link feature encoding e j of the link and the node feature encoding v i of the node are respectively feature-normalized, they are pre-extracted through a fully connected layer, and then input into the EGAT model for further feature interaction.
[0083] As an example, the calculation formula of the reward function value R in step three is designed as follows:
[0084]
[0085] where: R is the reward value, which is the return in the current state; R terminal is the termination reward, which is a fixed reward value returned when the task is successfully completed or an exception occurs; D is the network diameter, which represents the maximum length of the shortest path between any two nodes in the network; P is the penalty value, which is calculated according to the partial spectrum usage (such as the number of zero blocks); S is the number of steps, which represents the number of steps that the current task has proceeded; C is the condition value, which calculates the relative progress of the current node from the target node. A positive value indicates approaching the target, and a negative value indicates moving away from the target;
[0086] As an example, the calculation formula of the condition value C is designed as follows:
[0087] C = μ · (γ · len(P last.dst ) - len(P current.dst ))
[0088] where μ is the proportionality factor or scaling factor, which is used to adjust the influence degree of the potential energy function; it controls the contribution degree of the path length difference to the reward;
[0089] As an example, in practical applications, the μ can be adjusted according to the task requirements.
[0090] γ: A constant used to adjust the weight of the path length difference, which is used as the discount factor for controlling the long-term reward in reinforcement learning; in this formula, it is used to calculate the weighted shortest path length from the previous node to the target node; P current.dst : The shortest path length from the current node to the target node;
[0091] As an example, the termination condition refers to: reaching the end point, having no selectable next-hop route, or the common available wavelength or frequency gap not meeting the service bandwidth requirement.
[0092] Step 4. Repeat Step 2 and Step 3 until the termination condition of the current service is met. Determine whether the current service can be successfully routed according to the termination condition. If successful, select a wavelength or frequency gap according to the First Fit strategy; wavelengths and frequency gaps represent the allocable resources in wavelength-switched optical networks and elastic optical networks respectively. The present invention is not limited to a specific optical network type, and both types of optical networks can adopt this method for resource allocation.
[0093] As an example, the termination conditions include: having no selectable next-hop node; there is no longer a common available wavelength or frequency gap that can meet the service bandwidth requirement;
[0094] As an example, the First Fit strategy means: arranging the wavelengths or frequency gaps in a certain order (from large to small). Each time a wavelength or frequency gap is allocated, traverse all wavelengths or frequency gaps in order and select the first available wavelength or frequency gap block that meets the service bandwidth requirement;
[0095] Step 5. Record the input state S I generated in each iteration, the action selection result, and the reward function value during the above process for Agent training;
[0096] Through multiple rounds of training, the intelligent agent continuously adjusts its decisions; after each training, use the cross-entropy loss function to calculate the difference between the predicted action and the actual action, and update the weights of the model through the backpropagation algorithm; during the training process, adopt a learning rate scheduler to optimize the learning rate and improve the convergence speed of the model.
[0097] As an example, the Agent trained on a 24-node topology can directly perform result inference on a larger optical network topology and ensure the blocking rate performance without re-training or parameter fine-tuning, such as an optical network topology with a size of 150 nodes. The practical application effect of the trained Agent model in large-scale optical networks ensures that it can effectively solve the problem of spectrum resource allocation in large-scale optical networks.
[0098] In a second aspect, the present application discloses an electronic device, which includes: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in any of the above aspects.
[0099] In a third aspect, the present application discloses a non-transitory computer-readable storage medium, which when the instructions stored in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method described in any of the above aspects.
[0100] In a fourth aspect, the present application discloses a computer program product, which when the instructions in the computer program product are executed by a processor of an electronic device, enables the electronic device to execute the method described in any of the above aspects.
[0101] To better illustrate the design principle of the present invention, specific embodiments are exemplified as follows:
[0102] Embodiment 1;
[0103] For step 2, each node generates a 6-dimensional node feature encoding, which includes the number of services passing through the current node, the total number of wavelengths or frequency slots occupied by the services passing through the current node, the spectral connectivity, the spectral connection block, the shortest path length to the end point of the current service, and the node label. The node label is a boolean variable indicating whether the node is a start / end point or has already been in a partial solution. The spectral connectivity and spectral connection block are calculated as follows:
[0104]
[0105] SAB n : The average number of available spectral blocks of node n;
[0106] SC n : The average spectral connectivity of node n;
[0107] count: The total number of neighbor pairs considered;
[0108] c block (i, j): The number of consecutive zero blocks between node n and neighbors i and j;
[0109] c simi (i, j): The spectral similarity between node n and neighbors i and j;
[0110] For each link, a 6-dimensional state feature encoding is generated, recording the total number of available wavelengths or frequency slots, the starting position of the first large available wavelength or frequency slot, the size of the largest continuous available wavelength or frequency slot, the number of available wavelengths or frequency slots meeting the requirements, the link fragmentation entropy, and whether there is a wavelength or frequency slot block meeting the conditions, and its content includes:
[0111] ①Total number of available wavelengths or frequency slots (Total Zeros, T z ): The total number of available wavelengths or frequency slots on the link.
[0112] ②Starting position of the first large block of available wavelengths or frequency slots (First Large Zero Block Index, I fl ): The starting index position of the first available wavelength or frequency slot on the link that meets the specified size.
[0113] ③Size of the largest consecutive available wavelengths or frequency slots (Max Consecutive Zeros, M c ): The maximum length of consecutive available wavelengths or frequency slots on the link.
[0114] ④Number of available wavelengths or frequency slots that meet the requirements (Count of Large Zero Blocks, N b ): The number of available wavelengths or frequency slots on the link that meet the specified size requirements.
[0115] ⑤Link fragmentation entropy (Entropy Metric, H): Based on the distribution of wavelength or frequency slot occupancy and idle states, calculate the fragmentation entropy of the link, indicating the degree of fragmentation of the link's wavelengths or frequency slots.
[0116] ⑥Whether there is a frequency wavelength or frequency slot block that meets the conditions (Has Large Zero Block, F b ): Whether there is at least one available wavelength or frequency slot on the link that meets the specified conditions, used to quickly judge the availability of the link's wavelength or frequency slot resources.
[0117] It should be noted that for the method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by this application.
[0118] Optionally, the embodiments of this application also provide an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0119] Embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above method embodiments and can achieve the same technical effects. To avoid repetition, details are not described here again. Among them, the computer-readable storage medium includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0120] Figure 2 FIG. is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0121] Referring to Figure 2 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0122] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0123] The memory 804 is configured to store various types of data to support the operation of the device 800; examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0124] The power supply component 806 provides power for various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0125] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0126] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0127] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0128] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects of the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and a change in the temperature of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0129] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0130] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0131] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0132] Figure 3It is a block diagram of an electronic device 1900 shown in the present application. For example, the electronic device 1900 can be provided as a server.
[0133] Referring to Figure 3 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0134] The electronic device 1900 may also include a power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0135] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0136] The embodiments of the present application have been described above in conjunction with the accompanying drawings, but the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all belong to the protection scope of the present application.
[0137] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0138] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0139] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0140] If the described functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0141] The above are only the preferred embodiments of the present invention. It should be understood that the description of the above embodiments is only used to help understand the method and its core idea of the present invention, and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for allocating resources in a large-scale optical network based on iterative learning, characterized in that Including: Step 1, initialize the optical network topology, background service distribution, and current service request; Step 2: According to the current optical network state, calculate the node feature encodings of all nodes and the link feature encodings of all links {v i}, {e j}, and the mask vector M. The above together constitute the input state S of the decision-making Agent I ; Step 3: According to the input state S I , the Agent filters out invalid adjacent nodes through the action mask mechanism, selects one of the legal adjacent nodes as the action selection result, that is, the next-hop route, and calculates the reward function value; Step 4, repeat Step 2 and Step 3 until the termination condition of the current service is met. Determine whether the current service can be successfully routed according to the termination condition. If successful, select a wavelength or frequency slot according to the First Fit strategy; Step 5. In the above process, record the input state S I generated in each iteration, the action selection result, and the reward function value for Agent training.
2. The method for allocating large-scale optical network resources based on iterative learning according to claim 1, wherein The network topology includes: the number of nodes, the number of links, the topological connection relationship, and the number of available wavelengths or frequency slots on the links; the background service distribution refers to filling a certain number of random background services in the optical network, and the number of background services is determined by the network load rate; the current service request includes: source node, destination node, and bandwidth requirement.
3. A large-scale optical network resource allocation method based on iterative learning according to claim 1, characterized in that The process of calculating the node feature encodings of all nodes and the link feature encodings of all links {v i}, {e j}, and the mask vector M is as follows: According to the topological structure relationship, the background service distribution, and the current service request status, dynamically update the node feature encoding v of each node i and the link feature encoding e of each link j ; According to the topological situation, calculate a mask vector N is the number of current topological nodes; The Agent refers to the agent of reinforcement learning.
4. A method for allocating large-scale optical network resources based on iterative learning according to claim 3, characterized in that The node feature encoding v of each node i means that each node has a K-dimensional node feature encoding i ∈ {1, 2,..., N}, where N is the number of current topological nodes, and each dimension corresponds to a feature of the node; The node feature encoding includes: the number of services passing through the current node, the total number of time slots occupied by the services passing through the current node, spectral connectivity, spectral connectivity block, the shortest path length from the current service to the end point, and node label; The link feature encoding e of each link j means that each link has an L-dimensional link feature encoding j ∈ {1, 2,..., E}, where E is the number of edges in the entire topology; among them: each dimension corresponds to a link state feature, which is customized according to the network state and service distribution requirements; The link feature encoding includes: the total number of available wavelengths or frequency slots, the starting position of the first large available spectrum block, the size of the largest continuous available spectrum block, link fragmentation entropy, and whether there is a spectrum block that meets the conditions.
5. The method for allocating large-scale optical network resources based on iterative learning according to claim 3, characterized in that The dynamic update means that: according to the dynamic changes in the selection of Agent nodes, the v is updated in real time i and e j . Whenever a new node joins the partial solution, the network state is adjusted accordingly, and the usage of spectrum resources on each link is updated according to the current service request; by calculating the common available wavelengths or frequency gaps on the links of the selected path, the v i and e j are updated to ensure that the input state S of the Agent I reflects the latest network resource situation; The partial solution refers to: according to the current service request, a node sequence composed of the nodes that have been selected by the decision-making agent, but has not yet formed a complete route from the source node to the destination node.
6. A large-scale optical network resource allocation method based on iterative learning according to claim 1, characterized in that The model in the Agent adopts an edge feature-enhanced graph neural network structure, and uses a multi-layer edge feature-enhanced graph attention network EGAT architecture to extract the complex coupling relationship between the link feature encoding and the node feature encoding, and capture the input state S I , the relationship and variation law among the action selection result and the reward function value.
7. A method for allocating resources in a large-scale optical network based on iterative learning according to claim 6, characterized in that, In the multi-layer edge feature enhanced graph attention network EGAT, the link feature encoding e j and the node feature encoding v i are respectively used as link features and node features and fed into a neural network for update; Through the graph attention network EGAT enhanced by multi-layer edge features, the node features and link features are iteratively updated, and the information of adjacent nodes and links is combined to optimize the path selection.
8. A large-scale optical network resource allocation method based on iterative learning according to claim 1, characterized in that The calculation formula of the reward function value R in Step 3 is designed as follows: Where: R is the reward value, which is the return in the current state; R terminal is the termination reward, which is a fixed reward value returned when the task is successfully completed or an exception occurs; D is the network diameter, which represents the maximum length of the shortest path between any two nodes in the network; P is the penalty value, which is a penalty value calculated based on the partial spectrum usage; S is the number of steps, which represents the number of steps that the current task has taken; C is the condition value, which calculates the relative progress of the current node from the target node. A positive value indicates approaching the target, and a negative value indicates moving away from the target; The calculation formula of the condition value C is designed as follows: C = μ·(γ·len(P last,dst ) - len(P current,dst )) Where: μ: is a proportionality factor or scaling factor, used to adjust the influence degree of the potential energy function; it controls the contribution degree of the path length difference to the reward; γ: A constant used to adjust the weight of the path length difference, which is used as the discount factor for controlling the long-term return in reinforcement learning; in this formula, it is used to weight and calculate the shortest path length from the previous node to the target node; P current,dst : The shortest path length from the current node to the target node.
9. A large-scale optical network resource allocation method based on iterative learning according to claim 1, characterized in that The termination condition in Step 4 refers to: reaching the end point, having no selectable next-hop node; there is no longer a common available wavelength or frequency slot that can meet the service bandwidth requirement.
10. A large-scale optical network resource allocation method based on iterative learning according to claim 1, characterized in that Step 5 also includes: through multiple rounds of training, the agent continuously adjusts the decision; after each training, use the cross-entropy loss function to calculate the difference between the predicted action and the actual action, and update the weights of the model through the backpropagation algorithm; during the training process, use a learning rate scheduler to optimize the learning rate and improve the convergence speed of the model.
Citation Information
Patent Citations
FSO adaptive routing method based on deep reinforcement learning
CN116320841A
Resource allocation method based on deep reinforcement learning in space division multiplexing elastic optical network
CN116707698A
Network path computation method, device and system
WO2017004747A1