A Satellite Virtual Network Mapping Method Based on Deep Reinforcement Learning
By applying a virtual network mapping method with deep reinforcement learning in satellite networks, the problem of difficult to adapt to the dynamic nature of satellite networks and low resource allocation efficiency in the prior art is solved, and more efficient virtual network request processing and resource utilization are achieved.
Patent Information
- Application Number
- CN202211138369.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-09-19
AI Technical Summary
The prior art is difficult to adapt to the dynamics of satellite networks and cannot effectively allocate network resources for virtual network requests, resulting in low virtual network mapping efficiency.
The satellite virtual network mapping method based on deep reinforcement learning is adopted, and the satellite network mapping area is divided through sliding windows, and the DRL agent is used to dynamically map and resource allocation of nodes and links according to the environmental state and reward function.
It improves the acceptance rate of virtual network requests and the utilization rate of network resources, reduces the average delay, and achieves more efficient resource allocation.
Smart Images

Figure CN115550970B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of 6G satellite networks, and particularly to a satellite virtual network mapping method based on deep reinforcement learning. When the satellite network is used as an underlying communication service facility, deep reinforcement learning is used to allocate resources for virtual network requests. Background Art
[0002] With the wide deployment of 5G communication networks globally, the era of the Internet of Everything is approaching. While the scale of the Internet of Things (IOTs) is expanding day by day, huge data traffic and various service requests are also emerging, so a vast number of terrestrial networks are established to meet user needs. However, due to geographical environment limitations, some specific areas (such as the deep sea, polar regions, and mountains) cannot be equipped with traditional cellular base stations. In addition, due to deployment cost and technological limitations, it is unrealistic to widely establish terrestrial networks worldwide. Therefore, communication coverage and quality are key issues for future 6G communication networks.
[0003] In the above scenario, satellite networks can serve as a powerful supplement to terrestrial networks, providing global coverage and greater network capacity. Since satellites operate in space at a very high orbital altitude, they are not restricted by geography, and the corresponding infrastructure is not as vulnerable to natural disasters as terrestrial networks. Therefore, satellite networks have unique advantages in disaster emergency communication. In addition, satellite networks have lower power losses and shorter communication delays by using optical communication technology. According to the vision of 6G wireless network systems, satellite networks are expected to become an important part of future 6G Space-Ground Integrated Networks (SGIN).
[0004] Different from terrestrial networks, due to the relative motion of satellites in different orbits, Inter-Satellite Links (ISL) are periodically connected and the links are only established when in use. In addition, satellite nodes are also limited in terms of weight, volume, and energy consumption, which results in limited processing capabilities of satellites and the available resources they carry. Therefore, satellite networks require efficient network configuration and resource allocation schemes to complete various service requests.
[0005] Network virtualization is a key technology that abstracts underlying hardware resources into software, enabling multiple network application requests to share the same physical network resources. As the core technology of network virtualization, the Virtual Network Embedding (VNE) algorithm can make full use of underlying hardware resources, satisfy more Virtual Network Requests (VNRs), and ensure various Quality of Service (QoS) requirements and the revenue of network operators.
[0006] Since the VNE problem is an NP-hard problem, most existing studies have proposed many heuristic VNE algorithms. To reduce complexity, these algorithms have formulated a series of rules and constraints and cannot optimize themselves. Therefore, the solutions obtained by these heuristic algorithms may not be the most ideal, and the experimental results are not particularly convincing. In addition, most of them are designed based on static terrestrial networks and cannot adapt to the high dynamicity and specific constraints of satellite networks. For the satellite virtual network mapping problem, there is less existing research work and all the proposed mapping algorithms are heuristic. Therefore, new methods should be sought for the VNE problem on satellite networks.
[0007] With the development of machine learning and artificial intelligence, many intelligent algorithms have been used to solve network resource allocation problems. In particular, as the combination of Deep Learning (DL) and Reinforcement Learning (RL), the Deep Reinforcement Learning (DRL) method inherits its powerful perception ability and decision-making ability. By interacting with the environmental state, the DRL agent can perceive the underlying network characteristics and perform corresponding actions according to the reward function. To obtain the maximized cumulative reward, the DRL agent can dynamically optimize the strategy through training and self-learning. Therefore, the DRL method shows great potential in solving high-dimensional continuous decision-making problems. Inspired by these, this application introduces the DRL method to solve the VNE problem of satellite networks. Summary of the Invention
[0008] Aiming at the technical problem that the existing heuristic mapping methods for virtual network requests are difficult to adapt to the dynamicity of satellite networks and cannot meet the efficient allocation of network resources for virtual network requests in 6G satellite networks, the present invention proposes a satellite virtual network mapping method based on deep reinforcement learning. By interacting with the environment through reinforcement learning, the virtual nodes and virtual links in the virtual network request are mapped to appropriate satellite nodes and inter-satellite links, and the underlying network resources are reasonably allocated, improving the efficiency of resource allocation.
[0009] To achieve the above object, the technical solution of the present invention is implemented as follows: A satellite virtual network mapping method based on deep reinforcement learning, the steps are as follows:
[0010] Step 1: Model both the underlying satellite network and the virtual network request as undirected graphs. For an incoming virtual network request, use a sliding window method to divide the mapping area of the entire satellite network, and select the mapping area with the smallest load factor;
[0011] Step 2: Perform the node mapping process: Extract the feature matrix from the mapping area with the smallest load factor as the environmental state to input into the DRL agent, and use the DRL method to give the probability of a satellite node being mapped and select a satellite node for resource allocation;
[0012] Step 3: Determine whether the node mapping is successful. If it is not successful, the virtual network request mapping fails and jumps to Step 4; if the node mapping is successful, enter the link mapping process: Measure the inter-satellite link status according to the cost metric, and select an inter-satellite link for resource allocation of the virtual link according to the shortest path algorithm;
[0013] Step 4: Determine whether the link mapping is successful, and calculate the reward value using the reward function of the DRL agent;
[0014] Step 5: Determine whether it is time for batch update rounds. If it is not time for update rounds, repeat Steps 1 to 4; if it is time for update rounds, calculate the loss function through the reward value to update the network parameters and gradients of the DRL agent.
[0015] The modeling method for the underlying satellite network and the virtual network request is as follows:
[0016] Represent the underlying satellite network as an undirected graph G s (N s ,L s ), where N s represents the set of satellite nodes, and L s represents the set of inter-satellite links; for a satellite node the corresponding attribute set is where represents the total CPU resources owned by this satellite node , represents the available CPU resources of this satellite node , represents the number of virtual nodes already embedded in this satellite node ; for an inter-satellite link the corresponding attribute set is represents the total bandwidth resources owned by this inter-satellite link , Indicates the inter-satellite link The available bandwidth resources Indicates the inter-satellite link The average transmission delay Indicates the inter-satellite link The time required to establish Indicates the inter-satellite link The status, used to determine whether the inter-satellite link Has been established
[0017] Represents the virtual network request as an undirected graph G v (N v ,L v ), where N v And L v Respectively represent the virtual node set and virtual link set of the virtual network request; the CPU resource requirement of the virtual node Is represented as The bandwidth resource requirement of the virtual link Is represented as t s And t e Respectively represent the service start time and end time of the virtual network request
[0018] The method of the sliding window is: set the size of the sliding window to 3*3 and the sliding step to 1, and slide and divide the entire satellite network. Each slide forms a mapping area
[0019] The calculation method of the load factor is: the load factor Load of the resources in the kth mapping area k Is
[0020]
[0021] Among them Represents the satellite node set in the kth mapping area Is the satellite node set The ith satellite node in Represents the satellite node Available CPU resources Is the link set with the satellite node As the endpoints Is the jth inter-satellite link Represents the inter-satellite link The status Represents the inter-satellite link The total bandwidth resources owned Represents the inter-satellite link Available bandwidth resources, μ is the influence factor of the link status factor Represents the satellite node The maximum number of link connections that can be established represents a satellite node The remaining number of link connections that can be established
[0022] The DRL agent is set with a 4-layer neural network, namely the input layer, the convolutional layer, the normalization layer, and the output layer; in the input layer, a feature matrix extracted from the mapping area with the lowest load factor is used as the environmental state input, and then the convolutional layer obtains an n-dimensional vector through convolutional operations, representing the available resource state of the physical node; the normalization layer uses the softmax function to generate a probability vector of the mapped nodes; according to the mapping probability of the probability vector, different physical nodes are sequentially selected for each virtual node for resource allocation
[0023] The state space S of the DRL agent is a four-dimensional feature matrix: S=(F 1 ,F 2 ...,F n ) T ;
[0024] where n is the number of satellite nodes
[0025] The feature vector F i includes four network attributes and: F i ={C i ,B i ,D i ,L i};
[0026] where C i represents the CPU resource state after regularization, B i is the bandwidth resource after regularization, D i represents the degree of the node, and L i represents the load factor of the satellite node
[0027] Based on the input of the state space S, the DRL agent combines the convolutional layer and the normalization layer to generate a probability vector {P 1 ,P 2 ,...,P n}, where P i represents the mapping probability of the satellite node
[0028] The action space A of the DRL agent is: A=(E 1 ,E 2 ,...,E n ); E i represents whether the satellite node n i s is selected
[0029] The calculation formula for the reward function R of the DRL is as follows:
[0030]
[0031] Among them, the revenue function
[0032] The cost function
[0033] Among them, represents the number of hops of the virtual link mapped to the physical path, and L′ v represents the number of inter-satellite links that need to be newly established on the physical path, represents the time required for establishing the inter-satellite link.
[0034] The cost price The calculation formula is:
[0035]
[0036] Among them, α and β are the adjustment parameters of the risk coefficient and the bandwidth resource load; is the average delay of the inter-satellite link , is the risk coefficient of the inter-satellite link , is the bandwidth resource load.
[0037] The regularized CPU resource status C i is:
[0038] The regularized bandwidth resource B i is:
[0039] Among them, L is the set of links connected to the i-th satellite node , and S l represents the status of the inter-satellite link l;
[0040] The degree D of the node i is:
[0041] The satellite node The load factor L i is:
[0042] Among them, represents the number of virtual nodes that have been embedded, represents that the mapping paths of other virtual network requests pass through this satellite node The quantity;
[0043] The average delay represents as:
[0044] Wherein, represents the average transmission delay of the inter-satellite link, represents the state of the inter-satellite link, represents the time required to establish the inter-satellite link;
[0045] Risk coefficient as:
[0046] Wherein, represents the number of virtual links that have been embedded;
[0047] Bandwidth resource load as:
[0048] The loss function is:
[0049] Wherein, R γ is the cumulative reward in a batch and
[0050] Wherein, γ is the discount rate, P i represents the selection probability of the satellite node of, R i represents the reward value at the i-th time in a batch.
[0051] Compared with the prior art, the beneficial effects of the present invention are as follows: for the incoming virtual network requests, the entire satellite network is divided into multiple mapping areas and the resource status of each mapping area is evaluated. The virtual network mapping process will only be carried out in the mapping area with the lowest resource load; by extracting the satellite node attribute features in the mapping area with the lowest resource load to form a feature matrix, and using it as the environmental state to input to the intelligent agent of the DRL, the intelligent agent gives the probability that each satellite node can be embedded with a virtual node through neural network analysis, and selects a suitable node for mapping; when the node mapping is successful, the cost weight is used to measure each inter-satellite link, and the shortest path is selected to embed the virtual link. The present invention continuously updates the network through reinforcement learning, selects suitable satellite nodes and inter-satellite links to be allocated to different virtual network requests, uses the DRL method to automatically optimize the mapping strategy, improves the acceptance rate of virtual network requests, the utilization rate of network resources, and reduces the average delay. Brief Description of the Drawings
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0053] Figure 1 It is a schematic flowchart of the present invention.
[0054] Figure 2 It is a schematic diagram of the mapping area division of the present invention.
[0055] Figure 3 It is a schematic diagram of the deep learning method of the present invention.
[0056] Figure 4 It is the simulation result of the present invention. Among them, (a) is the acceptance rate, (b) is the resource utilization rate, and (c) is the average delay. Detailed implementation manners
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0058] As Figure 1 shown, a satellite virtual network mapping method based on deep reinforcement learning has the following steps:
[0059] Step 1: Model the underlying satellite network as an undirected graph G s , and model the virtual network request as an undirected graph G v ; for an incoming virtual network request G v , use a sliding window method to divide the mapping area of the entire satellite network, and select the mapping area with the smallest load factor for mapping the virtual network request G v .
[0060] For the convenience of representation, first model the underlying satellite network and the virtual network request. For the underlying satellite network, represent it as an undirected graph G s (N s , L s ), where N s represents the set of satellite nodes, and L s represents the set of inter-satellite links. For a satellite node its corresponding attribute set is where Indicates the total CPU resources owned by the satellite node Indicates the available CPU resources of the satellite node Indicates the satellite node The number of virtual nodes that have been embedded. For an inter-satellite link Its corresponding set of attributes is Indicates the inter-satellite link The total bandwidth resources it owns Indicates the inter-satellite link Available bandwidth resources Indicates the inter-satellite link Average transmission delay Indicates the inter-satellite link Time required to establish Indicates the inter-satellite link Status, used to determine whether the inter-satellite link has been established. If the inter-satellite link has been established, the corresponding status Is set to 0, otherwise set to 1
[0061] For a virtual network request, it is also represented as an undirected graph G v (N v ,L v ), where N v and L v Represent the virtual node set and virtual link set of the virtual network request respectively. The CPU resource requirement of a virtual node Is expressed as The bandwidth resource requirement of a virtual link Is expressed as In addition, t s and t e Are used to represent the service start time and end time of the virtual network request respectively
[0062] As Figure 2 Shown, a sliding window method is used for sliding division. The size of the sliding window is set to 3*3, and the sliding step is 1. Each sliding forms a mapping area. For these mapping areas, a load factor is used to measure the network resource status in the mapping area. For the k-th mapping area, the load factor Load k Of its corresponding resources is expressed as follows
[0063]
[0064] Among them Indicates the set of satellite nodes in the k-th mapping area Is the set of satellite nodes the $i$-th satellite node in is a set of links with the satellite node as the endpoint $\mathcal{L}_i$ is the $j$-th inter-satellite link. $\mu$ is the influence factor of the link state factor and is set to 0.5. denotes the maximum number of link connections that the satellite node can establish. denotes the remaining number of link connections that the satellite node can establish. The load factor comprehensively considers the available CPU resources, bandwidth resources, and link resources in this mapping area. Obviously, the lower the load factor Load k is, the better the resource state of this area is, and the higher the probability that the virtual network request can be successfully mapped in this area. Therefore, this application only maps the virtual network request in the mapping area with the lowest resource load factor.
[0065] Step 2: Perform the node mapping process: Extract the feature matrix from the mapping area with the smallest load factor as the environmental state input to the DRL agent, and use the DRL method to give the probability that the satellite node is mapped and select the satellite node for resource allocation.
[0066] The virtual network mapping process is divided into two steps: node mapping and link mapping. Only when both steps are completed can the request be considered successfully mapped. In the node mapping process, a DRL agent is set up, and the idea of the PG (Policy Gradient) reinforcement learning algorithm is adopted. As Figure 3 shown, the DRL agent sets up a four-layer neural network, namely the input layer, convolutional layer, normalization layer, and output layer. In the input layer, the feature matrix extracted from the mapping area with the lowest resource load factor is used as the environmental state input, and then an $n$-dimensional vector is obtained through convolutional operations, representing the available resource state of the physical node. The normalization layer uses the softmax function to generate the probability vector that the node is mapped. Finally, according to the mapping probability, several appropriate physical nodes are selected for resource allocation. For the DRL agent, it is necessary to set the environmental state space $S$, action space $A$, and reward function $R$.
[0067] State space $S$: Since the state of the underlying satellite network is very complex, the feature matrix is used as the state space of the DRL agent. The feature matrix contains the feature vectors of each satellite node in the mapping area with the lowest load factor, and each feature vector $F$ i includes four main network attributes, as follows:
[0068] $F$ i $ = \{C$ i , B i , D i , L i} (2)
[0069] Among them, represents the i-th satellite node. C i represents the regularized CPU resource status, and its corresponding formula is as follows:
[0070]
[0071] B i is the regularized bandwidth resource, and its calculation formula is as follows:
[0072]
[0073] Among them, L is the set of links connected to the i-th satellite node and S l represents the status of the inter-satellite link l.
[0074] D i represents the degree of the node. Higher connectivity means that more paths can be selected in the subsequent link mapping stage, and its calculation formula is as follows:
[0075]
[0076] L i represents the load factor of the satellite node . The number of virtual nodes that have been embedded and the number of mapping paths of other virtual network requests passing through this node are mainly considered, and are represented by and respectively:
[0077]
[0078] Among them, represents the number of virtual nodes that have been embedded in the satellite node , and represents the number of mapping paths of other virtual network requests passing through this satellite node .
[0079] The final state space S is a four-dimensional feature matrix:
[0080] S = (F 1 , F 2 ..., F n ) T (7)
[0081] Among them, n is the number of satellite nodes.
[0082] Action space A: Based on the input of the state space S, the DRL agent combines the convolutional layer and the normalization layer to generate a probability vector {P 1 , P2 ,..., P n}, where P i represents the mapping probability of the satellite node . The convolutional layer can obtain an n-dimensional vector through convolutional operations and can further extract features from the environmental state space; the normalization layer uses the softmax function to generate a probability vector for node mapping, which can make the mapping probability of each node distributed between 0 and 1, providing assistance for subsequent action selection. According to this probability vector, some appropriate satellite nodes will be selected for mapping, and E i represents whether the satellite node is selected. Therefore, the action space A can be represented as follows:
[0083] A = (E 1 , E 2 ,..., E n ) (8)
[0084] The state space S is a necessary prerequisite for action selection. The intelligent agent needs to analyze and extract the feature information of the state space to select physical nodes for each virtual node from the action space A for mapping. The selection strategy is to select with the help of the generated mapping probability. The mapping probability P i reflects the suitability of this physical node for mapping. Therefore, the higher the probability, the greater the possibility of the node being mapped. And the size of the action space is consistent with the number of nodes in the lowest resource-negative mapping area. The i-th physical node can be marked as selected or not selected by E i .
[0085] Reward function R: The reward function helps the DRL intelligent agent determine whether the executed plan can bring greater rewards. This application uses the benefit-cost ratio as the reward function, and the benefit function is calculated as follows:
[0086]
[0087] The cost function is calculated as follows:
[0088]
[0089] where represents the number of hops of the virtual link mapped to the physical path, and L′ v represents the number of new inter-satellite links that need to be established on the physical path. represents the time required for establishing the inter-satellite link.
[0090] The final calculation formula of the reward function R is as follows:
[0091]
[0092] It should be noted that the reward function is used to evaluate the entire mapping process. If one of the node mapping and link mapping fails, the obtained reward is 0.
[0093] Step 3: Determine whether the node mapping is successful. If it is not successful, the request mapping fails, and then jump to Step 4; if it is successful, enter the link mapping process: measure the inter-satellite link status according to the cost metric, and select an inter-satellite link using the shortest path algorithm for resource allocation of the virtual link.
[0094] In the link mapping process, this application uses the cost metric to measure the status of the inter-satellite link, comprehensively considering the average delay, risk coefficient, and bandwidth resource load of the link.
[0095] For an inter-satellite link Its average delay is expressed as It includes the link establishment delay and transmission delay, as follows:
[0096]
[0097] Among them, the establishment delay depends on the link status, represents the average transmission delay of the inter-satellite link, represents the status of the inter-satellite link.
[0098] Inter-satellite link The risk coefficient of represents the possible degree of failure of this link mapping. The higher the risk coefficient of the link, the fewer resources it may have, and thus the greater the risk of mapping failure. And
[0099]
[0100] Among them, represents the number of virtual links that have been mapped to this link. This variable is dynamically updated during implementation. If a virtual link of a certain virtual network request is mapped to this physical link, it is incremented by 1. When this virtual network request is completed and the resources are released, it is decremented by 1.
[0101] Bandwidth resource load represents the resource availability of the inter-satellite link. If the available bandwidth resource of the link is less, its bandwidth resource load will be higher, thus increasing the possibility of link mapping failure.
[0102]
[0103] Cost The calculation formula is defined as follows, where α and β are the adjustment parameters of the risk coefficient and the bandwidth resource load, and are generally set to 0.5.
[0104]
[0105] After measuring the link, the shortest path algorithm is used to select a path for resource allocation of the virtual link.
[0106] The cost metric comprehensively considers the resource load, risk, delay, etc. of the current physical link, and then measures the suitability of the link to be mapped. The shortest path algorithm calculates using the cost metric as the weight value to obtain a path with the minimum total cost.
[0107] Step 4: Determine whether both node mapping and link mapping are successfully completed. If one of the processes fails, the reward value is 0. Otherwise, calculate the reward value of the reward function according to formula (11).
[0108] Step 5: Determine whether it is time for batch update. If it is not time for the update round, repeat Steps 1 to 4. If it is time for the update round, calculate the loss function to update the network parameters and gradients.
[0109] Batch update can avoid updating the neural network every time it is trained. Frequent calculations will slow down the convergence process. The introduction of batches can reduce the computational pressure and converge faster.
[0110] The loss function is defined as follows:
[0111]
[0112] where R γ is the cumulative reward in a batch and
[0113]
[0114] where γ is the discount rate. P i represents the selection probability of the satellite node R i represents the reward value at the i-th time in a batch. The former is calculated by the neural network, and the latter is calculated by the reward function. These values are stored in the stack during the training process for easy extraction to calculate the loss function. Defining the loss function in this way can more comprehensively consider the cumulative reward value in a batch.
[0115] Calculate the gradient according to the loss function, and perform backpropagation with the set learning rate step size to update the network parameters. Figure 4For the simulation results, where DRL-LBVNE is the algorithm proposed in the present invention, and the rest are comparative algorithms. DQN is the Deep Q-Network, LPM is the Link Priority Mapping Algorithm, SN-VNE is the Satellite Virtual Network Mapping Algorithm, and Risk-VNE is the Virtual Network Mapping Algorithm Based on Risk Level. These three types of algorithms are all rule-based heuristic algorithms and do not adopt the idea of artificial intelligence algorithms. Through Figure 4 It can be seen that, compared with other comparative algorithms, the algorithm proposed in the present invention effectively improves the acceptance rate of requests and resource utilization rate, and at the same time ensures a relatively low average delay.
[0116] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A satellite virtual network mapping method based on deep reinforcement learning, characterized in that, the steps are as follows: Step 1: Model both the underlying satellite network and virtual network requests as undirected graphs. For an incoming virtual network request, use a sliding window method to divide the mapping area of the entire satellite network, and select the mapping area with the minimum load factor; Step 2: Perform the node mapping process: Extract the feature matrix from the mapping area with the minimum load factor as the environmental state input to the DRL agent, and use the DRL method to give the probability of satellite nodes being mapped and select satellite nodes for resource allocation; Step 3: Judge whether the node mapping is successful. If it is not successful, the virtual network request mapping fails, and jump to Step 4; If the node mapping is successful, enter the link mapping process: Measure the inter-satellite link status according to the cost metric, and select an inter-satellite link for resource allocation of the virtual link according to the shortest path algorithm; Step 4: Judge whether the link mapping is successful, and calculate the reward value using the reward function of the DRL agent; Step 5: Judge whether it is time for batch update rounds. If it is not time for update rounds, repeat Steps 1 to 4; If it is time for update rounds, calculate the loss function through the reward value to update the network parameters and gradients of the DRL agent.
2. The satellite virtual network mapping method based on deep reinforcement learning according to claim 1, characterized in that, the modeling method of the underlying satellite network and virtual network requests is: Represent the underlying satellite network as an undirected graph G s (N s ,L s ), where N s represents the set of satellite nodes, and L s represents the set of inter-satellite links; for a satellite node the corresponding attribute set is where represents the total CPU resources owned by this satellite node , represents the available CPU resources of this satellite node , represents the number of virtual nodes already embedded in this satellite node ; for an inter-satellite link the corresponding attribute set is represents the total bandwidth resources owned by this inter-satellite link , represents the available bandwidth resources of the inter-satellite link , represents the average transmission delay of the inter-satellite link , represents the time required to establish the inter-satellite link , represents the status of the inter-satellite link , used to determine whether the inter-satellite link has been established; Represent the virtual network request as an undirected graph G v (N v , L v ), where N v and L v represent the virtual node set and virtual link set of the virtual network request respectively; the CPU resource requirement of the virtual node is represented as the bandwidth resource requirement of the virtual link is represented as t s and t e represent the service start time and end time of the virtual network request respectively.
3. The satellite virtual network mapping method based on deep reinforcement learning according to claim 1 or 2, characterized in that, the sliding window method is: Set the size of the sliding window to 3*3 and the sliding step to 1 to slide and divide the entire satellite network, and each slide forms a mapping area.
4. The satellite virtual network mapping method based on deep reinforcement learning according to claim 3, characterized in that, The calculation method of the load factor is: the load factor Load of the resources in the k-th mapping area k is Among them, represents the set of satellite nodes in the k-th mapping area, is the i-th satellite node in the satellite node set , represents the available CPU resources of this satellite node ; is the set of links with the satellite node as the endpoint, is the j-th inter-satellite link; represents the state of the inter-satellite link , represents the total bandwidth resources owned by this inter-satellite link ; represents the available bandwidth resources of the inter-satellite link , and μ is the influence factor of the link state factor. represents the maximum number of link connections that the satellite node can establish, represents the remaining number of link connections that the satellite node can establish.
5. The satellite virtual network mapping method based on deep reinforcement learning according to claim 2 or 4, characterized in that, the DRL agent is set with a 4-layer neural network, namely the input layer, convolutional layer, normalization layer and output layer; In the input layer, extract the feature matrix from the mapping area with the lowest load factor as the environmental state input, and then the convolutional layer obtains an n-dimensional vector through convolutional operations, representing the available resource status of physical nodes; The normalization layer uses the softmax function to generate the probability vector of node mapping; Filter according to the mapping probability of the probability vector, and select different physical nodes for each virtual node in turn for resource allocation.
6. The satellite virtual network mapping method based on deep reinforcement learning according to claim 5, characterized in that, The state space S of the DRL agent is a four-dimensional feature matrix: S = (F 1 , F 2 ..., F n ) T ; where n is the number of satellite nodes; Feature vector F i comprises four network attributes and: F i = {C i , B i , D i , L i}; Among them, C i represents the regularized CPU resource status, B i is the regularized bandwidth resource, D i represents the degree of the node, L i represents the satellite node 's load factor; Based on the input of the state space S, the DRL agent combines the convolutional layer and the normalization layer to generate a probability vector {P 1 , P 2 ,..., P n}, where P i represents the mapping probability of the satellite node . The action space A of the DRL agent is: A = (E 1 , E 2 ,..., E n ); E i represents whether the satellite node is selected.
7. The satellite virtual network mapping method based on deep reinforcement learning according to claim 6, characterized in that, the calculation formula of the reward function R of the DRL is: Among them, the revenue function Cost function Among them, represents the hop count of the virtual link mapped to the physical path, L v′ represents the number of inter-satellite links that need to be newly established on the physical path, represents the time required for establishing the inter-satellite links.
8. The satellite virtual network mapping method based on deep reinforcement learning according to claim 6 or 7, characterized in that, The cost price The calculation formula is as follows: Among them, α and β are the adjustment parameters of the risk coefficient and the bandwidth resource load; is the inter-satellite link of the average delay, is the inter-satellite link of the risk coefficient, is the bandwidth resource load.
9. The satellite virtual network mapping method based on deep reinforcement learning according to claim 8, characterized in that, The regularized CPU resource status C i is as follows: The regularized bandwidth resource B i is as follows: where L is the set of links connected to the i-th satellite node and S l represents the state of the inter-satellite link l; Degree D of the node i is Satellite node The load factor L i is as follows: Among them, represents the number of virtual nodes that have been embedded, represents the number of mapping paths of other virtual network requests passing through this satellite node ; The average latency represents as follows: Among them, represents the average transmission delay of the inter-satellite link, represents the state of the inter-satellite link, represents the time required to establish the inter-satellite link; Risk coefficient is as follows: Among them, represents the number of virtual links that have been embedded; Bandwidth resource load is as follows:
10. The satellite virtual network mapping method based on deep reinforcement learning according to claim 8, characterized in that, The loss function is as follows: wherein, R γ is the cumulative reward in a batch and Among them, γ is the discount rate, P i represents the selection probability of the satellite node , and R i represents the reward value at the i-th time in a batch.