Virtual wireless sensor network resource allocation method, system, electronic device and computer-readable storage medium based on reinforcement learning
By adopting a resource allocation method based on reinforcement learning in the virtualized wireless sensor network, dynamically selecting physical nodes and wireless links, the problem of low resource allocation efficiency in the virtualized wireless sensor network is solved, and more efficient resource utilization and cost reduction are achieved.
Patent Information
- Application Number
- CN202210782127.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-30
AI Technical Summary
The resource allocation efficiency in existing virtualized wireless sensor networks is low, and network resources cannot be effectively utilized. The physical sensor network has limited energy and is difficult to bear multiple virtual sensor network requests.
Using reinforcement learning methods, we use the resource requirements of virtual wireless sensor network requests in advance, establish a deployment utility maximization model, select candidate physical nodes and wireless links from the underlying wireless sensor network, and use the Markov decision model and evolutionary algorithm meta-learning reinforcement learning algorithm to dynamically allocate resources to meet virtual network requests.
It improves the resource utilization of virtual sensor networks, reduces the deployment cost of service providers, and extends the life cycle of physical sensor networks, and can more effectively carry multiple virtual sensor network requests.
Smart Images

Figure CN115243377B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and more specifically, to a virtual wireless sensor network resource allocation method, system, electronic device and computer-readable storage medium based on reinforcement learning. Background Art
[0002] Currently, as a key component of the Internet of Things, Wireless Sensor Network (WSN) has been widely used in environmental monitoring, industrial production, smart medical care, smart cities and other fields. In recent years, WSN has combined a variety of information technologies in the fields of communication and computers, and can simultaneously monitor, process and communicate data, and has extended to new WSNs including intelligent sensor nodes, mobile sensor nodes and multimedia sensor nodes. However, traditional WSNs are mainly oriented to specific fields and established tasks, and their network deployment is tightly coupled with services. When the network is idle, it cannot be used for business requests from other users. For new user services, WSNs need to be redeployed, resulting in low network resource utilization.
[0003] In order to solve the "rigidity" problem of traditional WSNs, while also taking into account the possibility of a single sensor node serving different application tasks, the design trend of WSNs has begun to shift from independent, application-specific deployments to highly integrated wireless sensor systems that support a heterogeneous ecosystem of services and applications, allowing multiple application tasks to run on the same underlying infrastructure, and virtualization is an effective way to achieve this goal.
[0004] Virtualization allows actual physical resources to be abstracted into logical units, so that multiple application tasks can coexist on the same WSN infrastructure. Current WSN virtualization research is mainly divided into two categories: node-level virtualization and network-level virtualization. WSN node-level virtualization allows multiple applications to run their tasks concurrently on a single sensor node. The sensor node can basically become a multifunctional device, and its application tasks can be executed simultaneously or sequentially. WSN network-level virtualization supports virtual sensor networks, where the virtual sensor network is composed of a subset of sensor nodes in the physical sensor network. The subset is dedicated to a specific task at a given time, and the remaining sensor nodes can still be used for other application tasks or network functions. Multiple virtual sensor networks can coexist on the same WSN infrastructure.
[0005] As the number of IoT services continues to increase, resource allocation in virtualized WSNs has attracted widespread attention. Delgado C, Canales M, Ortin J proposed an optimization framework for joint admission control and network slicing in virtualized WSNs in “Joint Application Admission Control and Network Slicing in Virtual Sensor Networks” [IEEE Internet of Things Journal, 2017, 5(1): 28-43]. They considered the resource management problem of virtualized WSNs from the perspective of network infrastructure providers and introduced a joint optimization framework to solve the admission control of application tasks and WSN network slicing problems. Kaiwartya O, Abdullah AH, Cao Yue proposed a fault-tolerant optimization framework in virtualized WSNs in “Virtualization in Wireless Sensor Networks: Fault Tolerant Embedding for Internet of Things” [IEEE Internet of Things Journal, 2017, 5(1): 28-43]. They considered that the failure of communication links in WSNs would affect the network running IoT services, formulated a multi-objective optimization problem of maximizing fault tolerance and minimizing communication delay, and proposed a genetic algorithm solution based on non-dominated sorting. Li Mingyan, Chen Cailian, and Hua Cunqing proposed an intelligent latency-aware virtual network mapping scheme for industrial WSN in "Intelligent Latency-Aware Virtual Network Embedding for Industrial Wireless Networks" [IEEE Internet of Things Journal, 2019, 6(5): 7484-7496]. They used arbitrary path routing technology to map virtual links and utilized the broadcast characteristics of wireless channels to reduce the resources consumed by retransmission. In addition, in order to meet the resource and latency requirements of the virtual network, deep Q learning is used to provide intelligent latency sensing and timely forwarding adjustments to address the dynamic changes in link quality and network workload.Katona R, Cionca V, and Orshea D proposed a QoS- and QoI-aware WSN virtual resource mapping algorithm in "Virtual Network Embedding for Wireless Sensor Networks TimeEfficient QoS / QoI Aware Approach" [IEEE Internet of Things Journal, 2021, 8(2): 916-926], which takes into account the perception accuracy of sensor nodes, the reliability of communication paths, and the impact of multi-access link interference on the overall consumption of communication resources. The proposed algorithm adopts an offline mapping method, processes the resource requirements of the virtual network requests in sequence, and selects the best physical resources in the node mapping and link mapping stages respectively.
[0006] Although the above virtualized WSN resource allocation algorithms have been studied from different perspectives such as cost-benefit, service quality and energy efficiency, there are many shortcomings in the existing work. For example, due to the random arrival and departure of virtual sensor network requests, the underlying physical sensor network resources also change. Therefore, it is necessary to reasonably select virtual sensor network requests according to resource conditions and network status to allocate resources for them. Secondly, the energy of physical sensor networks is limited. How to balance network energy consumption while deploying virtual sensor network requests and extend the network life cycle to carry more subsequent virtual sensor network requests requires further research.
[0007] Therefore, how to provide a virtual wireless sensor network resource allocation method, system, electronic device and computer-readable storage medium based on reinforcement learning has become a technical problem that urgently needs to be solved in this field. Summary of the invention
[0008] The purpose of the present invention is to provide a virtual wireless sensor network resource allocation method, system, electronic device and computer-readable storage medium based on reinforcement learning.
[0009] The first aspect of the present invention discloses a virtual wireless sensor network resource allocation method based on reinforcement learning; the method comprises:
[0010] Step S1, pre-acquire the resource requirements of the virtual wireless sensor network request arriving at any time;
[0011] Step S2: selecting a candidate physical node set from the underlying wireless sensor network according to the required resource requirements and a pre-established deployment utility maximization model, and obtaining a corresponding candidate wireless link set according to the candidate physical node set, wherein the deployment utility maximization model is pre-established based on the sensing information quality requirements and resource requirement constraints;
[0012] Step S3, substituting the candidate physical node set and the candidate wireless link set into the Markov decision model, and using the reinforcement learning algorithm to select a candidate physical node from the candidate physical node set as the best physical node, and selecting a corresponding candidate wireless link from the candidate wireless link set as the best wireless link based on the best physical node, so as to request allocation of corresponding communication bandwidth and sensor node computing resources for the virtual wireless sensor network.
[0013] According to the method of the first aspect of the present invention, after step S3, the method further comprises:
[0014] Step S4: Map the virtual node to the best physical node selected by the reinforcement learning algorithm.
[0015] According to the method of the first aspect of the present invention, in step S4, if the virtual node is successfully mapped to the best physical node, a mapping success prompt is output; if the virtual node is not successfully mapped to the best physical node, the best physical node is removed from the candidate physical node set, and the mapping number of the virtual node is determined. If the mapping number is greater than a preset number threshold, a mapping failure prompt is output; if the mapping number is less than the preset number threshold, the process proceeds to step S3 again.
[0016] According to the method of the first aspect of the present invention, in step S2, the objective function of the deployment utility maximization model is expressed as:
[0017] max U(t)---(1);
[0018]
[0019] here, Indicates the location coordinates of the area of interest in the business request, c re (n s ) represents the node processing capacity threshold, b re (l s ) represents the threshold of system resources, b(l v ) represents the bandwidth resource demand of the virtual link. In addition, Represents the quality requirement of sensing information. is the Euclidean distance between the sensing area locations, Request distance threshold for the service.
[0020] Furthermore, c1 indicates that the processing resources occupied by the virtual sensor node on any physical sensor node should be less than the processing resources available to the physical sensor node; c2 indicates that the bandwidth resources occupied by the virtual link should be less than the bandwidth resources of its corresponding physical link; c3 indicates the quality requirement of the sensing information; c4 and c5 indicate the value range constraints of the variables; the objective function is U(t)=R(t)-C(t), where R(t) and C(t) represent the benefit and cost, respectively.
[0021] According to the method of the first aspect of the present invention, in step S3, the virtual wireless sensor network request is defined as a multi-tuple M=<S, A, P, R> through a Markov decision model;
[0022] Where S is the state space, A is the action space, P is the transition probability, R is the reward function, and s t ∈S is the system state at time t, including the arrival and departure of virtual wireless sensor network requests, the resource requirements of virtual wireless sensor network requests, and the available resources of the current underlying wireless sensor network; action a t ∈A is based on the current network status, sensor node processing and link resource allocation; in state s t Execute action a t After that, the deployment of the virtual wireless sensor network request at the current moment is completed, and an immediate response R is obtained. t =U(t); state transition probability pr(s t+1 ∈S t+1 |s t ,a t ), indicating that in state s t Take action a t Transfer to t+1 probability.
[0023] According to the method of the first aspect of the present invention, in step S3, the reinforcement learning algorithm is a meta-learning reinforcement learning algorithm combined with an evolutionary algorithm. By designing an inner and outer loop, the outer loop mainly performs algorithm performance evaluation, and selects the batch of algorithms with the largest reward returns from a pile of candidate algorithms; while the inner loop mainly performs algorithm screening, and after each algorithm is trained to the best state, the outer loop is used to perform performance evaluation, so as to automatically select the optimal virtual wireless sensor network resource allocation algorithm.
[0024] The second aspect of the present invention discloses a virtual wireless sensor network resource allocation system based on reinforcement learning; the system comprises:
[0025] The first processing module is configured to pre-acquire the resource requirements of the virtual wireless sensor network request arriving at any time;
[0026] A second processing module is configured to select a set of candidate physical nodes from the physical network according to the required resource requirements and a pre-established deployment utility maximization model, and obtain a corresponding set of candidate wireless links according to the set of candidate physical nodes, wherein the deployment utility maximization model is pre-established based on the sensing information quality requirements and the resource requirement constraints;
[0027] The third processing module is configured to substitute the candidate physical node set and the candidate wireless link set into the Markov decision model, and use the reinforcement learning algorithm to select a candidate physical node from the candidate physical node set as the best physical node, and select a corresponding candidate wireless link from the candidate wireless link set as the best wireless link according to the best physical node, so as to allocate corresponding communication bandwidth and sensor node computing resources for the virtual wireless sensor network request.
[0028] According to the system of the second aspect of the present invention, the system further comprises:
[0029] The fourth processing module is configured to map the virtual node to the best physical node selected by the reinforcement learning algorithm.
[0030] According to the system of the second aspect of the present invention, the fourth processing module is specifically configured to output a mapping success prompt if the virtual node is successfully mapped to the best physical node; if the virtual node is not successfully mapped to the best physical node, remove the best physical node from the candidate physical node set, and determine the number of mappings of the virtual node. If the mapping number is greater than a preset number threshold, output a mapping failure prompt; if the mapping number is less than the preset number threshold, re-enter the third processing module.
[0031] According to the system of the second aspect of the present invention, in the second processing module, the objective function of the deployment utility maximization model is expressed as:
[0032] max U(t)---(1);
[0033]
[0034] here, Indicates the location coordinates of the area of interest in the business request, c re (n s ) represents the node processing capacity threshold, b re (l s ) represents the threshold of system resources, b(l v ) represents the bandwidth resource demand of the virtual link. In addition, Represents the quality requirement of sensing information. is the Euclidean distance between the sensing area locations, Request distance threshold for the service.
[0035] Furthermore, c1 indicates that the processing resources occupied by the virtual sensor node on any physical sensor node should be less than the processing resources available to the physical sensor node; c2 indicates that the bandwidth resources occupied by the virtual link should be less than the bandwidth resources of its corresponding physical link; c3 indicates the quality requirement of the sensing information; c4 and c5 indicate the value range constraints of the variables; the objective function is U(t)=R(t)-C(t), where R(t) and C(t) represent the benefit and cost, respectively.
[0036] According to the system of the second aspect of the present invention, in the third processing module, the virtual wireless sensor network request is defined as a multi-tuple M=<S, A, P, R> through a Markov decision model;
[0037] Where S is the state space, A is the action space, P is the transition probability, R is the reward function, and s t ∈S is the system state at time t, including the arrival and departure of virtual wireless sensor network requests, the resource requirements of virtual wireless sensor network requests, and the available resources of the current underlying wireless sensor network; action a t ∈A is based on the current network status, sensor node processing and link resource allocation; in state s t Execute action a t After that, the deployment of the virtual wireless sensor network request at the current moment is completed, and an immediate response R is obtained. t =U(t); state transition probability pr(s t+1 ∈S t+1 |s t ,a t ), indicating that in state s t Take action a t Transfer to t+1 probability.
[0038] According to the system of the second aspect of the present invention, in the third processing module, the reinforcement learning algorithm is a meta-learning reinforcement learning algorithm combined with an evolutionary algorithm. By designing an inner and outer loop, the outer loop mainly performs algorithm performance evaluation, and selects the batch of algorithms with the largest reward returns from a bunch of candidate algorithms; while the inner loop mainly performs algorithm screening, and after each algorithm is trained to the best state, the outer loop is used to perform performance evaluation, so as to automatically select the optimal virtual wireless sensor network resource allocation algorithm.
[0039] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps in any one of the virtual wireless sensor network resource allocation methods based on reinforcement learning in the first aspect of the present disclosure are implemented.
[0040] The fourth aspect of the present invention discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the virtual wireless sensor network resource allocation methods based on reinforcement learning in the first aspect of the present disclosure are implemented.
[0041] According to the technical content disclosed in the present invention, the following beneficial effects are achieved:
[0042] The solution proposed in the present invention aims at the problem of effective utilization of virtualized wireless sensor network resources. From the perspective of service requests of virtualized wireless sensor networks, a virtualized wireless sensor network resource allocation strategy based on reinforcement learning is proposed. First, a virtual sensor network request deployment utility maximization model based on business sensing information quality requirements and resource capacity constraints is established; then, considering that business requests are random and the network environment changes dynamically, its resource optimization problem is converted into a Markov decision process, and finally, a meta-learning reinforcement learning algorithm based on an evolutionary algorithm is used to continuously interact with the virtualized wireless sensor network environment to obtain the optimal resource allocation strategy; this method effectively improves the service provider's revenue and reduces the virtual sensor network request deployment cost while meeting resource requirements.
[0043] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0045] Figure 1 A virtual wireless sensor network resource allocation method based on reinforcement learning is provided according to an embodiment of the present invention. Figure 1 ;
[0046] Figure 2 A virtual wireless sensor network resource allocation method based on reinforcement learning is provided according to an embodiment of the present invention. Figure 2 ;
[0047] Figure 3 A reinforcement learning algorithm framework for meta-learning of evolutionary algorithms;
[0048] Figure 4A specific implementation flow chart of a virtual wireless sensor network resource allocation method based on reinforcement learning provided according to an embodiment;
[0049] Figure 5 The structure of a virtual wireless sensor network resource allocation system based on reinforcement learning according to an embodiment of the present invention is Figure 1 ;
[0050] Figure 6 The structure of a virtual wireless sensor network resource allocation system based on reinforcement learning according to an embodiment of the present invention is Figure 2 ;
[0051] Figure 7 The figure is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0052] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention unless otherwise specifically stated.
[0053] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.
[0054] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.
[0055] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0056] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0057] Embodiment 1:
[0058] The invention discloses a virtual wireless sensor network resource allocation method based on reinforcement learning. Figure 1 FIG. 1 is a flowchart of a virtual wireless sensor network resource allocation method based on reinforcement learning according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0059] Step S1, pre-acquire the resource requirements of the virtual wireless sensor network request arriving at any time;
[0060] Step S2: selecting a candidate physical node set from the underlying wireless sensor network according to the required resource requirements and a pre-established deployment utility maximization model, and obtaining a corresponding candidate wireless link set according to the candidate physical node set, wherein the deployment utility maximization model is pre-established based on the sensing information quality requirements and resource requirement constraints;
[0061] Step S3, substituting the candidate physical node set and the candidate wireless link set into the Markov decision model, and using the reinforcement learning algorithm to select a candidate physical node from the candidate physical node set as the best physical node, and selecting a corresponding candidate wireless link from the candidate wireless link set as the best wireless link according to the best physical node, so as to allocate corresponding communication bandwidth and sensor node computing resources for the virtual wireless sensor network request.
[0062] In step S2, a set of candidate physical nodes is selected from the underlying wireless sensor network according to the required resource requirements and a pre-established deployment utility maximization model, and a corresponding set of candidate wireless links is obtained according to the candidate physical node set, wherein the deployment utility maximization model is pre-established based on the sensing information quality requirements and resource requirement constraints.
[0063] In some embodiments, in step S2, the objective function of the deployment utility maximization model is expressed as:
[0064] max U(t)---(1);
[0065]
[0066] here, Indicates the location coordinates of the area of interest in the business request, c re (n s ) represents the node processing capacity threshold, b re (l s ) represents the threshold of system resources, b(l v ) represents the bandwidth resource demand of the virtual link. In addition, Represents the quality requirement of sensing information. is the Euclidean distance between the sensing area locations, Request distance threshold for the service.
[0067] Furthermore, c1 indicates that the processing resources occupied by the virtual sensor node on any physical sensor node should be less than the processing resources available to the physical sensor node; c2 indicates that the bandwidth resources occupied by the virtual link should be less than the bandwidth resources of its corresponding physical link; c3 indicates the quality requirement of the sensing information; c4 and c5 indicate the value range constraints of the variables; the objective function is U(t)=R(t)-C(t), where R(t) and C(t) represent the benefit and cost, respectively.
[0068] Specifically, the system architecture of the embodiment of the present invention is a virtualized WSN with a single Sink node. The underlying physical network is a WSN composed of a group of sensor nodes with limited perception, processing and communication resources, wherein the WSN controller collects and maintains the network status, where the network status includes the topology of the WSN, node location, logical connectivity, and available node and link resources. For the virtual sensor network request (VSNR) submitted by the user, the WSN controller deploys its request service in the underlying physical network according to its resource requirements and service quality, and allocates resources to it to meet the VSNR service requirements.
[0069] The underlying physical network WSN consists of a weighted undirected graph G s =(N s ,L s ) indicates that N s represents the set of WSN sensor nodes, L s Represents the set of links between WSN sensor nodes. For each WSN sensor node represents the processing capability of the sensor node, and the location information is Represents; for each WSN link l s ∈L s , the available link capacity is given by b(l s )express.
[0070] Virtual sensor network requests are represented by a weighted directed graph G v =(N v ,L v ) indicates that N v and L v They represent the virtual sensor nodes and virtual link sets in the virtual sensor network respectively. For each virtual sensor node Indicates its processing capacity requirements, Indicates the location of the area of interest for the service request. For each virtual link l v ∈L v , b(l v ) represents the bandwidth resource requirement of the virtual link. In addition, each VSNR has a corresponding service quality requirement, that is, the sensing information quality requirement
[0071] Generally, the quality of the sensed information can be modeled using different parameters, such as accuracy, frequency, freshness, and effectiveness. Represented as sensor node Geolocation and Virtual Sensor Nodes Request the Euclidean distance between sensor region locations And the service request distance threshold is inversely proportional to each other, that is
[0072] In a virtualized wireless sensor network environment, VSNSP leases sensor node processing resources and link bandwidth resources from WSNInP according to user requests, deploys VNSR, and thus provides services to users and obtains benefits. Therefore, the benefits of VSNSR are closely related to the user's resource demand. In the embodiment of the present invention, the benefits of resource allocation are modeled as a non-decreasing convex function, expressed as:
[0073]
[0074] Among them, c(n v ) and b(l v ) represent the processing resources and link bandwidth resources allocated by the underlying physical network to VSNR respectively.
[0075] The VSNR deployment cost is generated by the overhead required by VSNSP to lease resources from WSNInP, including the physical sensor node processing resource cost and link bandwidth resource cost, which is expressed as:
[0076]
[0077] in, Represents the physical sensor node n s Is it a virtual sensor node n v Allocate processing resources, if n s To n v Allocate processing resources, then otherwise Indicates physical link l s Is it a virtual sensor link? v Allocate bandwidth resources, if l s To l v Allocate bandwidth resources. otherwise α(n s ) represents the physical sensor node n s The unit price of processing resources, β(l s ) indicates physical link l s The unit price of bandwidth resources, α(n s ) and β(l v ) is expressed as:
[0078]
[0079]
[0080] Among them, c re (n s ) is the processing resource currently available at the physical sensor node, b re (l s ) is the currently available bandwidth resource of the physical link, η and δ are weight coefficients, which are used to adjust the unit price of processing resources and bandwidth resources. For any physical sensor node or physical link, the larger the load, the smaller the available resources, and the higher the corresponding unit resource cost, which reflects the network load and ensures efficient and balanced utilization of physical resources.
[0081] In summary, the utility U(t) obtained by VSNSP deploying VNSR at time t is expressed as:
[0082] U(t)=R(t)-C(t)---(7);
[0083] The optimization goal of the present invention is to maximize the VSNSP utility under the premise of satisfying the virtual sensor network request resource and service quality constraints. The objective function is expressed as:
[0084] max U(t)---(8);
[0085]
[0086] here, Indicates the location coordinates of the area of interest in the business request, c re (n s ) represents the node processing capacity threshold, b re (l s ) represents the threshold of system resources, b(l v ) represents the bandwidth resource demand of the virtual link. In addition, Represents the quality requirement of sensing information. is the Euclidean distance between the sensing area locations, Request distance threshold for the service.
[0087] Furthermore, c1 indicates that the processing resources occupied by the virtual sensor node on any physical sensor node should be less than the processing resources available to the physical sensor node; c2 indicates that the bandwidth resources occupied by the virtual link should be less than the bandwidth resources of its corresponding physical link; c3 indicates the quality requirement of the sensing information; c4 and c5 indicate the value range constraints of the variables; the objective function is U(t)=R(t)-C(t), where R(t) and C(t) represent the benefit and cost, respectively.
[0088] In step S3, the candidate physical node set and the candidate wireless link set are substituted into the Markov decision model, and a reinforcement learning algorithm is used to select a candidate physical node from the candidate physical node set as the best physical node, and a corresponding candidate wireless link is selected from the candidate wireless link set as the best wireless link according to the best physical node, so as to allocate corresponding communication bandwidth and sensor node computing resources for the virtual wireless sensor network request.
[0089] In some embodiments, in step S3, the virtual wireless sensor network request is defined as a multi-tuple M=<S, A, P, R> through a Markov decision model;
[0090] Where S is the state space, A is the action space, P is the transition probability, R is the reward function, and s t ∈S is the system state at time t, including the arrival and departure of virtual wireless sensor network requests, the resource requirements of virtual wireless sensor network requests, and the available resources of the current underlying wireless sensor network; action a t ∈A is based on the current network status, sensor node processing and link resource allocation; in state s t Execute action a t After that, the deployment of the virtual wireless sensor network request at the current moment is completed, and an immediate response R is obtained. t =U(t); state transition probability pr(s t+1 ∈S t+1 |s t ,a t ), indicating that in state s t Take action a t Transfer to t+1 probability.
[0091] In step S3, the reinforcement learning algorithm is a meta-learning reinforcement learning algorithm combined with an evolutionary algorithm. By designing an inner and outer loop, the outer loop mainly performs algorithm performance evaluation, and selects the algorithm with the largest reward return from a bunch of candidate algorithms; while the inner loop mainly performs algorithm screening, and after each algorithm is trained to the best state, the outer loop is used for performance evaluation. Finally, the system can automatically select the optimal virtual wireless sensor network resource allocation algorithm.
[0092] Specifically, the embodiment of the present invention aims at the scenario where the VSNR is random and the network status changes dynamically, and establishes the resource allocation problem of maximizing the utility of VSNR deployment as a Markov decision process (MDP). MDP is defined as a multi-tuple M=<S,A,P,R>, where S is the state space, A is the action space, P is the transition probability, and R is the reward function.t ∈S is the system status at time t, including the arrival and departure of VSNR, the resource demand of VSNR and the available resources of the current underlying WSN. Action a t ∈A is the sensor node processing and link resource allocation according to the current network status. t Execute action a t After that, the VSNR deployment at the current moment is completed, and the system will get an immediate response R t =U(t). State transition probability pr(s t+1 ∈S t+1 |s t ,a t ), indicating that in state s t Take action a t Transfer to t+1 probability.
[0093] Generally, reinforcement learning agents evaluate and improve strategies based on value functions, where the value function is the cumulative reward function obtained by following the strategy, and the state-action value function is expressed as the expected value of the cumulative reward starting from the current state and taking an action, and then using a given strategy to select an action:
[0094]
[0095] Where E{·} represents the expected value. β∈(0,1) is the discount factor used to measure current or future decisions. π The iterative form of the (s,a) function is the Bellman Equation, as shown below:
[0096] Q π (s,a)=E{R t +βQ π (s t+1 ,a t+1 )}---(11);
[0097] For finite-state MDPs, the value function is usually stored in a lookup table and can be learned recursively. However, for continuous state and action spaces, it is not practical to compute and store the value function for each specific state-action pair. Therefore, the goal of reinforcement learning is to find a policy π that maximizes the following objective function:
[0098] J(π)=E π {Q π (s,a)}=∫ S d π (s)∫ A π(a|s)Q π(s,a)dads---(12);
[0099] where d π (s) is the state distribution function under strategy π.
[0100] The strategy found here through reinforcement learning maximizes the function J, that is, R(t) is maximized in the deployment utility maximization model.
[0101] The embodiment of the present invention adopts a meta-learning reinforcement learning algorithm based on an evolutionary algorithm to solve the MDP decision problem, see Figure 3 As shown in Figure 2. The inner loop method Eval(L,ε) evaluates the reward return of the reinforcement learning algorithm L on a given environment ε. The goal of the outer loop optimization is to select the reinforcement learning algorithm environment with high training return in a set of training environments.
[0102] The strategy used by the agent can be parameterized as π θ (a t |s t ), indicating that the agent performs action a in the environment ε in each time slot t , and get reward r t , then the environment switches to s t+1 , θ is the parameter of the Q-value function, and the strategy π θ is obtained from the Q-value function using the ε-greedy method. Finally, the agent saves the state transition probability (s t ,s t+1 ,a t ,r t ) to the experience replay buffer D and continuously update the strategy L(s) by minimizing the loss function through the gradient descent algorithm t ,a t ,r t ,s t+1 ,θ,γ), the training will be carried out for M rounds. In each round, the reward function obtained by the agent can be expressed as The average training return after normalization of the algorithm's performance in a given environment is expressed as Where R max and R min represents the maximum and minimum rewards returned by the environment. The process of this inner loop evaluation is outlined in Algorithm 1. To improve the efficiency of the algorithm, the algorithm is scored using the normalized average training return instead of the final behavior policy return. The goal of the meta-learner is to find the optimal loss function L(s t ,a t ,r t ,s t+1 ,θ,γ) to optimize the strategy π θ, which has the largest normalized average training return on the set of training environments. The complete optimization objective of the meta-learner is:
[0103]
[0104] The specific steps of the two algorithms are as follows:
[0105]
[0106]
[0107]
[0108] Figure 2 FIG. 4 is another flow chart of a virtual wireless sensor network resource allocation method based on reinforcement learning according to an embodiment of the present invention. Figure 2 As shown, the method also includes step S4: mapping the virtual node to the best physical node selected by the reinforcement learning algorithm.
[0109] In some embodiments, if the virtual node is successfully mapped to the best physical node, a mapping success prompt is output; if the virtual node is not successfully mapped to the best physical node, the best physical node is removed from the candidate physical node set, and the number of mappings of the virtual node is determined. If the mapping number is greater than a preset number threshold, a mapping failure prompt is output; if the mapping number is less than the preset number threshold, the process returns to step S3.
[0110] Specifically, in the embodiment of the present invention, the resources required by the virtual sensor network request arriving at any time are first known, and then the candidate sensor nodes and wireless link sets that meet the virtual sensor network request are selected according to the resource constraints and information quality constraints. Finally, the candidate sensor nodes and wireless link sets are substituted into the Markov decision process, and the meta-learning reinforcement learning algorithm based on the evolutionary algorithm is used to obtain the optimal resource allocation scheme. For the specific process, see Figure 4 As shown, after selecting the best physical node according to the meta-learning reinforcement learning method of the evolutionary algorithm, the virtual node needs to be mapped to the physical node, that is, the best physical node. If it is determined that the virtual node is successfully mapped to the physical node, the mapping is successful. Otherwise, the currently mapped physical node (that is, the best physical node) is removed from the candidate physical node set, and it is determined whether the number of virtual node mappings exceeds K times (k is a preset number threshold, k is an integer greater than 1, and the value is determined according to actual needs and is not specifically limited here). If it exceeds k times, the mapping fails, otherwise it returns to step S3 and selects the next best physical node again.
[0111] In summary, the solution proposed in the present invention aims at the problem of effective utilization of virtualized wireless sensor network resources. From the perspective of service requests of virtualized wireless sensor networks, a virtualized wireless sensor network resource allocation strategy based on reinforcement learning is proposed. First, a virtual sensor network request deployment utility maximization model based on business sensing information quality requirements and resource capacity constraints is established; then, considering the randomness of business requests and the dynamic changes of the network environment, its resource optimization problem is transformed into a Markov decision process, and finally, a meta-learning reinforcement learning method based on an evolutionary algorithm is adopted to continuously interact with the virtualized wireless sensor network environment to obtain the optimal resource allocation strategy; this method effectively improves the service provider's revenue and reduces the virtual sensor network request deployment cost while meeting resource requirements.
[0112] Embodiment 2:
[0113] The invention discloses a virtual wireless sensor network resource allocation system based on reinforcement learning. Figure 5 is a structural diagram of a virtual wireless sensor network resource allocation system based on reinforcement learning according to an embodiment of the present invention; Figure 5 As shown, the system 100 includes:
[0114] The first processing module 101 is configured to pre-acquire the resource requirements of a virtual wireless sensor network request arriving at any time;
[0115] The second processing module 102 is configured to select a candidate physical node set from the underlying wireless sensor network according to the required resource requirements and a pre-established deployment utility maximization model, and obtain a corresponding candidate wireless link set according to the candidate physical node set, wherein the deployment utility maximization model is pre-established based on the sensing information quality requirements and the resource requirement constraints;
[0116] The third processing module 103 is configured to substitute the candidate physical node set and the candidate wireless link set into the Markov decision model, and use the reinforcement learning algorithm to select a candidate physical node from the candidate physical node set as the best physical node, and select a corresponding candidate wireless link from the candidate wireless link set as the best wireless link according to the best physical node, so as to allocate corresponding communication bandwidth and sensor node computing resources for the virtual wireless sensor network request.
[0117] According to the system of the second aspect of the present invention, see Figure 6 As shown, the system also includes:
[0118] The fourth processing module 104 is configured to map the virtual node to the best physical node selected by the reinforcement learning algorithm.
[0119] According to the system of the second aspect of the present invention, the fourth processing module 104 is specifically configured to output a mapping success prompt if the virtual node is successfully mapped to the best physical node;
[0120] If the virtual node is not successfully mapped to the best physical node, the best physical node is removed from the candidate physical node set, and the mapping times of the virtual node are determined. If the mapping times are greater than the preset times threshold, a mapping failure prompt is output; if the mapping times are less than the preset times threshold, the third processing module 103 is reentered.
[0121] According to the system of the second aspect of the present invention, in the second processing module 102, the objective function of the deployment utility maximization model is expressed as:
[0122] max U(t)---(1);
[0123]
[0124] here, Indicates the location coordinates of the area of interest in the business request, c re (n s ) represents the node processing capacity threshold, b re (l s ) represents the threshold of system resources, b(l v ) represents the bandwidth resource demand of the virtual link. In addition, Represents the quality requirement of sensing information. is the Euclidean distance between the sensing area locations, Request distance threshold for the service.
[0125] Furthermore, c1 indicates that the processing resources occupied by the virtual sensor node on any physical sensor node should be less than the processing resources available to the physical sensor node; c2 indicates that the bandwidth resources occupied by the virtual link should be less than the bandwidth resources of its corresponding physical link; c3 indicates the quality requirement of the sensing information; c4 and c5 indicate the value range constraints of the variables; the objective function is U(t)=R(t)-C(t), where R(t) and C(t) represent the benefit and cost, respectively.
[0126] According to the system of the second aspect of the present invention, in the third processing module 103, the virtual wireless sensor network request is defined as a multi-tuple M=<S, A, P, R> through a Markov decision model;
[0127] Where S is the state space, A is the action space, P is the transition probability, R is the reward function, and s t∈S is the system state at time t, including the arrival and departure of virtual wireless sensor network requests, the resource requirements of virtual wireless sensor network requests, and the available resources of the current underlying wireless sensor network; action a t ∈A is based on the current network status, sensor node processing and link resource allocation; in state s t Execute action a t After that, the deployment of the virtual wireless sensor network request at the current moment is completed, and an immediate response R is obtained. t =U(t); state transition probability pr(s t+1 ∈S t+1 |s t ,a t ), indicating that in state s t Take action a t Transfer to t+1 probability.
[0128] According to the system of the second aspect of the present invention, in the third processing module 103, the reinforcement learning algorithm is a meta-learning reinforcement learning algorithm combined with an evolutionary algorithm, and the outer and inner loops are designed. The outer loop mainly performs algorithm performance evaluation, and selects the batch of algorithms with the largest reward returns from a pile of candidate algorithms; while the inner loop mainly performs algorithm screening, and after each algorithm is trained to the best state, the outer loop is used for performance evaluation. Finally, the system can automatically select the optimal virtual wireless sensor network resource allocation algorithm.
[0129] The specific functions of each module of the virtual wireless sensor network resource allocation system based on reinforcement learning disclosed in the present invention have been described in detail in a virtual wireless sensor network resource allocation method based on reinforcement learning, and will not be repeated here.
[0130] Embodiment 3:
[0131] The present invention discloses an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a virtual wireless sensor network resource allocation method based on reinforcement learning in any one of the embodiments 1 disclosed in the present invention are implemented.
[0132] Figure 7 is a structural diagram of an electronic device according to an embodiment of the present invention, such as Figure 7As shown, the electronic device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, near field communication (NFC) or other technologies. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covered on the display screen, or a button, a trackball or a touch pad set on the housing of the electronic device, or an external keyboard, touch pad or mouse, etc.
[0133] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a structural diagram of the part related to the technical solution of the present disclosure, and does not constitute a limitation on the electronic device to which the technical solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0134] Embodiment 4:
[0135] The present invention discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the virtual wireless sensor network resource allocation methods based on reinforcement learning in Embodiment 1 of the present invention are implemented.
[0136] Please note that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification. The above embodiments only express several implementation methods of the present application, and their descriptions are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, without departing from the concept of the present application, several variations and improvements can be made, which all belong to the scope of protection of the present application. Therefore, the scope of protection of the patent in this application shall be based on the attached claims.
[0137] The embodiments of the subject matter and functional operations described in this specification may be implemented in the following: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules in computer program instructions encoded on a tangible non-temporary program carrier to be executed by a data processing device or to control the operation of the data processing device. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information and transmit it to a suitable receiver device for execution by a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0138] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform corresponding functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuits, such as FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits), and the apparatus can also be implemented as special purpose logic circuits.
[0139] Computers suitable for executing computer programs include, for example, general and / or special microprocessors, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory and / or a random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, the computer will also include one or more large-capacity storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or the computer will be operably coupled to this large-capacity storage device to receive data from it or to transmit data to it, or both. However, the computer does not necessarily have such a device. In addition, the computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0140] Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0141] Although this specification includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of the specific embodiments of specific inventions. Certain features described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may work in certain combinations as described above and even initially claim protection, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of a sub-combination.
[0142] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or requiring that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
[0143] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.
[0144] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
[0145] Although some specific embodiments of the present invention have been described in detail by way of example, it will be appreciated by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It will be appreciated by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A virtual wireless sensor network resource allocation method based on reinforcement learning, characterized in that: The method comprises: Step S1, pre-acquire the resource requirements of the virtual wireless sensor network request arriving at any time; Step S2: selecting a candidate physical node set from the underlying wireless sensor network according to the required resource requirements and a pre-established deployment utility maximization model, and obtaining a corresponding candidate wireless link set according to the candidate physical node set, wherein the deployment utility maximization model is pre-established based on the sensing information quality requirements and resource requirement constraints; Step S3, substituting the candidate physical node set and the candidate wireless link set into a Markov decision model, and using a reinforcement learning algorithm to select a candidate physical node from the candidate physical node set as the best physical node, and selecting a corresponding candidate wireless link from the candidate wireless link set as the best wireless link according to the best physical node, so as to allocate corresponding communication bandwidth and sensor node computing resources to the virtual wireless sensor network request; After step S3, the method further includes: Step S4: Map the virtual node to the best physical node selected by the reinforcement learning algorithm; In step S4, if the virtual node is successfully mapped to the best physical node, a successful mapping prompt is output; if the virtual node is not successfully mapped to the best physical node, the best physical node is removed from the candidate physical node set, and the mapping times of the virtual node are determined. If the mapping times are greater than a preset times threshold, a mapping failure prompt is output; if the mapping times are less than the preset times threshold, the process proceeds to step S3 again; In step S2, the objective function of the deployment utility maximization model is expressed as: max U(t)---(1); here, Indicates the location coordinates of the area of interest in the business request, c re (n s ) represents the node processing capacity threshold, b re (l s ) represents the threshold of system resources, b(l v ) represents the demand for bandwidth resources of the virtual link; in addition, represents the sensing information quality requirement; d is the Euclidean distance between the sensing area locations, Request distance threshold for service; Furthermore, c1 indicates that the processing resources occupied by the virtual sensor node on any physical sensor node should be less than the processing resources available to the physical sensor node; c2 indicates that the bandwidth resources occupied by the virtual link should be less than the bandwidth resources of its corresponding physical link; c3 indicates the quality requirement of the sensing information; c4 and c5 indicate the value range constraints of the variables; the objective function is U(t)=R(t)-C(t), where R(t) and C(t) represent the benefit and cost respectively; In step S3, the virtual wireless sensor network request is defined as a multi-tuple M=<S,A,P,R> ; Where S is the state space, A is the action space, P is the transition probability, R is the reward function, and s t ∈S is the system state at time t, including the arrival and departure of virtual wireless sensor network requests, the resource requirements of virtual wireless sensor network requests, and the available resources of the current underlying wireless sensor network; action a t ∈A is based on the current network status, sensor node processing and link resource allocation; in state s t Execute action a t After that, the deployment of the virtual wireless sensor network request at the current moment is completed, and an immediate response R is obtained. t =U(t); state transition probability pr(s t+1 ∈S t+1 |s t ,a t ), indicating that in state s t Take action a t Transfer to t+1 The probability of In step S3, the reinforcement learning algorithm is a meta-learning reinforcement learning algorithm combined with an evolutionary algorithm. By designing an inner and outer loop, the outer loop mainly performs algorithm performance evaluation, and selects the batch of algorithms with the largest reward returns from a pile of candidate algorithms; while the inner loop mainly performs algorithm screening. After each algorithm is trained to the best state, the outer loop is used for performance evaluation to automatically select the optimal virtual wireless sensor network resource allocation algorithm.
2. A virtual wireless sensor network resource allocation system based on reinforcement learning, the system adopts the method of claim 1, characterized in that: The system comprises: The first processing module is configured to pre-acquire the resource requirements of the virtual wireless sensor network request arriving at any time; A second processing module is configured to select a set of candidate physical nodes from the physical network according to the required resource requirements and a pre-established deployment utility maximization model, and obtain a corresponding set of candidate wireless links according to the set of candidate physical nodes, wherein the deployment utility maximization model is pre-established based on the sensing information quality requirements and the resource requirement constraints; A third processing module is configured to substitute the candidate physical node set and the candidate wireless link set into a Markov decision model, and select a candidate physical node from the candidate physical node set as an optimal physical node using a reinforcement learning algorithm, and select a corresponding candidate wireless link from the candidate wireless link set as an optimal wireless link according to the optimal physical node, so as to allocate corresponding communication bandwidth and sensor node computing resources to the virtual wireless sensor network request; The fourth processing module is configured to map the virtual node to the best physical node selected by the reinforcement learning algorithm.
3. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps in the virtual wireless sensor network resource allocation method based on reinforcement learning described in claim 1 are implemented.
4. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps in the virtual wireless sensor network resource allocation method based on reinforcement learning described in claim 1 are implemented.
Citation Information
Patent Citations
A low-cost industrial wireless sensor selection method
CN109089323A
Mapping method and device for virtualized wireless sensor network, and storage medium
CN110933728A