A space flight information system resource dynamic allocation method based on reinforcement learning
By employing a dynamic resource allocation method based on reinforcement learning, the problems of uneven business distribution and different resource demands among service areas in the space flight information system were solved, realizing intelligent resource allocation and improving the system's resource utilization and service performance.
Patent Information
- Application Number
- CN202310214571.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-02-28
AI Technical Summary
In existing space flight information systems, uneven distribution of services and different resource requirements among service areas lead to severe resource interference. Existing static resource allocation methods are inflexible and unadaptable, making it difficult to effectively solve the resource allocation problem of complex and ever-changing digital twin systems.
A dynamic resource allocation optimization model based on reinforcement learning is constructed, with the objective function of minimizing the user terminal blocking rate and the constraint of resource allocation rationalization. The dynamic resource allocation optimization model is solved by reinforcement learning to achieve intelligent resource allocation.
It improved resource utilization, effectively avoided resource usage interference, enhanced system service performance, and reduced the business blocking rate of the digital twin system.
Smart Images

Figure CN116347603B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of resource allocation, and particularly relates to a space flight information system resource dynamic allocation method based on reinforcement learning. BACKGROUND
[0002] In recent years, with the development and maturity of the Internet and mobile Internet, online and virtual application scenarios in the social and entertainment fields are becoming more and more abundant. After 2020, people's life is increasingly dependent on online communication, which has promoted the rapid development of related technologies such as digital twin, artificial intelligence, big data, and virtual reality, and has given birth to the development and evolution of the "metaverse" concept facing future comprehensive virtual and real integration. On the other hand, the state promotes the digital transformation of the industry and the digital upgrading of important industries in the national economy, and on the basis of the digital twin technology empowering the industry, in the future, the industrial and industrial metaverse will be created for the industrial production and manufacturing field with more complex and demanding resource use requirements, so that in terms of technology iteration and comprehensive management, it will develop towards more efficient and more abundant aspects.
[0003] Therefore, the virtual world composed of digital resources will rapidly expand in the next few decades, and digital twins of various industries and applications will continue to emerge, and the digital resources based on computer, network and communication technology are limited, which puts forward higher requirements for the comprehensive and efficient use of digital resources.
[0004] At the same time, in the case of sharing resources by digital twins of various industries and applications, resource reuse without mutual interference is a solution to efficient resource utilization. Moreover, the resource demand of various applications changes dynamically, and intelligent resource allocation strategy is a feasible way to improve resource utilization.
[0005] At present, the digital twin obtained by mapping the space flight information system mostly uses static (frequency or power, etc.) resource reuse method, the adjacent service areas use different resources to provide services to the users in their coverage areas, and the coverage areas using the same resources are as far away from each other as possible, so as to suppress the interference between the reused same resources; but as the number of system users increases, the overlap of the coverage areas of multiple service providing nodes will be higher and higher, resulting in more and more serious interference caused by the reuse of the same resources, although in extreme cases some service areas can be closed, but when the overlapping area gradually increases, the interference between resource use is also an important problem to be solved. In addition, when the coverage areas of digital twins of different types of systems overlap, they may interfere with the use of resources of the existing system.
[0006] Current methods for dealing with resource interference mostly focus on dealing with frequency resource reuse, i.e. co-channel interference. The following gives the current existing schemes for dealing with co-channel interference:
[0007] One is to enhance inter-cell interference coordination. For the same frequency interference, eICIC (enhanced inter-cell interference coordination) technology can be considered to suppress interference. Specifically, in the downlink direction, ABSF (Almost Blank Subframe) technology is used to prohibit multiple interfering resource use areas from transmitting downlink service data at the same time, which is a kind of time division multiplexing idea; in the uplink direction, different users of different service areas are allocated non-overlapping frequency domain resources for uplink service data, which is a kind of frequency division multiplexing idea. This method has two problems. One is the complexity of scheduling. A complex network digital twin needs to model and simulate a multi-level time-varying network. The same frequency interference changes over time, and the number of same frequency cells that interfere with each other changes. The application of eICIC technology under this condition needs to be studied. In addition, due to the staggered transmission in time, time synchronization needs to be maintained between different coverage areas.
[0008] Another one is based on related research, which shows that multi-carrier spread spectrum technology can resist narrowband interference like traditional spread spectrum technology, and is more effective in resisting partial wideband interference. When multi-carrier spread spectrum technology is used for multiple access transmission, the system is not as sensitive to changes in the number of users as traditional CDMA technology, which is particularly advantageous in environments with fluctuating traffic demand. OFDM technology itself has the ability to resist multipath fading, and when combined with spread spectrum technology, the system performance is better than traditional CDMA technology using RAKE reception, and the implementation is simpler. There are several ways to implement multi-carrier spread spectrum technology, and the most suitable way is MC-CDMA technology. However, the problem with this method is that in the specific design, when the number of overlapping resource use areas increases, the spread spectrum gain obtained by this method may not be able to maintain the correct transmission performance. Therefore, the spread spectrum gain needs to be increased, and a frame structure with lower bandwidth needs to be designed while avoiding the introduction of excessive overhead, which further increases the complexity of system design. In addition, the selection of spread spectrum codes is also an important problem. If completely orthogonal spread spectrum codes such as Walsh codes are used, time synchronization needs to be maintained between different service areas to ensure orthogonality. If non-orthogonal spread spectrum codes are used, the impact of interference introduced by different service areas on system performance needs to be carefully studied.
[0009] From this, it can be seen that the above static resource allocation method has poor flexibility, complex design, and poor adaptability to complex and variable digital twin system resource allocation. SUMMARY
[0010] In view of the above analysis, the embodiments of the present application aim to provide a spatial flight information system resource dynamic allocation method based on reinforcement learning, to solve the problem of uneven distribution of services between service areas and different resource demands in existing spatial flight information systems.
[0011] The application discloses a space flight information system resource dynamic allocation method based on reinforcement learning, which comprises the following steps:
[0012] The space flight information system is mapped into a digital twin system, and all available resources, service areas and user terminals in the digital twin system are acquired;
[0013] Based on the digital twin system, a dynamic resource allocation optimization model is constructed, with minimization of user terminal blocking rate as an objective function and resource allocation rationalization as a constraint condition;
[0014] When a service request of a user terminal is received, the dynamic resource allocation optimization model is solved based on a reinforcement learning mode, so that a dynamic resource allocation strategy of the space flight information system is obtained.
[0015] On the basis of the above scheme, the application further makes the following improvements:
[0016] Further, the objective function is:
[0017] max r=R max *(1-U block / U all ) (1)
[0018] Wherein, r represents the objective function, R max represents an optimized blocking rate reward coefficient; U block represents the total number of user terminals in a blocking state in the digital twin system, and U all represents the total number of user terminals sending service requests in the digital twin system.
[0019] Further, the constraint condition is:
[0020]
[0021] Wherein, p n =[p n,1 ,p n,2 ,...,p n,m ,...,p n,M ] T , p n,m represents the power of the available resource m allocated to the service area n; the maximum power of the digital twin system is recorded as P tot , and the maximum power of each service area is recorded as P b ; H represents a Hamiltonian transpose, the service area set B={n|n=1,2,...,N}, N represents the total number of service areas; the available resource set C={m|m=1,2,...,M}, M represents the total number of available resources;
[0022] Cu,m denotes the capacity of the available resource m allocated to the user terminal u, C th denotes the capacity threshold; each user terminal is uniquely identified by a binary tuple set U = {u | u = (n, k), n ∈ B, k ∈ Z}, where the user terminal u represents the ID of the user terminal accessing the nth service area, and the user terminal ID set Z = {k | k = 1, 2,..., K};
[0023] d i,j denotes the distance between the service area i and the service area j, L d denotes the minimum resource reuse distance; w i,m , w j,m respectively represent the occupation state of the available resource m by the service area i and j.
[0024] Further,
[0025] C u,m = C subc · log2 (1 + SINR u,m ) (3)
[0026] wherein C subc denotes the capacity of each available resource; SINR u,m denotes the useful-to-noise signal ratio of the user terminal u receiving the available resource m:
[0027]
[0028] wherein p b,m denotes the power of the available resource m allocated to the service area b, denotes the channel noise of the user terminal u;
[0029] The channel transmission loss matrix E = {e u,n | u ∈ U, u = (n, k), n ∈ B} between the resource provider and the user terminal; wherein e u,n denotes the channel transmission loss of the user terminal u transmitting the available resource accessing the service area n, e u,b denotes the channel transmission loss of the user terminal u transmitting the available resource accessing the service area b.
[0030] Further,
[0031] E = O · G U · G B (5)
[0032] wherein O denotes the path loss matrix caused by free space; G B denotes the power gain matrix of the resource provider; G U denotes the power gain of the user terminal.
[0033] Further, O = diag{o1, o2,..., o u ,...,o U}, o u represents the path loss of the user terminal u due to free space;
[0034] G B = {g u,n | u e U, u = (n, k), n e B}, g u,n represents the power gain provided by the resource provider to the user terminal u in the access service area n;
[0035] G U = diag{g1, g2,..., g u ,...,g U}, g u represents the power gain of the user terminal u in the access service area n.
[0036] Further, when receiving the service request of the user terminal, the dynamic resource allocation optimization model is solved based on the reinforcement learning mode, including:
[0037] Step S31: parameter initialization; initialize the learning rate a, the discount factor g, the optimization period T,
[0038] Initialize the exploration probability e = e init , and let t = 0; initialize s0 = {W ad (0), P ad (0)}; W ad (t) represents the resource occupation state of the available resources at P ad (t) time, and P ad (t) represents the power allocation information respectively; if W ad (0) contains 0 elements, execute step S32;
[0039] Step S32: update the exploration probability e = max (e - e gap , e f );
[0040] Let state s t = {W ad (t), P ad (t)}, calculate the feasible action set A (s t ) according to the state s t ;
[0041] With e probability, randomly select action a t e A (s t ),
[0042] Otherwise, select at = argmax a Q(s t ,a t );
[0043] performing action a t , updating the environment to state s t = {W ad (t+1), P ad (t+1)} at time t+1, and obtaining reward r(t) = R max *[1-U block (t) / U all (t)] at time t;
[0044] updating the Q value, Q(s t+1 ,a t+1 )←Q(s t ,a t )+α[r t +γmaxQ(s t ,a t )-Q(s t ,a t )].
[0045] determining whether W ad (t+1) contains no 0 elements,
[0046] if yes, jumping to step S33;
[0047] if no, determining whether t=T is true,
[0048] if yes, jumping to step S33;
[0049] if no, t=t+1, and jumping to step S32;
[0050] step S33: taking the action corresponding to argmax a Q(s,a) as the optimal resource allocation strategy.
[0051] Further, the service request of the user terminal includes: a resource occupation request of the user terminal, and a power allocation request of a service area in which the user terminal is located.
[0052] W re represents a resource occupation state matrix corresponding to the resource occupation request of the user terminal; and P re represents a power allocation matrix corresponding to the power allocation request of the service area in which the user terminal is located.
[0053] W ad (t) is expressed in the form of the resource occupation state matrix W; and P ad (t) is expressed in the form of the power allocation matrix P.
[0054] Resource occupancy state matrix W = [w1, w2, ..., w n ,...,w N ], where w n =[w n,1 ,w n,2 ,...,w n,m ,...,w n,M ] T ;w n,m This indicates the occupancy status of service area n for available resource m.
[0055] Power allocation matrix P = [p1, p2, ..., p n ,...,p N ].
[0056] Furthermore, w n,m =1 indicates that service area n occupies available resources m, w n,m =0 indicates that the service area n does not occupy the available resource m.
[0057] Furthermore, the statement based on state s t Calculate the set of feasible actions A(s) t ),implement:
[0058] According to W at time t ad (t) and P ad (t), determine the resource occupancy state matrix W corresponding to unoccupied resources. no (t) and the power allocation matrix P corresponding to the unallocated power no (t);
[0059] From W no Select from (t) those that simultaneously satisfy W re Given all feasible resource allocation methods and constraint d4, where feasible resource allocation methods are represented in the form of a resource occupancy state matrix;
[0060] From P no Select from (t) those that simultaneously satisfy P re Given all feasible power allocation methods for constraints d1-d3, where feasible power allocation methods are represented in the form of a power allocation matrix;
[0061] Each feasible resource allocation method and power allocation method is combined into a feasible action. All feasible actions are then aggregated to form a set of feasible actions A(s). t ).
[0062] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0063] The application provides a space flight information system resource dynamic allocation method based on reinforcement learning.
[0064] In the implementation process, the decision-making ability of the digital twin for resource allocation is improved through dynamic intelligent learning and decision-making. The application proposes a system framework for interaction between an agent and an environment, feedback of state, action and decision information, and allocation process and steps based on the target and constraints of dynamic resource allocation. The application analyzes the dynamic resource allocation problem in the system business scenario, establishes an optimization problem of maximizing resource utilization, and gives the effectiveness of the intelligent model in solving the resource allocation problem, and gives the detailed design of the dynamic resource allocation algorithm. Simulation shows that the resource allocation method proposed in the application can reduce the digital twin system business blocking rate and improve the system resource utilization.
[0065] The above technical solutions can be combined with each other to achieve more preferred combination solutions. Other features and advantages of the application will be described in the subsequent specification, and some advantages will become apparent from the specification, or will be understood by implementing the application. The purpose and other advantages of the application can be achieved and obtained from the specific embodiments described in the specification and the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0066] The accompanying drawings are included to provide a further understanding of the application and are incorporated herein and constitute a part of the application. The application will be described with reference to the accompanying drawings, and the same reference numerals in the drawings denote the same elements.
[0067] Figure 1 A flowchart of a digital twin resource allocation method based on reinforcement learning provided by an embodiment of the application;
[0068] Figure 2 A comparison diagram of the resource dynamic allocation method of the application and the resource fixed allocation method in the prior art provided by an embodiment of the application. DETAILED DESCRIPTION
[0069] The preferred embodiments of the present application will be described in detail below with reference to the drawings, which form a part of this application. The drawings illustrate embodiments of the application and, together with the description, serve to explain the principles of the application, and are not intended to limit the scope of the present application.
[0070] One specific embodiment of the present application discloses a method for dynamic allocation of resources of a space flight information system based on reinforcement learning, a flowchart of which is shown in Figure 1 The method comprises the following steps:
[0071] Step S1: mapping the space flight information system into a digital twin system, and obtaining all available resources, service areas and user terminals in the digital twin system;
[0072] The space flight information system is an information system formed by multiple aircrafts distributed in space as resource providers, and ground user terminals distributed in multiple service areas as resource users. Specifically, the resource providers are used to provide available resources, which mainly include power resources in this embodiment, i.e., the allocation of transmission signal power resources in the service areas. The service area is one or more areas that use different available resources to realize a certain type of business function for the user terminal. The user terminal is a business requester distributed in different service areas, which accesses the resource provider through the available resources allocated in the respective service area, and uses the corresponding available resources provided by the resource provider.
[0073] In the process of mapping the space flight information system into the digital twin system, the available resources in the space flight information system are mapped into the available resources in the digital twin system; the service areas in the space flight information system are mapped into the service areas in the digital twin system; and the user terminals in the space flight information system are mapped into the user terminals in the digital twin system, thereby forming the digital twin system.
[0074] Step S2: based on the digital twin system, constructing a dynamic resource allocation optimization model with the minimization of the user terminal blocking rate as the objective function and the rationalization of resource allocation as the constraint condition;
[0075] Specifically, in this embodiment, the dynamic resource allocation optimization model of the digital twin system is constructed in the following manner:
[0076] The digital twin system includes N service areas, which are represented by a service area set B = {n | n = 1, 2,..., N}. The available resource set C = {m | m = 1, 2,..., M} in the digital twin system, M represents the total number of available resources. Among them, the available resources do not overlap with each other and do not interfere with each other, the capacity of each available resource is equal, and the capacity of each available resource can be represented as C subc=C totc M, C totc denotes the total capacity of all available resources in the digital twin system.
[0077] The user terminals are distributed in different service area ranges, and according to the resource use requirements of the user terminals, resource access is carried out in the respective service areas, and each user terminal has a unique corresponding relationship with the accessed resources. It is assumed that the digital twin system includes K user terminals, and is represented by the user terminal ID set Z = {k | k = 1, 2,..., K}, and k represents the ID of the user terminal. Each user terminal can be uniquely identified by the binary tuple set U = {u | u = (n, k), n ∈ B, k ∈ Z}, and the user terminal u represents that the ID of the user terminal accessing the nth service area is k.
[0078] In the model construction process, attention is focused on whether the resources are occupied, and the power allocation of each service area. Based on this, the following related parameters are designed:
[0079] The resource occupation state vector of the service area n is denoted as w n = [w n,1 ,w n,2 ,…,w n,m ,…,w n,M ] T , where w n,m represents the occupation state of the available resource m of the service area n, w n,m = 1 indicates that the service area n occupies the available resource m, and w n,m = 0 indicates that the service area n does not occupy the available resource m. At this time, the resource occupation vectors of all service areas in the digital twin system constitute the resource occupation state matrix W = [w1, w2,..., w n ,…,w N ] of the digital twin system. And define the power allocation matrix P = [p1, p2,..., p n ,...,p N ] with the service area as the basic unit, where p n = [p n,1 ,p n,2 ,...,p n,m ,...,p n,M ] T , p n,m represents the power of the available resource m allocated to the service area n. The maximum power of the digital twin system is denoted as P tot , the maximum power of each service area is denoted as P b , and P b = P tot / N.
[0080] In the embodiment, the channel transmission loss matrix E = {e u,n |u∈U,u=(n,k),n∈B}, where e u,n represents the channel transmission loss of the user terminal u transmitting the available resource in the service area n. In practical application, E can be calculated by the following formula:
[0081] E = O · G U · G B (1)
[0082] where O represents the path loss matrix caused by free space (such as atmospheric attenuation, etc.); specifically, O = diag{o1, o2,..., o u ,...,o U}, o u represents the path loss of the user terminal u caused by free space;
[0083] G B represents the power gain matrix of the resource provider; G B = {g u,n |u∈U,u=(n,k),n∈B}, g u,n represents the power gain of the resource provider provided to the user terminal u in the service area n;
[0084] G U represents the power gain of the user terminal; G U = diag{g1, g2,..., g u ,...,g U}, g u represents the power gain of the user terminal u in the service area n.
[0085] At this time, the signal to interference and noise ratio (SINR) of the user terminal u receiving the available resource m can be represented by the following formula: u,m
[0086]
[0087] where p b,m represents the power of the available resource m allocated to the service area b, that is, the interference of other service areas to the service area n on the same resource, e u,b represents the channel transmission loss of the user terminal u transmitting the available resource in the service area b, represents the channel noise of the user terminal u (i.e. noise not caused by free space). Further, the capacity C u,m of the available resource m allocated to the user terminal u can be calculated by the Shannon formula, as shown in formula (3):
[0088] C u,m =C subc ·log2(1+SINR u,m ) (3)
[0089] To ensure that the service quality requirements of the user terminal can be met, the capacity under the allocated resources should be at least guaranteed not to be lower than the capacity threshold C th , the value of C th is related to the type of transmission service and the performance of the receiving end, and can be set according to the actual application scenario. When C u,m ≥ C th (the threshold size needs to be designed according to the actual service requirements), the service quality requirements of the user terminal can be met.
[0090] In this embodiment, the dynamic resource allocation optimization model with the minimum user terminal blocking rate as the objective function and the resource allocation rationalization as the constraint condition is shown in formulas (4) and (5):
[0091] max r=R max *(1-U block / U all ) (4)
[0092]
[0093] Formula (4) is the objective function with the minimum user terminal blocking rate as the target; wherein, r represents the objective function; R max represents the optimization blocking rate reward coefficient, R max is a scalar positive value; U block represents the total number of user terminals in the blocking state in the digital twin system, U all represents the total number of user terminals that send service requests in the digital twin system, U block / U all represents the user terminal blocking rate.
[0094] Formula (5) is the constraint condition; wherein, the constraint condition d1 indicates that the resource allocation scheme should meet that the allocated power should not exceed the total power of the resource provider, H represents the Hamilton transpose. The constraint condition d2 indicates that the power of each service area should not exceed the maximum transmission power of each service area. The constraint condition d3 indicates that the current resource allocation scheme will not affect the service quality of the existing services, and also avoids the influence of the interference problem. The constraint condition d4 is the interference limiting condition, w i,m , w j,m respectively represent the occupation state of the available resource m of the service area i and j; d i,j represents the distance between the service area i and the service area j; that is, the minimum resource reuse distance L dInside, only one service area is allowed to use one available resource.
[0095] After determining the required resource demand of the service request of the user terminal, the candidate resource allocation mode set and the candidate power allocation mode set can be selected by solving the constraint conditions in the dynamic resource allocation optimization model.
[0096] Step S3: When receiving the service request of the user terminal, the dynamic resource allocation strategy of the space flight information system is obtained by solving the dynamic resource allocation optimization model based on the reinforcement learning mode.
[0097] In the embodiment, the reinforcement learning model is established by taking the resource provider as an agent, taking the resource usage as an environment, and taking the allocation of available resources as an action. The agent observes the environment s t Decide to perform action a t , a t influences the environment to become s t+1 , and the agent gets the instant feedback of the environment r t . Specifically, the main contents of the reinforcement learning model are as follows:
[0098] (1) State
[0099] The state S is an abstraction of the environment and is also the basis for determining the action performed. In the embodiment, the state s t at time t is s ad (t) = {W ad (t), P t (t)}, s ad ∈ S; wherein W ad (t) represents the resource occupied state of the available resource at time t, which is represented in the form of resource occupation state matrix W; P ad (t) represents the power allocation information at time t, which is represented in the form of power allocation matrix P.
[0100] When the state W t (t) at a certain time does not contain 0 elements, it means that all resources are occupied, reaching the termination state, i.e. there is no available resource for the current user terminal.
[0101] (2) Action
[0102] The action is the output of the agent to the environment, which, in this embodiment, refers to allocating the available resources and their power to the user terminal initiating the service request. In the specific implementation process, according to the service request of the user terminal at time t and the state of the agent, the feasible action set A(s tChoose the action with the largest Q value to execute a(t).
[0103] a t ={(n,m)|n,m∈A(s)} t ),n∈B,m∈M}
[0104] The action a t Indicates: In state s t The set of feasible actions A(s) t In ), the available resources m will be converted to power p. n,m The available resources of the agent are provided to the service area n. By selecting different actions based on the policy learned in Q-learning under different states, the agent allocates the available resources to the user terminals in each service area.
[0105] User terminal service requests include: resource usage requests from the user terminal, and power allocation requests from the service area where the user terminal is located; among which, W re This represents the resource usage state matrix corresponding to the resource usage requests of user terminals; P re This represents the power allocation matrix corresponding to the power allocation request of the service area where the user terminal is located.
[0106] Specifically, in this embodiment, the set of feasible actions A(s) at time t is determined in the following manner. t ):
[0107] According to W at time t ad (t) and P ad (t), determine the resource occupancy state matrix W corresponding to unoccupied resources. no (t) and the power allocation matrix P corresponding to the unallocated power no (t);
[0108] From W no Select from (t) those that simultaneously satisfy W re Given all feasible resource allocation methods and constraint d4, where feasible resource allocation methods are represented in the form of a resource occupancy state matrix;
[0109] From P no Select from (t) those that simultaneously satisfy P re Given all feasible power allocation methods for constraints d1-d3, where feasible power allocation methods are represented in the form of a power allocation matrix;
[0110] Each feasible resource allocation method and power allocation method is combined into a feasible action. All feasible actions are then aggregated to form a set of feasible actions A(s). t ).
[0111] (3) Rewards
[0112] The reward corresponds to the objective function of the dynamic resource allocation optimization model.
[0113] The reward is the feedback from the environment in the process of interaction between the agent and the environment, and is the evaluation after determining the corresponding action in the state. Whether the value is designed reasonably directly determines the size of the long-term income of the agent, that is, the performance of the solution to the dynamic resource allocation problem. In the dynamic resource allocation problem, the optimization goal is to maximize the utility of the system. Taking the blocking rate as an example, the optimization goal is to minimize the number of blocked users of the system.
[0114] r=R max *(1-U block / U all )
[0115] It can be seen that the fewer the number of blocked users in the digital twin system, the more the reward obtained, and the higher the overall utility performance of the digital twin system. Because the agent pays more attention to the reward when reaching the final state, the immediate reward in the state transition process can be set to 0.
[0116] Based on the foregoing definitions of the environment, state, action, and reward, the dynamic resource allocation algorithm based on reinforcement learning is as follows:
[0117] Whenever a service request of a user terminal is received, the following is performed:
[0118] Step S31: parameter initialization; initialize the learning rate a, the discount factor g, the optimization period T,
[0119] Initialize the exploration probability e = e init , and let t = 0; initialize s0 = {W ad (0), P ad (0)}; if W ad (0) contains 0 elements, perform step S32;
[0120] Step S32: update the exploration probability e = max (e - e gap , e f );
[0121] Let the state s t = {W ad (t), P ad (t)}, calculate the feasible action set A(s t ) according to the state s t ;
[0122] Randomly select an action a t e A(s t ) with a probability of e,
[0123] otherwise, select at = argmax a Q(s t ,a t );
[0124] perform action a t , update the environment to state s t = {W ad (t+1), P ad (t+1)} at time t+1, and obtain the reward r t at time t; r(t) = R max *[1-U block (t) / U all (t)].
[0125] update the Q value, Q(s t+1 ,a t+1 ) <- Q(s t ,a t )+a[r t +g max Q(s t ,a t )-Q(s t ,a t )].
[0126] determine whether W ad (t+1) contains no 0 elements,
[0127] if yes, jump to step S33;
[0128] if no, determine whether t = T is true,
[0129] if yes, jump to step S33;
[0130] if no, t = t+1, and jump to step S32;
[0131] step S33: take the action corresponding to argmax a Q(s,a) as the optimal resource allocation strategy.
[0132] The action selection strategy in the above algorithm flow adopts an e-greedy strategy, i.e., randomly selecting an action with a probability e e [0,1], or otherwise selecting and executing the action with the maximum Q value. For example, the exploration probability e is 30%, and through a computer program random number generator mechanism, it is determined whether the current loop round is randomly selected (exploration) or the action with the maximum Q value is selected. In theory, after infinite times of simulation, the number of loop rounds in which the action is selected (exploration) accounts for 30% of the total number of loop rounds. "Select an action a t e A(s t ) with a probability e, or otherwise, select a t = argmaxa Q(s t ,a t The relationship between the two approaches is as follows: Randomly selecting an action ("exploration") means randomly choosing one action from the action space to execute, and its Q value is not necessarily the largest. On the other hand, the choice following "otherwise" has the largest Q value, but as will be discussed later, this may "get stuck in a local optimum".
[0133] Therefore, in this embodiment, an ε-greedy strategy is chosen to strike a balance between exploration and exploitation. Exploration refers to making decisions based on currently available information, thus fully utilizing historical experience; exploration refers to discarding currently available information and randomly trying a new method, thus avoiding getting trapped in local optima and searching for feasible globally optimal solutions. During training, the exploration probability ε should gradually decrease. This scheme adopts a linear descent criterion, with the decay factor denoted as ε. gap , from the initial exploration probability ε init Decay to the final exploration probability ε f .
[0134] Example 2
[0135] To further illustrate the beneficial effects of the present invention, the resource allocation system and method proposed in this invention are further simulated and calculated below.
[0136] Different business distributions were selected as simulation scenarios, and compared with a fixed resource allocation method. The fixed resource allocation algorithm divides the available resources into several subsets, and each service area selects one set from the subsets as the available resource allocation set.
[0137] Table 1 Simulation parameters of resource allocation algorithm
[0138]
[0139] The user terminal service arrival model used in the simulation follows a Poisson distribution with parameter λ. The service duration follows a negative exponential distribution with parameter μ.
[0140] Figure 2 The system blocking rate performance of two allocation methods under different service arrival rates λ is presented. The service duration is constant at μ = 3 minutes. Figure 2 As shown, under the same service arrival rate, the resource allocation algorithm based on reinforcement learning proposed in this invention can achieve a lower blocking rate compared with the fixed allocation algorithm.
[0141] The results show that the blocking rate increases with the increase of the service arrival rate, which is mainly because with the increase of the service, due to the fixed number of available channels, more services will be blocked due to the inability to obtain service. Under the same service arrival rate, the method of the present application can achieve lower blocking rate compared with the fixed resource allocation method. For example, when the service arrival rate λ = 80, the blocking rates of the fixed and dynamic methods are 0.31 and 0.09 respectively. At the same time, when the system blocking rate performance is 0.10, the load bearing capacity of the fixed and dynamic methods is λ = 43 and λ = 82 respectively, that is, the algorithm proposed in the present application can improve the load bearing capacity by one time compared with the fixed allocation method.
[0142] Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a computer readable storage medium. The computer readable storage medium includes a magnetic disk, an optical disk, a read-only memory, a random access memory, etc.
[0143] The above description is only the preferred embodiment of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application.
Claims
1. A method for dynamic resource allocation in a space flight information system based on reinforcement learning, characterized in that, include: The space flight information system is mapped into a digital twin system, and all available resources, service areas and user terminals in the digital twin system are obtained; Based on the aforementioned digital twin system, a dynamic resource allocation optimization model is constructed with the objective function of minimizing the user terminal blocking rate and the constraint of resource allocation rationalization. When a service request is received from a user terminal, a dynamic resource allocation optimization model is solved based on reinforcement learning to obtain the dynamic resource allocation strategy of the space flight information system. The objective function is: maxr=R max *(1-U block U all ) (1) Where r represents the objective function, R max U represents the reward coefficient for optimizing the blocking rate; block U represents the total number of user terminals in a blocked state in the digital twin system. all This represents the total number of user terminals that issue service requests in the digital twin system; The constraints are as follows: Where, p n =[p n,1 ,p n,2 ,...,p n,m ,...,p n,M ] T p n,m This represents the power allocated from available resource m to service area n; the maximum power of the digital twin system is denoted as P. tot The maximum power in each service area is denoted as P. b H represents Hamiltonian transpose; the service area set B = {n | n = 1, 2, ..., N}, where N represents the total number of service areas; the available resource set C = {m | m = 1, 2, ..., M}, where M represents the total number of available resources; C u,m C represents the capacity of available resource m allocated to user terminal u. th The capacity threshold is represented by a set of binary tuples U = {u | u = (n, k), n ∈ B, k ∈ Z}, where each user terminal is uniquely identified by a set of binary tuples U = {u | u = (n, k), n ∈ B, k ∈ Z}. User terminal u represents the user terminal with ID k that accesses the nth service area. The set of user terminal IDs Z = {k | k = 1, 2, ..., K}. d i,j L represents the distance between service region i and service region j. d Indicates the minimum resource reuse distance; w i,m w j,m These represent the occupancy status of service regions i and j for available resource m, respectively.
2. The method for dynamic resource allocation in a space flight information system based on reinforcement learning according to claim 1, characterized in that, C u,m =C subc ·log2(1+SINR u,m ) (3) Among them, C subc SINR represents the capacity of each available resource. u,m This represents the ratio of useful to useless signals when user terminal u receives available resources m: Where, p b,m This represents the power that available resource m is allocated to service area b. The channel noise of user terminal u; The channel transmission loss matrix E = {e} between the resource provider and the user terminal u,n |u∈U,u=(n,k),n∈B}; where, e u,n e represents the channel transmission loss of available resources for user terminal u to access service area n. u,b This represents the channel transmission loss of available resources for user terminal u to access service area b.
3. The method for dynamic resource allocation in a space flight information system based on reinforcement learning according to claim 2, characterized in that, E=O·G U ·G B (5) Where O represents the path loss matrix caused by free space; G B G represents the power gain matrix of the resource provider; U This indicates the power gain of the user terminal.
4. The method for dynamic resource allocation in a space flight information system based on reinforcement learning according to claim 3, characterized in that, O = diag{o1, o2, ..., o u ,...,o U }, o u This represents the path loss of user terminal u due to free space limitations; G B ={g u,n |u∈U,u=(n,k),n∈B},g u,n This represents the power gain provided by the resource provider to the user terminal u in the access service area n; G U =diag{g1,g2,...,g u ,...,g U }, g u This represents the power gain of user terminal u in the access service area n.
5. The method for dynamic resource allocation in a space flight information system based on reinforcement learning according to any one of claims 1-4, characterized in that, When a service request is received from a user terminal, a dynamic resource allocation optimization model is solved based on reinforcement learning, including: Step S31: Parameter initialization; initialize learning rate α, discount factor γ, and optimization period T. Initialize the exploration probability ε = ε init And, let t = 0; initialize s0 = {W ad (0),P ad (0)};W ad (t) represents P ad The resource occupancy status of available resources at time (t), P ad (t) represent the power allocation information; if W ad If (0) contains the element 0, proceed to step S32; Step S32: Update the exploration probability ε = max(ε - ε gap ,ε f );ε gap ε is the attenuation factor. f For the final exploration probability; Let state s t ={W ad (t),P ad (t)}, according to state s t Calculate the set of feasible actions A(s) t ); Action a is randomly selected with probability ε. t ∈A(s t ), Otherwise, choose a. t =argmax a Q(s t ,a t ); Perform action a t Update the environment to the state s at time t+1. t ={W ad (t+1),P ad (t+1)}, and obtain the reward r(t) = R at time t. max *[1-U block (t) / U all (t)]; Update the Q value, Q(s) t+1 ,a t+1 )←Q(s t ,a t )+α[r t +γmaxQ(s t ,a t )-Q(s t ,a t )]; Determine W ad Does (t+1) contain no 0 elements? If so, proceed to step S33; If not, determine whether t = T is true. If true, proceed to step S33; If not true, t = t + 1, then jump to step S32; Step S33: Set argmax a The action corresponding to Q(s,a) is taken as the optimal resource allocation strategy.
6. The method for dynamic resource allocation in a space flight information system based on reinforcement learning according to claim 5, characterized in that, The service requests of the user terminal include: resource usage requests of the user terminal and power allocation requests of the service area where the user terminal is located. W re This represents the resource usage state matrix corresponding to the resource usage requests of user terminals; P re This represents the power allocation matrix corresponding to the power allocation request of the service area where the user terminal is located; W ad (t) is represented in the form of the resource occupancy state matrix W; P ad (t) is represented in the form of a power allocation matrix P; Resource occupancy state matrix W = [w1, w2, ..., w n ,…,w N ], where w n =[w n,1 ,w n,2 ,...,w n,m ,...,w n,M ] T ;w n,m This indicates the occupancy status of service area n for available resource m. Power allocation matrix P = [p1, p2, ..., p n ,...,p N ].
7. The method for dynamic resource allocation in a space flight information system based on reinforcement learning according to claim 6, characterized in that, w n,m =1 indicates that service area n occupies available resources m, w n,m =0 indicates that the service area n does not occupy the available resource m.
8. The method for dynamic resource allocation in a space flight information system based on reinforcement learning according to claim 7, characterized in that, According to state s t Calculate the set of feasible actions A(s) t ),implement: According to W at time t ad (t) and P ad (t), determine the resource occupancy state matrix W corresponding to unoccupied resources. no (t) and the power allocation matrix P corresponding to the unallocated power no (t); From W no Select from (t) those that simultaneously satisfy W re Given all feasible resource allocation methods and constraint d4, where feasible resource allocation methods are represented in the form of a resource occupancy state matrix; From P no Select from (t) those that simultaneously satisfy P re Given all feasible power allocation methods for constraints d1-d3, where feasible power allocation methods are represented in the form of a power allocation matrix; Each feasible resource allocation method and power allocation method is combined into a feasible action. All feasible actions are then aggregated to form a set of feasible actions A(s). t ).
Citation Information
Patent Citations
Multi-unmanned aerial vehicle air charging and task scheduling method based on deep reinforcement learning
CN114048689A
Federal learning freshness optimization method and system based on digital twinning assistance
CN115481748A