Spectrum resource allocation method for high dynamic communication network based on multi-domain information fusion

By optimizing spectrum resource allocation through multi-domain information fusion and deep deterministic strategy gradient algorithm, the problem of insufficient multi-domain information parsing in rail transit wireless communication system is solved, realizing efficient utilization of spectrum resources and improvement of network performance.

CN119815521BActive Publication Date: 2025-11-21SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411846350.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-11-21
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

Existing technologies lack the ability to analyze multi-domain information in rail transit wireless communication systems, which limits the improvement of spectrum resource utilization, especially in the high-dynamic scenario of high-speed railways where spectrum resource allocation is not optimized.

Method used

A spectrum resource allocation method based on multi-domain information fusion is adopted. It uses residual neural networks and long short-term memory networks to predict electromagnetic wave propagation loss and traffic demand, and combines deep deterministic policy gradient algorithm to optimize spectrum resource allocation. The dynamic adjustment of spectrum resources is achieved through multi-domain information fusion mechanism and intelligent agent.

Benefits of technology

It improves the efficiency of spectrum resource allocation, enhances network performance, improves the resource utilization of wireless communication systems, and optimizes network service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119815521B_ABST
    Figure CN119815521B_ABST
Patent Text Reader

Abstract

The application relates to a high-dynamic communication network spectrum resource allocation method based on multi-domain information fusion, and belongs to the technical field of wireless communication. The method comprises the following steps: acquiring communication environment data and historical traffic demand data sets; decoupling the correlation between the communication environment data and electromagnetic wave propagation loss by using a residual neural network to obtain a next-stage electromagnetic wave propagation loss prediction value; predicting the historical traffic demand data sets by using a long short-term memory network to obtain a next-stage traffic demand prediction value; fusing the electromagnetic wave propagation loss prediction value and the traffic demand prediction value to formulate a spectrum resource allocation strategy; applying the spectrum resource allocation strategy to a railway LTE network to obtain a dynamic spectrum resource allocation strategy; and fusing the dynamic spectrum resource allocation strategy with output information of a network state monitor to optimize the dynamic spectrum resource allocation strategy. The application enhances the multi-domain information analysis capability and improves the resource utilization rate of a wireless communication system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and in particular to a method for allocating spectrum resources in a highly dynamic communication network based on multi-domain information fusion. Background Technology

[0002] In wireless communication systems for rail transit, trains inevitably traverse complex environments such as tunnels, bridges, and tall buildings during operation. These environments often trigger severe multipath interference, affecting the transmission quality of wireless signals. For high-speed railway wireless networks, mobility management is crucial, with location management playing a central role. Location management tracks and updates the locations of mobile users on trains, ensuring seamless data transmission services even at high speeds. Furthermore, spectrum resource allocation, a key technology supporting the efficient operation and high reliability of wireless communication systems, is equally indispensable. Dynamic spectrum allocation allocates spectrum resources dynamically based on actual communication needs. This method utilizes spectrum resources more efficiently, avoids spectrum fragmentation, and can dynamically adjust according to the load of the communication system. As a critical component of spectrum resource management, it plays a vital role in improving the data transmission efficiency and Quality of Service (QoS) of railway communications. The rational allocation of spectrum resources requires comprehensive consideration of various heterogeneous information, such as equipment operating status, network topology, and environmental factors, and these information are temporally correlated, including changes in equipment location and network congestion.

[0003] In high-dynamic communication scenarios such as railways, spectrum resources are even more scarce, which places higher demands on spectrum resource allocation technologies. Most existing technologies consider modeling the spectrum resource allocation problem in complex high-dynamic scenarios like railways as a nonlinear programming problem to achieve the optimization objective under the specific scenario. They then use optimization algorithms such as heuristic methods, algebraic methods, and approximate iterative methods to solve the problem and achieve optimal spectrum resource allocation. However, these technologies often lack the ability to analyze multi-domain information, limiting the improvement of resource utilization in wireless communication systems. Summary of the Invention

[0004] Therefore, the technical problem to be solved by the present invention is to overcome the limitation of insufficient multi-domain information parsing capability in the prior art, which restricts the improvement of resource utilization of wireless communication systems.

[0005] Firstly, to address the aforementioned technical problems, this invention provides a method for allocating spectrum resources in a highly dynamic communication network based on multi-domain information fusion, comprising:

[0006] Acquire communication environment data and historical traffic demand dataset; use a residual neural network to decouple the correlation between the communication environment data and electromagnetic wave propagation loss to obtain the predicted value of electromagnetic wave propagation loss for the next stage; use a long short-time memory network to predict the historical traffic demand dataset to obtain the predicted value of traffic demand for the next stage.

[0007] The predicted electromagnetic wave propagation loss value and the predicted traffic demand value are fused to form an input information set; a spectrum resource allocation strategy is formulated based on the input information set; the spectrum resource allocation strategy is applied to the railway LTE network, and the spectrum resource allocation strategy is dynamically adjusted to obtain a dynamic spectrum resource allocation strategy.

[0008] The dynamic spectrum resource allocation strategy is fused with the output information of the network status monitor to optimize the dynamic spectrum resource allocation strategy; wherein, the output information includes scenario domain information and service domain information; in the process of optimizing the dynamic spectrum resource allocation strategy, a deep deterministic policy gradient algorithm is used to select a certain continuous action in the continuous action space, and the continuous action is mapped to a discrete action that can be scheduled for resources in the communication environment.

[0009] In one embodiment of the present invention, the process of applying the spectrum resource allocation strategy to a railway LTE network includes designing a multi-domain information fusion strategy using a weighted fusion model. The method for designing the multi-domain information fusion strategy using the weighted fusion model is as follows:

[0010] The information weights of each domain are determined based on the degree of influence of each domain's information on spectrum resource allocation.

[0011] The information weights of each domain are fused to obtain the real-time network revenue function.

[0012] In one embodiment of the present invention, the calculation formula for the network real-time revenue function is as follows:

[0013]

[0014] in, For real-time network benefits, α and β represent the network's focus on transmission efficiency and real-time performance, respectively. For transmission efficiency benefit function, For real-time performance benefit function, UE g (t) represents all users, g represents base stations, G represents the set of base stations, u represents users, and p and q represent different coordination factors.

[0015] In one embodiment of the present invention, the calculation formulas for the transmission efficiency benefit function and the real-time performance benefit function are as follows:

[0016]

[0017] in, This represents the real-time transmission efficiency benefit of user u belonging to base station g. This represents the historical queue transmission efficiency gains of user u belonging to base station g. The historical average transmission efficiency benefit of user u belonging to base station g. D is the service latency budget for user u belonging to base station g. g,u (t) represents the queue delay of user u belonging to base station g at time t.

[0018] In one embodiment of the present invention, the action space needs to satisfy a constraint, which is:

[0019]

[0020] Where g is a base station, G is the set of base stations, N is the total number of base stations in the network, u is a user, V is the total number of resource blocks, and BW g This refers to the upper limit of bandwidth for base station cells within the railway LTE network. To meet the remaining resource block set of base station cells in the railway LTE network after the private network service, UE g (t) represents all users, a g,u,r (t)∈{0,1}, 0 indicates no allocation, 1 indicates allocation, r is a time-frequency resource block, b RB Where is the standard bandwidth, and j is any resource block number in the time-frequency resource block.

[0021] In one embodiment of the present invention, the optimization of the dynamic spectrum resource allocation strategy includes updating the main network and the target network; the main network includes a policy network and an evaluation network, the policy network searching for actions that maximize the value of the evaluation network; the target network includes a target policy network and a target evaluation network.

[0022] In one embodiment of the present invention, the formula for calculating the number of input nodes in the policy network is:

[0023] N input = N·M·(2V+4);

[0024] Where, N input The input node number is N, the total number of base stations in the network is M, the number of users in the corresponding cell under each base station is M, and the number of resource blocks in the corresponding cell under each base station is V.

[0025] The formula for calculating the number of output nodes of the evaluation network is:

[0026] N output =N·M·V;

[0027] Where, N output The number of output nodes.

[0028] Secondly, to solve the above-mentioned technical problems, the present invention provides a high-dynamic communication network spectrum resource allocation system based on multi-domain information fusion, comprising:

[0029] The scenario domain cognition module and the service domain cognition module are used to receive scenario domain information and service domain information output by the network status monitor, respectively, and obtain the predicted scenario domain information and predicted service domain information for the next stage.

[0030] The multi-domain information fusion module is used to fuse the prediction scenario domain information and the prediction business domain information through a multi-domain information fusion mechanism to obtain fused information, and to calculate the real-time network revenue value through the fused information;

[0031] An intelligent agent is used to receive the real-time revenue value of the network and optimize the dynamic spectrum resource allocation based on the real-time revenue value of the network; wherein the intelligent agent includes at least two main networks and a target network.

[0032] In one embodiment of the present invention, the reward function of the agent is obtained by calculating the real-time revenue value of the network.

[0033] Secondly, in order to solve the above-mentioned technical problems, the present invention provides a computer program product, which includes a computer program that, when the computer program is run, causes the above-mentioned method to be executed.

[0034] Compared with the prior art, the above-described technical solution of the present invention has the following advantages:

[0035] (1) The spectrum resource allocation method for high-dynamic communication networks based on multi-domain information fusion described in this invention constructs a comprehensive input information set by integrating predicted values ​​of electromagnetic wave propagation loss and traffic demand, providing rich contextual information for the decision-making process. By deeply analyzing this contextual information, this invention can comprehensively analyze data from different domains, accurately understand and predict the behavior of wireless communication systems, and achieve efficient allocation of spectrum resources. By utilizing long short-term memory networks to deeply analyze and predict historical traffic demand data, this method not only improves the accuracy of prediction but also optimizes resource allocation, thereby enhancing network performance. Furthermore, this invention converts continuous actions into discrete actions, ensuring the feasibility of resource scheduling in actual communication environments. This invention solves the problem of low resource utilization in wireless communication systems caused by incomplete information consideration in traditional spectrum resource allocation methods, effectively avoiding the limitation imposed by insufficient multi-domain information parsing capabilities on improving the resource utilization of wireless communication systems.

[0036] (2) This invention utilizes an intelligent agent to continuously interact with the communication environment to acquire and predict multi-domain information of the network; combined with a multi-domain information fusion mechanism, it continuously updates the long-term cumulative reward for optimizing network performance, and obtains the optimal allocation matrix of time-frequency resource blocks within the network. The wireless communication network achieves optimal allocation of spectrum resources through the optimal allocation matrix, effectively improving its own resource utilization rate. Attached Figure Description

[0037] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein...

[0038] Figure 1 This is a flowchart of a high-dynamic communication network spectrum resource allocation method based on multi-domain information fusion in a preferred embodiment of the present invention.

[0039] Figure 2 This is a schematic diagram of the residual neural network structure in a preferred embodiment of the present invention;

[0040] Figure 3 This is a schematic diagram of the convolution process in a preferred embodiment of the present invention;

[0041] Figure 4 This is a schematic diagram of the residual block unit structure in a preferred embodiment of the present invention;

[0042] Figure 5 This is a schematic diagram of the internal structure of a cell in an LSTM network according to a preferred embodiment of the present invention;

[0043] Figure 6 This is a schematic diagram of a high-dynamic communication network spectrum resource allocation method based on multi-domain information fusion in a preferred embodiment of the present invention;

[0044] Figure 7 This is a schematic diagram of the strategy network structure in a preferred embodiment of the present invention;

[0045] Figure 8 This is a schematic diagram of the evaluation network structure in a preferred embodiment of the present invention;

[0046] Figure 9 This is a schematic diagram of a high-dynamic wireless communication network scenario used for simulation in a preferred embodiment of the present invention;

[0047] Figure 10 This is a curve showing the method reward value for different numbers of users in a rural roadside scenario according to a preferred embodiment of the present invention;

[0048] Figure 11 This is a curve showing the method reward value for different numbers of users in a high-rise building area scenario according to a preferred embodiment of the present invention;

[0049] Figure 12This is a statistical result of the average throughput in a rural roadside scenario with different numbers of users in a preferred embodiment of the present invention;

[0050] Figure 13 This is the average throughput statistical result for different numbers of users in a high-rise dense area scenario according to a preferred embodiment of the present invention;

[0051] Figure 14 This is the average latency statistical result for different numbers of users in a rural roadside scenario according to a preferred embodiment of the present invention;

[0052] Figure 15 This is the average latency statistical result for different numbers of users in a high-rise dense area scenario according to a preferred embodiment of the present invention;

[0053] Figure 16 This is a statistical result of the average packet loss rate in a rural roadside scenario with different numbers of users in a preferred embodiment of the present invention;

[0054] Figure 17 This is a statistical result of the average packet loss rate in a high-rise dense area scenario with different numbers of users in a preferred embodiment of the present invention. Detailed Implementation

[0055] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0056] Example 1

[0057] Reference Figure 1 As shown, this embodiment of the invention provides a method for allocating spectrum resources in a highly dynamic communication network based on multi-domain information fusion, including:

[0058] Acquire communication environment data and historical traffic demand datasets; use residual neural networks to decouple the correlation between communication environment data and electromagnetic wave propagation loss to obtain the predicted electromagnetic wave propagation loss value for the next stage; use long short-term memory networks to predict the historical traffic demand dataset to obtain the predicted traffic demand value for the next stage.

[0059] The predicted electromagnetic wave propagation loss and traffic demand are fused to form an input information set; a spectrum resource allocation strategy is formulated based on the input information set; the spectrum resource allocation strategy is applied to the railway LTE network, and the spectrum resource allocation strategy is dynamically adjusted to obtain a dynamic spectrum resource allocation strategy.

[0060] The dynamic spectrum resource allocation strategy is optimized by fusing the output information of the network status monitor. The output information includes scenario domain information and service domain information. In the process of optimizing the dynamic spectrum resource allocation strategy, a deep deterministic strategy gradient algorithm is used to select a certain continuous action in the continuous action space and map the continuous action to a discrete action that can be scheduled for resources in the communication environment.

[0061] This invention provides a spectrum resource allocation method for highly dynamic communication networks based on multi-domain information fusion. It fuses predicted electromagnetic wave propagation loss with predicted traffic demand to form a more comprehensive input information set, providing richer contextual information for decision-making. By analyzing the contextual information to parse the multi-domain information of the wireless communication system, this invention can accurately understand and predict the behavior of the wireless communication system by combining information from different domains, thereby achieving more efficient spectrum resource allocation. Utilizing a Long Short-Term Memory (LSTM) network to predict historical traffic demand datasets not only improves prediction accuracy but also optimizes resource allocation and enhances network performance. Furthermore, this invention maps continuous actions to discrete actions, ensuring the feasibility of resource scheduling in real-world communication environments. Simultaneously, the combination of scenario domain information and service domain information provides a broader perspective for network management and optimization. This method effectively overcomes the limitations imposed by insufficient multi-domain information parsing capabilities on improving the resource utilization of wireless communication systems, providing an efficient and economical solution for spectrum resource allocation in highly dynamic communication networks.

[0062] Specifically, a residual neural network is used to decouple the correlation between these parameters and the electromagnetic wave propagation loss in the communication scenario, thereby enabling the prediction of future electromagnetic wave propagation loss values. (Refer to...) Figure 2 The residual neural network structure includes an input layer, a convolutional layer, a residual block, a fully connected layer, and an output layer. This embodiment of the invention assumes that there are q parameters related to electromagnetic wave propagation loss in the communication scenario, denoted as X = x1, x2, ..., x... q The true value of electromagnetic wave propagation loss is expressed as PL. real .

[0063] In terms of the input layer, the input layer will take input data X = x1, x2, ..., x q Mapping to a higher dimension facilitates the extraction of data features by subsequent modules in the network. The calculation formula is as follows:

[0064]

[0065] Among them, H j For the output of the input layer node, w ij For the weights of the input layer nodes, The input layer node bias is f(·), and the activation function is f(·), calculated as follows:

[0066] f(δ) = max(0,δ); (2)

[0067] Here, δ is the input to the activation function f(·).

[0068] Regarding convolutional layers, they are the core component of convolutional neural networks. The operation of a convolutional kernel on a feature map is equivalent to a sliding window, sliding back and forth across the entire feature map with a certain stride to extract features. The convolution process can be found in [reference needed]. Figure 3 The mathematical representation of the operation process of a convolutional layer is as follows:

[0069]

[0070] in, This represents the result of the convolution kernel θ operating at position [i,j] in the input matrix H. This is for outputting feature maps.

[0071] After the convolution kernel performs a convolution on the entire feature map, the size of the output feature map is calculated using the following formula:

[0072]

[0073] Among them, h in ×w in h represents the input data size. out ×w out h is the output data size. kernel ×w kernel h is the kernel size. stride ×w stride The horizontal and vertical strides represent the kernel movement, and padding is the edge padding unit.

[0074] Regarding residual blocks, they address the vanishing and exploding gradient problems through cross-layer skip connections, improving the network's convergence speed and performance. Their unit structure diagram can be found in [reference needed]. Figure 4 . Figure 4 In this context, φ is the input matrix of the residual block unit, and F(·) is the residual function.

[0075] Regarding fully connected layers, they map the features learned by convolutional layers, pooling layers, and residual blocks to the sample space. The calculation formula is as follows:

[0076]

[0077] Among them, s j For the output layer input, O k For output layer nodes, μ jk For the weights of the output layer nodes, This sets the bias for the output layer nodes.

[0078] In terms of the output layer, the output layer activates the output of the fully connected layer to increase the linearity of the network. The calculation formula of the activation function is the same as that of equation (2).

[0079] Furthermore, a Long Short-Term Memory (LSTM) network is used to predict future real-time demand for user traffic within the network using a historical user traffic dataset. This embodiment of the invention assumes that the normalized historical traffic dataset within the wireless communication network is R = (r1, r2, r3, ..., r...). T ,). Where r t Let t represent the network user traffic demand in time slot t. The prediction model forecasts the network user traffic demand in the next time slot based on the user traffic demand in the previous 20 time slots. The internal structure of a typical LSTM unit can be found in [reference needed]. Figure 5 Compared to the traditional residual neural network (RNN) structure, it adds a forget gate, input gate, and output gate to control the flow of historical information. The specific calculation formula is as follows:

[0080] f t =σ(W f [H t-1 ,x t ]+b f (7)

[0081] i t =σ(W i [H t-1 ,x t ]+b i (8)

[0082] o t =σ(W o [H t-1 ,x t ]+b o (9)

[0083]

[0084] H t =o t *tanh C t (11)

[0085]

[0086] Among them, f t i t o t The outputs of the forget gate, input gate, and output gate are respectively, σ(x) = (1+e -x ) -1To control the activation function of the gate output, based on the characteristics of the activation function σ, these three gates control the degree of influence of historical information at time t. t It is the state of the LSTM cell at time t. For candidate memory cell states, H t It is the output of the LSTM unit at time t, tanh(x) = (e x -e -x )·(e x +e -x ) -1 This refers to the activation function that controls the input and output. f b i b o b C These represent different bias terms. f W i W o W C These represent different weight matrices. t This represents the input item. Through the interaction between the three gates, LSTM is able to learn temporal information more effectively.

[0087] Specifically, the resource block rate descriptor and resource block allocation descriptor are used to characterize the detailed allocation relationship of wireless resources in the time-frequency domain. The resource block rate descriptor is a g×u×r dimensional matrix that describes the rate relationship between different users, services, and resource blocks at a given time point. This descriptor helps embodiments of the present invention better understand the resource quality in the wireless communication network. Its specific mathematical form is:

[0088]

[0089] in, R represents the resource block rate descriptor at time point t. g,u,r (t) represents the maximum transmission rate that can be achieved when the time-frequency resource block r belonging to base station g is allocated to user u at time point t.

[0090] The resource block allocation descriptor describes the allocation relationship between different users, services, and resource blocks at a specific point in time. This descriptor helps the embodiments of the present invention better understand the resource allocation situation in the wireless communication network and how to optimize resource allocation based on the actual situation. Similarly, a g×u×r dimensional matrix is ​​used to represent the mathematical form of the resource block allocation descriptor, and its specific mathematical form is as follows:

[0091]

[0092] in, a represents the resource block allocation descriptor at time point t. g,u,r(t) indicates whether the time-frequency resource block r to which base station g belongs at time point t is allocated to user u, a g,u,r (t)∈{0,1}, where 0 indicates no assignment and 1 indicates assignment.

[0093] Furthermore, in order to achieve dynamic management and optimization of spectrum resources, avoid waste and idleness of spectrum resources, and improve the utilization efficiency of spectrum resources, this embodiment of the invention analyzes the intelligent adaptation of spectrum resources under the homogeneous network of railway LTE (LTE stands for Long Term Evolution, which can be abbreviated as LTE-R).

[0094] Specifically, consider a time slot along discrete time slots An LTE-R network operating under a specific communication scenario, where T is the total number of time slots. The set of base stations in the network is represented as:

[0095] g∈G={1,2,…,N}; (15)

[0096] Where G is the set of base stations and N is the total number of base stations in the network.

[0097] At time t, all users in the cell under base station g are represented as follows:

[0098] UE g (t) = {1, 2, ..., M}; (16)

[0099] Where M represents the number of users in the corresponding cell for each base station.

[0100] At time t, the set of resource blocks of the cell under base station g is represented as:

[0101]

[0102] in, V is the set of remaining resource blocks of base station g cell in the LTE-R network after the private network service is satisfied, and V is the number of resource blocks of the corresponding cell under each base station.

[0103] Furthermore, the uniqueness of resource block allocation is analyzed, and the resource block allocation descriptor is used. The following formula should be satisfied:

[0104]

[0105] Furthermore, the bandwidth limit BW of a cell within an LTE-R network base station g is analyzed. g , The following formula should be satisfied:

[0106]

[0107] Furthermore, analyzing the limitations of LTE-R network modulation / demodulation and carrier aggregation technologies, resource blocks can only be allocated continuously in the frequency domain. The following formula should be satisfied:

[0108]

[0109] Where j is any resource block number in the time-frequency resource block. The specific meaning of equation (20) is that if user u is not allocated resources in resource block j of cell within base station g, then resources will not be allocated in resource blocks with higher numbers (r≥j+1) either.

[0110] For LTE-R homogeneous networks, their time-frequency resource blocks have a standard bandwidth of b RB It can adaptively adjust its transmission capacity based on channel quality. Within each TTI, the LTE-R resource scheduler allocates all available resource blocks not allocated to private network services to users accessing each cell base station. This embodiment of the invention assumes a minimum scheduling period (TTI) of 1 ms and requires solving the problem of finding the optimal allocation descriptor for resource blocks in the entire LTE-R network.

[0111] Multi-Domain Information Fusion (MDIF) is a unique technique for the comprehensive diagnosis and evaluation of networks or systems. It integrates information from various domains to comprehensively evaluate the tested object according to specific criteria, thereby deriving a unified prediction and assessment. Compared to single-domain information processing techniques, MDIF can analyze the experimental object from multiple levels, aspects, and perspectives. This invention employs a weighted fusion model to design a MDIF strategy and fuses information according to the weights of each domain.

[0112] Specifically, the information weights are determined based on the degree to which the information in each domain affects spectrum resource allocation. In this way, a comprehensive set of information, namely the network's real-time revenue function, can be obtained. Used to guide the allocation of spectrum resources. The specific calculation formula is as follows:

[0113]

[0114] Here, α and β are two weighting parameters that represent the network's focus on transmission efficiency and real-time performance, respectively. Let be the transmission efficiency benefit function, which represents the transmission efficiency gain generated by base station g performing a resource block scheduling at time t; Let p be the real-time performance benefit function, representing the real-time performance gain generated by base station g performing a resource block scheduling at time t; p and q are different coordination factors. In this embodiment of the invention, the preferred values ​​for the two coordination factors are: p = 3, q ​​= 1.

[0115] Specifically, the transmission efficiency benefit function The specific calculation formula is as follows:

[0116]

[0117] in, This represents the real-time transmission efficiency benefit of user u belonging to base station g. This represents the historical queue transmission efficiency gains of user u belonging to base station g. This represents the historical average transmission efficiency benefit of user u belonging to base station g.

[0118] Furthermore, The factor that has the greatest impact on the user's online transmission experience, taking into account information from the scenario domain, business domain, and resource domain, is calculated using the following formula:

[0119]

[0120] in, Let g be the upper bound of the real-time transmission rate of the service of user u belonging to base station g. R is the lower bound of the real-time transmission rate of the service of user u belonging to base station g. g,u (t) represents the real-time transmission rate, R g,u The formula for calculating (t) is:

[0121]

[0122] Furthermore, The factor that has the greatest impact on user-reliable transmission, taking into account information from the scenario domain, business domain, and resource domain, is calculated using the following formula:

[0123]

[0124] Among them, Q g,i (t) represents the queue length of user u belonging to base station g at time t, and TTI represents the transmission time interval.

[0125] Furthermore, The factor that has the greatest impact on fairness among users is business domain information, and its calculation formula is as follows:

[0126]

[0127] in, Let g be the expected average historical transmission rate of user u belonging to base station g. Let g be the historical transmission rate of user u belonging to base station g at time t. and The calculation formula is:

[0128]

[0129] Among them, L g,u The size of the service data packet for user u to which base station g belongs. For L g,u The probability density distribution function of T; g,u The packet interval for the service data packets of user u to which base station g belongs. For T g,u The probability density distribution function; L g,u (τ) represents the size of the transmitted data of user u belonging to base station g at time τ; As an adjustable constant, its calculation formula is:

[0130]

[0131] Among them, E(T) g,u ) represents the average packet interval for the service data of user u to which base station g belongs.

[0132] Specifically, the real-time performance benefit function The calculation formula mainly considers business domain information and is as follows:

[0133]

[0134] in, D is the service latency budget for user u belonging to base station g. g,u (t) represents the queue delay of user u belonging to base station g at time t.

[0135] This invention proposes an intelligent allocation method for spectrum resources in high-speed railway mobile communication networks based on multi-domain information reinforcement learning. The method framework diagram can be found in the following figures. Figure 6 . Figure 6 It consists of two parts: a highly dynamic communication environment based on multi-domain information fusion and an intelligent agent. In the highly dynamic communication environment based on multi-domain information fusion, the scene domain cognitive model and the service domain cognitive model receive scene domain information and service domain information output by the network state monitor to predict future information. The multi-domain information fusion mechanism calculates the network's real-time benefit value r by fusing the information. t The evaluation data is then fed back to the agent to improve performance. The agent has two main networks and two target networks. The main networks consist of a policy network μ and an evaluation network Q, with network parameters θ and θ, respectively. μ and θ QThe goal of the policy network is to find the action that maximizes the Q-value of the evaluation network, and then use this action to guide the high-dynamic communication network to achieve intelligent allocation of time-frequency resource blocks. The target network consists of a target policy network μ′ and a target evaluation network Q′, with network parameters θ. μ′ and θ Q′ The main network and the target network have identical structures, but the network parameters are updated at different frequencies. The main network updates once per step, while the target network updates once after the main network has been updated multiple times.

[0136] The considered state includes the user's transmission capacity information on each time-frequency resource block within the base station at the current time, user queue length information, user historical transmission rate information, user queue latency information, and the user's traffic demand information and transmission capacity information on each time-frequency resource block within the base station at the next time step. State s at time t. t Mathematical form is expressed as:

[0137]

[0138] Among them, R max This provides users with the ability to transmit time-frequency resource blocks under optimal channel conditions. The historical maximum queue length of user M belonging to base station N. This represents the predicted service traffic value for user M belonging to base station N.

[0139] Specifically, this embodiment of the invention utilizes the characteristics of the deep deterministic policy gradient algorithm to select a specific action from a continuous action space, rather than calculating the specific probabilities of all possible actions. Then, this continuous action is mapped to discrete actions that can be scheduled for resources in the communication environment. This method leverages the continuous decision-making characteristic of the deep deterministic policy gradient algorithm, thereby effectively solving the problem of an excessively large action space. At time t, action a... t Mathematical form is expressed as:

[0140]

[0141] Where, δ N,M,V The value of (t) indicates whether user M, to which base station N belongs, occupies its V-th time-frequency resource block, and its value range is {0,1}; 1 indicates occupation, and 0 indicates non-occupation. It should be noted that the action space should satisfy the constraints (18)(19)(20) mentioned above.

[0142] Furthermore, the network real-time revenue function calculated by equation (21) is... As a reward function, the reward function is defined as follows:

[0143]

[0144] Furthermore, the design of the policy (actor) network and the evaluation (critic) network is based on deep neural network technology, aiming to optimize the policy and evaluate the decision-making process. The structures of the actor and critic networks can be found in [reference needed]. Figure 7 and Figure 8 .

[0145] For the actor network, a fully connected network architecture is adopted, taking the environment state s as input and outputting the corresponding action a. According to the definition of the state space in equation (31), the number of nodes in the input layer of the actor network can be determined, and the mathematical expression is:

[0146] N input =N·M·(2V+4); (34)

[0147] Based on the definition of the action space in equation (32), the number of output nodes of the actor network can be determined as follows:

[0148] N output = N·M·V; (35)

[0149] Reference Figure 7 The actor network has four hidden layers: h1, h2, h3, and h4, with 500, 1000, 1000, and 500 nodes respectively. Except for the output layer, the ReLU activation function is preferred for each layer. The calculation formula for the ReLU activation function is the same as in equation (2). The output layer preferably uses the Sigmoid function, calculated as follows:

[0150]

[0151] For the critic network, it is also a fully connected network. The input is a combination of environment state s and action a, and the output is the Q-value of that state-action pair. The number of nodes in the input layer of the critic network is equal to...

[0152] N input =N·M·(3V+4); (37)

[0153] Furthermore, the number of output nodes in the critic network is

[0154] N output =1; (38)

[0155] Reference Figure 8 The critic network has two hidden layers, h1 and h2, with 128 and 64 nodes respectively. Except for the output layer, all layers use the ReLU activation function, and the output layer uses the Sigmoid function.

[0156] In the process of finding the optimal intelligent spectrum resource allocation strategy, the agent selects action a from the policy network μ. t =μ(s) t )+n. Where n represents the noise from the action to avoid local optima, and n~N(0,σ). After the communication network performs this action, a new state s will be generated. t+1 and rewards r t and the quadruple samples (s) t ,a t ,r t ,s t+1 The data are placed into the experience replay pool. When the experience replay pool is full, Z sample data are taken out and input into the policy network and evaluation network for training and updating the network parameters.

[0157] Specifically, the network parameters θ of the critic network Q. Q The gradient formula is:

[0158]

[0159] Wherein, the target value y t The calculation formula is:

[0160] y t =r t +γQ′(s t+1 ,a t+1 (40)

[0161] Where γ is the discount factor.

[0162] Specifically, the network parameters θ of the actor network μ μ The gradient formula is as follows:

[0163]

[0164] in, It evaluates the action gradient of the network. It is the policy gradient of the policy network.

[0165] Furthermore, the target evaluation network and the target policy network update their network parameters using a soft update method, with the specific update formula as follows:

[0166] θ Q′ =δθ Q +(1-δ)θ Q′ (42)

[0167] θ μ′ =δθ μ +(1-δ)θ μ′ (43) where δ is the soft update coefficient.

[0168] Therefore, the target evaluation network θ that satisfies the conditions can be obtained using the iterative gradient descent method. Q′ and target policy network θ μ′ .

[0169] This invention addresses the problem of low resource utilization in wireless communication systems caused by incomplete information consideration in existing spectrum resource allocation methods. It proposes a highly dynamic spectrum resource allocation method for communication networks based on multi-domain information fusion. This method integrates multi-domain information technology and reinforcement learning algorithms to achieve intelligent allocation of spectrum resources. An intelligent agent continuously interacts with the communication environment to collect and predict multi-domain information of the network. Utilizing the multi-domain information fusion mechanism, it dynamically updates and optimizes the long-term cumulative reward for network performance, thereby obtaining the optimal allocation matrix for time-frequency resource blocks within the network. Based on this optimal allocation matrix, the wireless communication network can achieve efficient allocation of spectrum resources, significantly improving resource utilization and optimizing network service quality. Furthermore, this invention's embodiment of accurately capturing information based on contextual relationships has a significant effect on optimizing the performance of rail transit communication networks. In-depth analysis of contextual information and fusion analysis of multi-source data are particularly crucial, providing strategic guidance for the continuous optimization of rail transit network performance.

[0170] To more clearly illustrate the specific content of the embodiments of the present invention, the following examples are provided.

[0171] Example: A high-dynamic wireless communication network simulation platform was built based on the NS3 simulation platform. Users are randomly distributed within the base station coverage area, and each base station has a certain degree of overlapping coverage. User communication services are generated according to the relevant service characteristic information in Table 1. Considering network load, this example gradually increases the number of services from 6 to 72, and sets two different service ratios, namely service ratio one (Num... Volte :Num Video :Num URLLC =2:3:1) and business ratio two (Num Volte :Num Video :Num URLLC =1:1:1)(where Num Volte Num Video Num URLLC (These represent the number of VoLTE, Video, and URLLC services, respectively), to simulate different distributions of service demand. Considering scenario domain information, this example conducts experiments in two scenarios: densely populated high-rise areas and rural areas. Specific scenarios are as follows: Figure 9 As shown.

[0172] Table 1. Characteristics of Wireless Communication Services

[0173]

[0174] In the method described in this example, the design of the reward function is crucial. The reward function defines the agent's objective: to learn how to take actions to maximize its cumulative reward. In highly dynamic scenarios, the reward function is designed to reflect network performance, including metrics such as throughput and latency.

[0175] In the network simulation platform set up in this invention, simulation experiments are conducted to compare the changes in the reward value of the agent under different network load conditions and different scenarios. Figure 10 and Figure 11 The figures show the reward value changes during training of the proposed method under four scenarios: business ratio 1 and 2 in a rural roadside scenario, and business ratio 1 and 2 in a high-rise dense area scenario. The figures show that the reward value obtained by the proposed method eventually stabilizes when the number of users is 6, 12, 18, 24, 48, and 72, indicating that the method can reach convergence. The training curves show that the reward value decreases as the number of users increases. This is because when the network load is low, the agent can obtain a higher reward through a simple strategy, such as evenly distributing resources. However, as the network load increases, this strategy becomes ineffective. The agent needs to learn more complex strategies to obtain higher rewards. Furthermore, the reward obtained by the agent in the business ratio 1 scenario is slightly higher than that in the business ratio 2 scenario. This is because the data generation of video streaming services is much smaller than that of URLLC services. When the network load is too high, even if the agent adopts the optimal resource allocation strategy, the physical limitations of resources cannot meet the needs of all users, leading to a decline in network performance. Additionally, the reward obtained by the agent in the high-rise dense area scenario is slightly higher than that in the rural roadside scenario. This is because the communication environment in rural areas is more complex than in densely populated high-rise areas, and users are more mobile. As a result, the agent needs to make resource allocation decisions more frequently, which places higher demands on the agent's learning ability.

[0176] Based on the above analysis, it can be seen that the agent can learn to find effective resource allocation strategies under different scenarios and network load conditions. Although the agent's performance may decline slightly in some situations, such as in rural areas or when the network load is too high, the agent can still maintain high performance overall.

[0177] The trained proposed method is compared with the maximum carrier-to-interference ratio (MCR) method, proportional fairness algorithm, round-robin algorithm, and channel-aware algorithm. This example statistically analyzes the average throughput, average latency, and average packet loss rate of each algorithm under different user scenarios. Figures 12 to 17 As shown.

[0178] Figure 12 and Figure 13The graphs show the average throughput statistics for VoLTE, Video, and URLLC services for each algorithm in rural and high-rise dense area scenarios, respectively.

[0179] from Figure 12 and Figure 13 As can be seen, compared to the four algorithms of maximum carrier-to-interference ratio, proportional fairness, round-robin, and channel awareness, the proposed method improves the average throughput of each service to a certain extent. It is evident that, compared to the other four algorithms, the proposed method better guarantees the average throughput performance of URLLC services while maintaining the expected average throughput of VoLTE and Video services. This indicates that the proposed method can effectively improve the average throughput of the system compared to traditional classical methods. However, as the number of users increases, the improvement in average throughput of the proposed method compared to traditional classical methods decreases. This is because under high load conditions, the needs of each user become more complex, and the near-optimal allocation method given by the agent can only try to meet the communication needs of each user, making it difficult to find a precise numerical allocation matrix. Similarly, compared to the second service ratio scenario, the agent's improvement in network average throughput is more significant in the first service ratio scenario. Furthermore, although the reward value obtained by the agent in the rural roadside scenario is slightly lower than that in the high-rise dense area scenario, the spectrum allocation strategy it provides improves the average throughput more than that in the high-rise dense area scenario. This also indicates that the agent has strong adaptability to different scenarios.

[0180] Figure 14 and Figure 15 The figures show the average latency statistics for VoLTE, Video, and URLLC services for each algorithm in rural and high-rise dense area scenarios, respectively.

[0181] from Figure 14 and Figure 15As can be seen, compared to the four algorithms of maximum carrier-to-interference ratio, proportional fairness, round-robin, and channel-awareness, the proposed method improves the average latency for various services. It is evident that the proposed method outperforms the other four algorithms in handling average latency regardless of the number of users. This indicates that the proposed method has a significant advantage over traditional classical algorithms in reducing the system's average latency. Moreover, the proposed method optimizes average latency more effectively in rural roadside scenarios than in densely populated high-rise areas. However, it is worth noting that the advantage of the proposed method in reducing average latency gradually diminishes as the number of users increases. This may be attributed to the increased complexity of user demands in high-load network environments, causing the agent to provide only an approximately optimal resource allocation scheme, rather than finding a precise numerical allocation matrix. Furthermore, in the case of service ratio one, the agent performs more effectively in reducing the network's average latency compared to service ratio two. This result further confirms that under more concentrated service demands, the agent can more effectively optimize network performance and reduce average latency.

[0182] Figure 16 and Figure 17 The figures show the average packet loss rates for each algorithm's VoLTE, Video, and URLLC services in rural and high-rise densely populated areas, respectively.

[0183] from Figure 16 and Figure 17 As can be seen, compared to the four algorithms of maximum carrier-to-interference ratio, proportional fairness, round-robin, and channel awareness, the proposed method improves the average packet loss rate for various services. Clearly, compared to the other four algorithms, the proposed method performs well in terms of average packet loss rate for all three service types across different user numbers. This indicates that the proposed method is more advantageous than traditional classical algorithms in reducing the system's average packet loss rate. Similarly, the proposed method optimizes average latency better in rural roadside scenarios than in densely populated high-rise areas. However, as the number of users increases, the advantage of the proposed method in reducing the average packet loss rate gradually diminishes. Furthermore, compared to service ratio two, the agent performs more effectively in reducing the network's average packet loss rate in service ratio one. This further demonstrates that when service demands are more concentrated, the agent can more effectively optimize network performance and reduce the average packet loss rate.

[0184] comprehensive Figures 12 to 17As can be seen, the time-frequency resource block allocation strategy obtained by the method proposed in this invention will significantly improve the system performance of highly dynamic wireless communication systems compared to the other four types of algorithms. This is mainly because the method proposed in this embodiment considers the joint information of the scenario domain, service domain, and resource domain. It not only considers the current scenario and number of services but also the available resources. In this way, the method can achieve intelligent optimization allocation of spectrum resources under various scenarios and service quantities. In addition, the method also has adaptability. The intelligent agent can dynamically adjust its behavior according to different scenarios, different numbers of users, and different service distribution states to meet the needs of each user. That is, the intelligent agent can adjust its own behavior according to the real-time state of the network to optimize network performance.

[0185] Example 2

[0186] Based on the same inventive concept, this embodiment provides a high dynamic communication network spectrum resource allocation system based on multi-domain information fusion. The principle of solving the problem is similar to the high dynamic communication network spectrum resource allocation method based on multi-domain information fusion provided in Embodiment 1, and the repeated parts will not be described again.

[0187] This embodiment provides a high-dynamic communication network spectrum resource allocation system based on multi-domain information fusion, including:

[0188] The scenario domain cognition module and the service domain cognition module are used to receive scenario domain information and service domain information output by the network status monitor, respectively, and obtain the predicted scenario domain information and predicted service domain information for the next stage.

[0189] The multi-domain information fusion module is used to fuse prediction scenario domain information and prediction business domain information through a multi-domain information fusion mechanism to obtain fused information, and to calculate the real-time network revenue value through the fused information;

[0190] An intelligent agent is used to receive real-time network revenue values ​​and optimize dynamic spectrum resource allocation based on these values; the intelligent agent includes at least two main networks and a target network.

[0191] Specifically, the reward function of the intelligent agent is obtained by calculating the real-time network revenue value, and the calculation formula for the real-time network revenue value can be referred to formula (21) in Example 1.

[0192] In practical applications, the high-dynamic communication network spectrum resource allocation system based on multi-domain information fusion provided in this invention is suitable for high-speed railway wireless communication networks. When a train is traveling on the track, the system can collect key parameters of the communication environment in real time through sensors. These parameters include the train's precise location, the antenna's installation height, the specific scene type, and the currently used communication frequency band. Through real-time collection and analysis of this data, the system can more effectively manage and optimize spectrum resources in the high-speed railway wireless communication network, thereby improving communication efficiency and quality.

[0193] Example 3

[0194] This embodiment provides a computer program product, which includes a computer program that, when run, causes the method described in Embodiment 1 to be executed.

[0195] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0196] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0197] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0198] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0199] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for allocating spectrum resources in highly dynamic communication networks based on multi-domain information fusion, characterized in that... ,include: Acquire communication environment data and historical traffic demand dataset; use a residual neural network to decouple the correlation between the communication environment data and electromagnetic wave propagation loss to obtain the predicted value of electromagnetic wave propagation loss for the next stage; use a long short-time memory network to predict the historical traffic demand dataset to obtain the predicted value of traffic demand for the next stage. The predicted electromagnetic wave propagation loss and the predicted traffic demand are fused to form an input information set; a spectrum resource allocation strategy is formulated based on the input information set; the spectrum resource allocation strategy is applied to the railway LTE network, and the spectrum resource allocation strategy is dynamically adjusted to obtain a dynamic spectrum resource allocation strategy; wherein, the process of applying the spectrum resource allocation strategy to the railway LTE network includes designing a multi-domain information fusion strategy using a weighted fusion model, and the method for designing the multi-domain information fusion strategy using the weighted fusion model is as follows: The information weights of each domain are determined based on the degree of influence of each domain's information on spectrum resource allocation. The information weights of each domain are fused to obtain a real-time network revenue function; the real-time network revenue value is obtained based on the real-time network revenue function; and the dynamic spectrum resource allocation is optimized based on the real-time network revenue value. The dynamic spectrum resource allocation strategy is fused with the output information of the network status monitor to optimize the dynamic spectrum resource allocation strategy; wherein, the output information includes scenario domain information and service domain information; in the process of optimizing the dynamic spectrum resource allocation strategy, a deep deterministic policy gradient algorithm is used to select a certain continuous action in the continuous action space, and the continuous action is mapped to a discrete action that can be scheduled for resources in the communication environment.

2. The method for allocating spectrum resources in a high-dynamic communication network based on multi-domain information fusion according to claim 1, characterized in that, The formula for calculating the real-time revenue function of the network is as follows: in, For real-time network benefits, α and β represent the network's focus on transmission efficiency and real-time performance, respectively. For transmission efficiency benefit function, For real-time performance benefit function, UE g (t) represents all users, g represents base stations, G represents the set of base stations, u represents users, and p and q represent different coordination factors.

3. The method for allocating spectrum resources in a high-dynamic communication network based on multi-domain information fusion according to claim 2, characterized in that, The calculation formulas for the transmission efficiency benefit function and the real-time performance benefit function are as follows: in, This represents the real-time transmission efficiency benefit of user u belonging to base station g. This represents the historical queue transmission efficiency gains of user u belonging to base station g. The historical average transmission efficiency benefit of user u belonging to base station g. D is the service latency budget for user u belonging to base station g. g,u (t) represents the queue delay of user u belonging to base station g at time t.

4. The method for allocating spectrum resources in a high-dynamic communication network based on multi-domain information fusion according to claim 1, characterized in that, The action space needs to satisfy a constraint, which is: Where g is a base station, G is the set of base stations, N is the total number of base stations in the network, u is a user, V is the total number of resource blocks, and BW g This refers to the upper limit of bandwidth for base station cells within the railway LTE network. To meet the remaining resource block set of base station cells in the railway LTE network after the private network service, UE g (t) represents all users, a g,u,r (t)∈{0,1}, 0 indicates no allocation, 1 indicates allocation, r is a time-frequency resource block, b RB Where is the standard bandwidth, and j is any resource block number in the time-frequency resource block.

5. The method for allocating spectrum resources in a high-dynamic communication network based on multi-domain information fusion according to claim 1, characterized in that, The optimization of the dynamic spectrum resource allocation strategy includes the updating of the main network and the target network; the main network includes a policy network and an evaluation network, the policy network searches for actions that maximize the value of the evaluation network; the target network includes a target policy network and a target evaluation network.

6. The method for allocating spectrum resources in a high-dynamic communication network based on multi-domain information fusion according to claim 5, characterized in that, The formula for calculating the number of input nodes in the policy network is: N input =N·M·(2V+4); Where, N input The input node number is N, the total number of base stations in the network is M, the number of users in the corresponding cell under each base station is M, and the number of resource blocks in the corresponding cell under each base station is V. The formula for calculating the number of output nodes of the evaluation network is: N output =N·M·V; Where, N output The number of output nodes.

7. A computer program product, said computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Spectrum demand calculation method and apparatus

    CN105451242A

  • Air-ground joint multi-domain detection system and method for unmanned aerial vehicle group target

    CN115542318A