Network resource allocation method and device, electronic equipment and computer program product
Through the combination of variational modal decomposition and deep Q network reinforcement learning model, the high burden and high power consumption problems caused by direct training of drones received signals are solved, and efficient resource allocation and real-time performance optimization of 6G air-ground integrated network are achieved.
Patent Information
- Application Number
- CN202510664030.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, various signals received by drones are directly used as deep learning training data, resulting in problems such as long training time, high burden, high computing complexity and high power consumption, making it difficult to achieve real-time resource allocation in the 6G space-space integrated network.
Variable mode decomposition technology is used to decompose the historical network signals received by the drone into multiple historical eigenmodal function sub-signals, and combine long and short-term memory networks and deep Q network reinforcement learning models to build a network resource prediction model and dynamically optimize resource allocation.
It reduces the training complexity and training time of the network resource prediction model, reduces the power consumption of the drone, realizes efficient allocation of spectrum, time and space resources, and improves the overall performance of the drone-assisted 6G air-ground integrated network and the service quality of ground users.
Smart Images

Figure CN120434787A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of communication technology, and in particular relates to a network resource allocation method, device, electronic device, and computer program product. Background Art
[0002] With the large-scale commercialization of fifth-generation mobile communication systems and the increasing demand for new intelligent services such as connected vehicles and immersive extended reality, demand for extreme communication network performance has become increasingly prominent. Research on sixth-generation mobile communication systems is in full swing worldwide. 6G wireless communication networks are expected to achieve global, three-dimensional, and full-area coverage, meeting differentiated service needs across all scenarios and ultimately realizing the "intelligent connection of all things." To achieve this ambitious goal, integrated air-space-ground networks and spatial three-dimensional networks have become a hot topic in current network research. As a key component of 6G integrated air-space-ground networks, drones (UAVs) have demonstrated tremendous development potential and market demand in the communications field, particularly driven by the low-altitude economy. Drones can assist terrestrial wireless communication networks in expanding coverage and increasing capacity, providing emergency communications and flexible networking capabilities, and have become a new driving force for the development of 6G integrated air-space-ground networks. However, UAV-assisted 6G integrated air-space-ground communication networks also face numerous challenges, including improving resource utilization and energy efficiency, ensuring low latency and fairness, and implementing effective mobility management. The development of Shenzhen's eight future industries has put new demands on resource allocation algorithms for UAV-assisted 6G integrated air-space-ground networks. Especially with the rise of the low-altitude economy, the demand for the application of drones in logistics distribution, urban monitoring, environmental monitoring, etc. is increasing, and the research on resource allocation algorithms has urgent practical significance.
[0003] There has been extensive research on drones both domestically and internationally, with early efforts focusing on obstacle avoidance, path planning, and power conservation. With the advancement of AI, various AI technologies are being applied to future 6G networks. A multi-agent drone collaboration algorithm based on deep learning can use drones (UAVs) as aerial base stations (BSs), providing wireless communication services cost-effectively and on-demand. Each UAV automatically selects a ground user, power level, and subchannel to communicate with without exchanging any information with other UAVs. To model the dynamics and uncertainty of the environment, this technique formulates the long-term resource allocation problem as a stochastic game to maximize expected payoff. Each UAV becomes a learning agent, and each resource allocation solution corresponds to an action taken by the UAV. In a multi-agent reinforcement learning (MARL) framework, each agent learns to discover its optimal strategy based on its local observations. In this approach, all agents independently execute decision-making algorithms but share a common structure based on Q-learning. This technique improves the performance of the proposed multi-agent reinforcement learning-based resource allocation algorithm by properly setting exploration and exploitation parameters. Furthermore, compared to other techniques that rely on full information exchange between drones, the multi-agent reinforcement learning algorithm delivers acceptable performance with reduced overhead. Thus, this technique achieves a good balance between performance improvement and information exchange overhead.
[0004] Although existing technologies have improved the performance of communication between drones and ground users and achieved a good balance with information exchange overhead, in the future 6G network, there will be a variety of signals. Drones assist ground base stations and ground users in communication. If the various signals received by drones are directly used as training data for deep learning to train the model, this will undoubtedly greatly increase the burden on the machine learning model, and increase the training time of machine learning and the power consumption of the drone, resulting in problems such as high computational complexity, weak real-time processing capabilities and high power consumption of existing technologies. It is difficult to ensure that drones can dynamically adjust parameters according to the network environment for real-time processing. Summary of the Invention
[0005] The embodiments of the present application provide a network resource allocation method, apparatus, electronic device, and computer program product, which can solve the technical problems in the prior art of directly using various signals received by drones as training data for deep learning, resulting in long training time, high burden, high computational complexity, and high power consumption.
[0006] In a first aspect, an embodiment of the present application provides a network resource allocation method, including:
[0007] Based on variational mode decomposition, the historical network signal received by the UAV is decomposed into one or more historical intrinsic mode function sub-signals;
[0008] Performing training based on the data of the one or more historical intrinsic mode function sub-signals to obtain a network resource prediction model;
[0009] Inputting one or more real-time intrinsic mode function sub-signals into the network resource prediction model to obtain a network resource prediction value output by the network resource prediction model, wherein the one or more real-time intrinsic mode function sub-signals are obtained based on variational mode decomposition of real-time network signals received by the UAV, and the parameters of the network resource prediction model are adjusted according to variability parameters in the network environment and a deep Q network reinforcement learning model;
[0010] Network resources are allocated according to the predicted value of the network resources.
[0011] In a possible implementation of the first aspect, decomposing the historical network signal received by the drone into one or more historical intrinsic mode function sub-signals based on variational mode decomposition includes:
[0012] The goal of performing variational modal decomposition on the historical network signal is set to minimize bandwidth constraints and minimize reconstruction error constraints, wherein the bandwidth minimization constraint is used to constrain the instantaneous frequency fluctuation of each historical intrinsic mode function sub-signal to be minimized, and the reconstruction error minimization constraint is used to constrain the superposition of each historical intrinsic mode function sub-signal to be equal to the historical network signal;
[0013] Initialize the number of historical intrinsic mode function sub-signals, penalty factor, convergence threshold, historical intrinsic mode function sub-signals, center frequency, maximum frequency, and Lagrange multiplier;
[0014] Iteratively update the historical intrinsic mode function sub-signals, center frequencies, and Lagrange multipliers and perform convergence judgment based on the minimized bandwidth constraint and the minimized reconstruction error constraint. If the residual of one or more historical intrinsic mode function sub-signals is less than a convergence threshold or the number of iterative updates reaches a maximum number of iterations, terminate the iterative update and obtain one or more historical intrinsic mode function sub-signals corresponding to the historical network signal. Otherwise, return to iterative update.
[0015] In a possible implementation of the first aspect, the training based on data of the one or more historical intrinsic mode function sub-signals to obtain a network resource prediction model includes:
[0016] Performing data preprocessing on the data of the one or more historical intrinsic mode function sub-signals to obtain a data set of historical intrinsic mode function sub-signals after data preprocessing;
[0017] Constructing a training data matrix based on the data set of the historical intrinsic mode function sub-signal, constructing a time series, and dividing the data into a training set, a validation set, and a test set;
[0018] Construct an initial network resource prediction model, train and optimize the initial network resource prediction model, and obtain a network resource prediction model.
[0019] In a possible implementation of the first aspect, the deep Q network reinforcement learning model is trained based on the following steps:
[0020] Initialize the main network, target network, experience pool, and exploration rate;
[0021] A deep Q network reinforcement learning model is obtained by training in a cyclic training manner. The cyclic training includes:
[0022] Select actions, randomly select actions based on the exploration rate to promote the algorithm's exploration of the environment, or select the action with the largest current Q value;
[0023] Execute actions and store experience: execute the selected action, obtain the corresponding reward and next state, and store the current state, selected action, reward, and next state into the experience pool;
[0024] Training the network involves sampling a batch of data from the experience pool, calculating the target Q value and loss based on the sampled data, and using the optimizer to update the parameters of the main network so that the network continuously optimizes its prediction of the Q value.
[0025] Update the target network and copy the parameters of the main network to the target network.
[0026] In a second aspect, an embodiment of the present application provides a network resource allocation device, including:
[0027] A decomposition module is used to decompose the historical network signal received by the UAV into one or more historical intrinsic mode function sub-signals based on variational mode decomposition;
[0028] A training module, configured to perform training based on data of the one or more historical intrinsic mode function sub-signals to obtain a network resource prediction model;
[0029] a prediction module, configured to input one or more real-time intrinsic mode function sub-signals into the network resource prediction model to obtain a network resource prediction value output by the network resource prediction model, wherein the one or more real-time intrinsic mode function sub-signals are obtained based on variational mode decomposition of real-time network signals received by the UAV, and parameters of the network resource prediction model are adjusted based on variability parameters in the network environment and a deep Q network reinforcement learning model;
[0030] The allocation module is used to allocate network resources according to the predicted value of network resources.
[0031] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements a method as described in any one of the first aspects above.
[0032] In a fourth aspect, an embodiment of the present application provides a computer program product, which, when run, enables the method described in any one of the above-mentioned first aspects to be executed.
[0033] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method as described in any one of the above-mentioned first aspects is implemented.
[0034] Compared with the prior art, the embodiments of the first aspect of the present application have the following beneficial effects:
[0035] The embodiment of the present application proposes a method that can adapt to the changing parameters of the network and dynamically optimize resource allocation by comprehensively using adaptive signal data decomposition, network resource prediction models such as long short-term memory networks, and algorithm combination model iteration and optimization of deep reinforcement learning, thereby achieving efficient allocation of network resources such as spectrum, time and space, so as to maximize the overall performance of the drone-assisted 6G air-ground integrated network, while ensuring that the service quality of ground users is not affected. The embodiment of the present application is based on interdisciplinary fusion innovation and is expected to promote the research on drone-assisted wireless communication network resource allocation algorithms into a new stage. This application not only has in-depth theoretical research significance, but also has realistic practical application value. The research results of this application can provide strong support for the development of the low-altitude economy.
[0036] In response to the problem of large data volume in the 6G air-ground integrated network, the embodiment of the present application designs an adaptive signal decomposition technology to decompose the received 6G mixed signal into independent signals of their respective frequencies, and extract the characteristic values such as frequency, amplitude and modulation mode of the signals of different communication standards. These characteristic values are then used together with parameters such as the transmission power and height of the aerial base station as training data to train network resource prediction models such as deep learning models. This can reduce the training complexity and training time of network resource prediction models such as deep learning models, and reduce the power consumption of aerial base stations.
[0037] In response to the problem that network resource prediction models such as deep learning models may find it difficult to capture the complex spatiotemporal characteristics in the data, the embodiments of the present application propose a long short-term memory network model. By building a long short-term memory network model, the characteristic values such as frequency, amplitude and modulation mode of different communication standard signals in the 6G mixed signal extracted by adaptive signal decomposition technology, as well as parameters such as the transmission power and height of the aerial base station are used as training data to train the long short-term memory network model, construct a 6G air-ground integrated network resource prediction model, predict and allocate network resources, and then adjust the position of the aerial base station according to the prediction results to meet the needs of maximizing network capacity.
[0038] To address the problem of deep learning model prediction lag caused by the rapid changes in 6G signal characteristics, the embodiments of the present application build a deep Q network reinforcement learning model and utilize the real-time online learning capabilities of deep reinforcement learning to continuously optimize and adjust the deep learning decision-making process in real time according to device location, business demand changes, and changes in various channel environment parameters. It dynamically adjusts the allocation strategy of network resources and the location of aerial base stations in real time, which can better meet real-time requirements.
[0039] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0041] Figure 1 This is a flow chart of a network resource allocation method provided by an embodiment of the present application;
[0042] Figure 2 This is a schematic diagram of the structure of a network resource allocation device provided in one embodiment of the present application;
[0043] Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0045] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0046] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0047] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0048] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0049] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0050] Figure 1 This is a flowchart of a network resource allocation method provided in one embodiment of the present application.
[0051] The network resource allocation method provided in this application can solve the technical problems in the prior art of directly using various signals received by drones as training data for deep learning, which leads to long training time, high burden, high computational complexity and high power consumption.
[0052] S11, based on variational mode decomposition, decomposes the historical network signal received by the UAV into one or more historical intrinsic mode function sub-signals.
[0053] Historical network signal data from drone-assisted 6G air-ground integrated networks in a region can be collected to create a data set. Using variational mode decomposition (VMD), the historical 6G network millimeter-wave network signals received by drones are decomposed into several historical intrinsic mode function (IMF) sub-signals (referred to as modes, modal components, or modal component signals) in different frequency ranges. The bandwidth of each modal component signal is minimized, and the data of each historical IMF sub-signal is arranged from high to low frequency. The arranged historical IMF sub-signals are then filtered and fused.
[0054] Based on variational mode decomposition, the historical network signals received by the drone are decomposed into one or more historical intrinsic mode function sub-signals, which reduces the difficulty of training the subsequent network resource prediction model, alleviates the edge computing burden of the drone, and slows down the power consumption of the drone power supply.
[0055] S12: Perform training based on the data of the one or more historical intrinsic mode function sub-signals to obtain a network resource prediction model.
[0056] Here, a network resource prediction model can be built using a variety of artificial intelligence network models. This application often uses the Long Short-Term Memory (LSTM) network as an example of an artificial intelligence network model. Those skilled in the art will appreciate that building a network resource prediction model using other artificial intelligence network models also falls within the scope of this application.
[0057] A long short-term memory (LSTM) network, a deep learning framework, is built. Using the decomposed data of one or more historical intrinsic mode function (IMF) sub-signals, a drone-assisted 6G air-ground integrated network resource prediction model is constructed. This network resource prediction model can be used to predict and allocate network resources, adjusting drone positions based on the LSTM model's predictions to maximize network capacity.
[0058] S13, inputting one or more real-time intrinsic mode function sub-signals into the network resource prediction model to obtain a network resource prediction value output by the network resource prediction model, wherein the one or more real-time intrinsic mode function sub-signals are obtained based on variational mode decomposition of the real-time network signal received by the UAV, and the parameters of the network resource prediction model are adjusted according to the variability parameters in the network environment and the deep Q network reinforcement learning model.
[0059] After training and obtaining the network resource prediction model, the real-time network signal received by the drone can be subjected to variational mode decomposition to obtain one or more real-time intrinsic mode function sub-signals. These one or more real-time intrinsic mode function sub-signals are input into the network resource prediction model to obtain the network resource prediction values output by the network resource prediction model, such as the network resource prediction values for the drone bandwidth allocation, drone transmit power, drone location, and drone altitude for a future time period (e.g., 5 minutes later).
[0060] Since variable parameters in the 6G air-ground integrated network, such as user device location, business demand, and channel environment, will change in real time, if the network resource prediction values such as drone location can be adjusted dynamically in real time according to changes in variable parameters such as user device location, business demand changes, and various channel environments, the algorithm can have environmental perception capabilities, which will further improve the effect of network resource allocation.
[0061] Therefore, this application designs a deep Q network reinforcement learning model, which uses the real-time online learning capabilities of deep reinforcement learning to continuously optimize and adjust the deep learning decision-making process in real time according to changes in variable parameters such as user device location, business demand changes, and various channel environments. By adjusting the parameters of the network resource prediction model, the location of aerial base stations such as drones, as well as the allocation strategy and network resource prediction value of air-ground integrated network resources, can be dynamically adjusted in real time, giving the algorithm environmental perception capabilities. This application can autonomously make optimal decisions based on variable parameters such as the real-time location, speed, battery life, and surrounding environmental information of drones to predict network changes and adjust the location of aerial base stations such as drones in real time to cope with the impact of variable parameters such as changes in user behavior patterns, environmental interference, and the mobility of drones.
[0062] S14: Allocate network resources according to the predicted network resource values.
[0063] After obtaining the network resource prediction value output by the network resource prediction model, network resource allocation can be performed based on the network resource prediction value.
[0064] The embodiment of the present application proposes a method that can adapt to the changing parameters of the network and dynamically optimize resource allocation by comprehensively using adaptive signal data decomposition, network resource prediction models such as long short-term memory networks, and algorithm combination model iteration and optimization of deep reinforcement learning, thereby achieving efficient allocation of network resources such as spectrum, time and space, so as to maximize the overall performance of the drone-assisted 6G air-ground integrated network, while ensuring that the service quality of ground users is not affected. The embodiment of the present application is based on interdisciplinary fusion innovation and is expected to promote the research on drone-assisted wireless communication network resource allocation algorithms into a new stage. This application not only has in-depth theoretical research significance, but also has realistic practical application value. The research results of this application can provide strong support for the development of the low-altitude economy.
[0065] In one embodiment, the above step S11 includes steps S111 to S113.
[0066] S111, setting the goal of performing variational modal decomposition on the historical network signal to minimize a bandwidth constraint and a reconstruction error constraint, wherein the bandwidth constraint is used to constrain the instantaneous frequency fluctuation of each historical intrinsic mode function sub-signal to be minimized, and the reconstruction error constraint is used to constrain the superposition of each historical intrinsic mode function sub-signal to be equal to the historical network signal;
[0067] S112, initializing the number of historical intrinsic mode function sub-signals, penalty factor, convergence threshold, historical intrinsic mode function sub-signals, center frequency, maximum frequency, and Lagrange multiplier;
[0068] S113, iteratively update the historical intrinsic mode function sub-signals, center frequencies, and Lagrange multipliers and perform convergence judgment based on the minimization bandwidth constraint and the minimization reconstruction error constraint. If the residual of one or more historical intrinsic mode function sub-signals is less than the convergence threshold or the number of iterative updates reaches the maximum number of iterations, then terminate the iterative update and obtain one or more historical intrinsic mode function sub-signals corresponding to the historical network signal; otherwise, return to iterative update.
[0069] The embodiment of the present application constructs a high-precision UAV-assisted wireless communication network environment model, which comprehensively considers the location, speed, communication needs of the UAV and the distribution and service quality requirements of ground users.
[0070] The network environment model for UAV-assisted air-to-ground wireless communication differs from the propagation mode of terrestrial communication, which mainly depends on the propagation environment and transmission angle. The air-to-ground channel link can be modeled by a probabilistic path loss model, which includes both line-of-sight (LOS) and non-line-of-sight (NLOS) paths. The LOS and NLOS path losses from UAV base station n to device k are: and They can be expressed as the following equations (1) and (2).
[0071]
[0072] Among them, f c represents the carrier frequency, c represents the microwave speed, d n,k Represents the distance between the nth UAV base station and the kth device. It can be expressed as ψ n =(x n ,y n ,h n ) represents the position of the nth UAV base station, using ψ k =(x k ,y k ,h k ) represents the position of the kth device. μ LOS and μ NLOS are the additional losses of LOS and NLOS respectively.
[0073] The probability of a LOS link from the nth UAV base station to the kth device It is defined as the following calculation formula (3).
[0074]
[0075] Among them, v1 and v2 are constant parameters that depend on the UAV network, θ n,k is the transmission angle from drone n to device k, and exp(…) represents the natural exponential function.
[0076] Correspondingly, the NLOS probability is
[0077] Probabilistic path loss L from drone n to device k n,k It can be expressed as the following calculation formula (4).
[0078]
[0079] Then the 6G signal Y received by the nth drone n It can be expressed as the following calculation formula (5).
[0080]
[0081] where p n,k represents the transmission power from the nth drone to the kth device, L n,k represents the probabilistic path loss from drone n to device k, x k is the transmission signal of the kth device, represents the noise of the kth device.
[0082] The emerging 6G network in the future will adopt millimeter waves and will operate in the D band (110GHz-170GHz) and H band (220GHz-330GHz). The signal received by the drone is very complex. If the received millimeter wave original signal (i.e., the historical network signal) is directly input into the machine learning model, it will greatly increase the edge computing burden of the drone and accelerate the power consumption of the drone. Therefore, the embodiment of the present application first uses the variational mode decomposition VMD technology to adaptively decompose the received 6G network millimeter wave signal (i.e., the historical network signal). The core idea of the embodiment of the present application is to decompose the complex historical network signal into several historical intrinsic mode function (IMF) sub-signals (referred to as modes) with specific bandwidth and center frequency, and realize the adaptive decomposition of the signal by constructing a variational model and solving the optimal solution. Each mode (IMF) is an amplitude-frequency modulation signal with a limited bandwidth and oscillates around a center frequency. Its goal is to simultaneously determine all modes and their corresponding center frequencies through a variational optimization method, minimize the sum of the bandwidths of all modal components, and ensure that the sum of the decomposed modal components is equal to the original signal.
[0083] The goal of using variational mode decomposition technology in this embodiment of the application is to convert the 6G signal Y received by the nth drone in the above formula (5) into n Decomposed into K modes u k (t), satisfies the following calculation formula (6).
[0084]
[0085] Each modal component u k (t) can be expressed as an AM / FM signal, as shown in the following calculation formula (7).
[0086]
[0087] Among them, A k (t) is the amplitude function, φ k (t) is the phase function, and its instantaneous frequency is
[0088] The modal component can be converted into an analytical signal by Hilbert transform, the single-sided spectrum can be extracted, and its bandwidth can be approximated by the compact support characteristics after Gaussian smoothing. k The bandwidth of is defined as the second-order moment of its spectrum, as shown in the following calculation formula (8).
[0089]
[0090] Among them, δ(t)+j / πt is the impulse response of Hilbert transform, which is used to transform the modal component signal u k (t) is converted to an analytical signal form (i.e., the positive frequency components are retained), and j represents the imaginary part of the complex signal. The Hilbert transform converts the real signal into a complex signal in order to separate the amplitude and phase information. Represents the time derivative. The purpose of taking the derivative of the signal after Hilbert transform is to measure the instantaneous frequency fluctuation of each mode (IMF) and thus constrain the bandwidth of the mode. The derived signal is multiplied by the complex exponential term Modulate the signal to baseband (center frequency ω k Nearby), which is equivalent to extracting the signal at the center frequency ω k The instantaneous frequency fluctuates around .
[0091] The embodiment of the present application decomposes the received 6G original signal (i.e., the historical network signal) into several historical intrinsic mode function sub-signals (IMFs), each of which has a center frequency and a bandwidth. The embodiment of the present application decomposes the signal in the frequency domain by variational mode decomposition, so that the bandwidth of each mode is minimized. The optimization goal of the embodiment of the present application is to minimize the L of the above derivative. 2 Norm (i.e. Force each mode u k The energy of (t) is concentrated at the center frequency ω k The above optimization objective can be expressed as a minimization optimization problem with constraints, as shown in the following formula (9).
[0092]
[0093] In formula (9) The constraint in expressing the minimization optimization problem is that the sum of the decomposed mode functions is equal to the original mixed signal.
[0094] By introducing the Lagrange multiplier λ(t) and the quadratic penalty factor α (to control the strictness of the constraint), the above constrained problem can be transformed into an unconstrained optimization problem, as shown in the following calculation formula (10).
[0095]
[0096] in Represents the inner product term. If there is an error between the sum of the modes and the original signal The Lagrange multiplier (also known as the Lagrange multiplier) λ(t) imposes a penalty on the objective function through the inner product term <λ(t),·>, forcing the optimization process to satisfy the constraint as much as possible, that is, dynamically adjusting the mode u through λ(t) k (t), so that their sum Approximating the original 6G signal Y n (t), if the sum of the modes deviates from Y n (t), the value of the inner product term will increase, forcing the algorithm to correct the modal decomposition results.
[0097] It can be seen from formula (10) that the Lagrangian function of the modified mode decomposition proposed in the embodiment of the present application includes two parts:
[0098] 1) Minimizing bandwidth constraints is used to constrain the instantaneous frequency fluctuation of each historical intrinsic mode function sub-signal to be minimized (through time derivatives accomplish).
[0099] 2) Minimizing the reconstruction error constraint is used to constrain the superposition of all modes to be equal to the original signal (i.e., by accomplish).
[0100] In S111, the goal of performing variational modal decomposition on the historical network signal is set to minimize the bandwidth constraint and minimize the reconstruction error constraint, wherein the minimization of the bandwidth constraint is used to constrain the instantaneous frequency fluctuation of each historical intrinsic mode function sub-signal to be minimized, and the minimization of the reconstruction error constraint is used to constrain the superposition of each historical intrinsic mode function sub-signal to be equal to the historical network signal.
[0101] The embodiment of the present application sets the goal of performing variational modal decomposition on historical network signals to minimize bandwidth constraints and reconstruction error constraints, namely:
[0102]
[0103] In the first part of the calculation formula (11), the bandwidth minimization constraint is used to minimize the instantaneous frequency fluctuation of each historical intrinsic mode function sub-signal. By penalizing the bandwidth (frequency fluctuation) of the mode, the energy of each mode is forced to be concentrated. In the second part of the calculation formula (11), the reconstruction error minimization constraint is used to constrain the superposition of each historical intrinsic mode function sub-signal to be equal to the historical network signal. By penalizing the difference between the sum of the modes and the original signal, the completeness of the decomposition is ensured.
[0104] Subsequently, the embodiment of the present application adopts an iterative method to alternately update each modal component uk , center frequency ω k and the Lagrange multiplier λ.
[0105] In S112, the number of historical intrinsic mode function sub-signals, penalty factors, convergence thresholds, historical intrinsic mode function sub-signals, center frequencies, maximum frequencies, and Lagrange multipliers are initialized.
[0106] Set the number of historical intrinsic mode function sub-signals (decomposition mode number) K, penalty factor α, convergence threshold ∈, historical intrinsic mode function sub-signals Center frequency Lagrange multiplier λ 0 .
[0107] Since 6G receives signal Y n (t) is a mixed signal. In the embodiment of the present application, the mixed signal Y is first n (t) First perform a Fourier transform (FFT) to observe the number of "main peaks" in the spectrum where energy is concentrated. Each main peak corresponds to a latent mode. If the spectrum shows six independent main peaks, K is initially set to 6.
[0108] Due to 6G signal Y n (t) is the ultra-wideband signal, and we set the penalty factor α to twice the sampling frequency of the 6G system, that is, α = 2*f s (The sampling rate of 6G signals will be in the range of 20GSPS to 500GSPS, and the specific value depends on the comprehensive trade-off between hardware capabilities, signal characteristics and scenario requirements).
[0109] The convergence threshold ∈ can be set to 10 -5 .
[0110] Initialize the modal is a zero vector, as calculated in the following formula (12).
[0111]
[0112] Initialize center frequency The signal spectrum is uniformly sampled as shown in the following calculation formula (13).
[0113] ω max is the highest frequency of the signal (13)
[0114] Calculate the highest frequency of the signal ω max The signal Y can be calculated n (t) to determine the highest frequency ω max .
[0115] Initialize the Lagrange multiplier λ 0 =0.
[0116] In S113, the historical intrinsic mode function sub-signals, center frequencies, and Lagrange multipliers are iteratively updated and convergence judgment is performed based on the minimization bandwidth constraint and the minimization reconstruction error constraint. If the residual of one or more historical intrinsic mode function sub-signals is less than the convergence threshold or the number of iterative updates reaches the maximum number of iterations, the iterative update is terminated and one or more historical intrinsic mode function sub-signals corresponding to the historical network signal are obtained. Otherwise, the process returns to iterative update.
[0117] Update the historical intrinsic mode function sub-signal (mode function) u k (Frequency domain). Fix the other historical eigenmode function sub-signals and frequencies, and find the frequency domain optimal solution of a single historical eigenmode function sub-signal, as shown in the following calculation formula (14).
[0118]
[0119] in and Represents u k The Fourier transform of (t) and λ(t), Indicates the historical network signal Y received n The Fourier transform of (t) is shown in the following calculation formula (15).
[0120] Update center frequency ω k . Calculate the frequency point (peak frequency) where the energy of the current historical intrinsic mode function sub-signal is concentrated.
[0121]
[0122] Update the Lagrange multiplier (also called the Lagrange multiplier). Adjust the Lagrange multiplier according to the residual to ensure the accuracy of signal reconstruction, as shown in the following calculation formula (16).
[0123]
[0124] The parameter τ is the step size of the Lagrange multiplier update, which is used to control the update amplitude of the Lagrange multiplier λ. In the embodiment of the present application, it can be set to
[0125] Convergence judgment: the residual is less than the threshold ∈ or reaches the set maximum number of iterations, as shown in the following calculation formula (17).
[0126]
[0127] If the convergence condition of formula (17) is satisfied, the iteration ends; otherwise, the process returns to step S113 and continues iterating until the set maximum number of iterations is reached.
[0128] The embodiment of the present application uses VMD technology to decompose the received 6G mixed signal into K (such as K = 5 corresponding to 5 communication standards) independent historical intrinsic mode function sub-signals of each frequency point, and extracts the characteristic values such as frequency, amplitude and modulation mode of the different communication standard signals in the 6G mixed signal. The frequency is obtained by calculating the instantaneous frequency mean of each IMF through Hilbert transform; the amplitude is obtained by calculating the envelope mean of each IMF; the modulation mode is identified by the modulation type (QPSK, 16QAM, etc.) based on the time-frequency distribution of the IMF (such as Wigner-Ville distribution). This achieves the adaptive decomposition of the 6G mixed signal.
[0129] Some machine learning models used in the prior art directly use the received 6G mixed signal as training data to train the machine learning model when making predictions on the 6G network. Since the 6G mixed signal contains multiple standard signals such as 6G, WiFi, and drone communication, and the data volume is large, directly using the 6G mixed signal as training data to train the machine learning model increases the complexity and training time of the machine learning model training, and also increases the power consumption of the drone. The solution of this application first uses VMD technology to decompose the received 6G mixed historical network signal into independent historical intrinsic mode function sub-signals of their respective frequency points, and extracts the characteristic values such as frequency, amplitude and modulation mode of different communication standard signals in the 6G mixed historical network signal. These characteristic values are then used together with parameters such as the drone's transmission power and altitude as training data to train the machine learning model. This can reduce the training complexity and training time of the machine learning model and reduce the power consumption of the drone.
[0130] In one embodiment, the above step S12 includes steps S121 to S123.
[0131] S121 , performing data preprocessing on the data of the one or more historical intrinsic mode function sub-signals to obtain a data set of historical intrinsic mode function sub-signals after data preprocessing.
[0132] S122, constructing a training data matrix based on the data set of the historical intrinsic mode function sub-signals, constructing a time series, and dividing the data into a training set, a validation set, and a test set.
[0133] S123: construct an initial network resource prediction model, train and optimize the initial network resource prediction model, and obtain a network resource prediction model.
[0134] In existing technologies, some machine learning models face other challenges when performing 6G network predictions, in addition to the high training complexity and time required by directly using 6G mixed signals for training. First, wireless communication data is highly nonlinear and time-dependent, making it difficult for traditional machine learning models to capture the complex spatiotemporal characteristics of the data. Second, system capacity is affected by multiple factors, including device mobility, changes in service demand, and channel environment parameters. The complex relationships between these factors require more precise modeling. Finally, traditional machine learning models may face the problem of vanishing or exploding gradients, which affects the training and convergence speed of the model.
[0135] Long Short-Term Memory (LSTM) is a special type of recurrent neural network that is highly effective in processing time series data. It is designed to address the vanishing and exploding gradient problems that traditional recurrent neural networks face when processing long sequences. LSTM is an improved recurrent neural network specifically designed to process sequence data with long-term dependencies. Compared to traditional recurrent neural networks, LSTM controls the flow of information by introducing memory cells (cell states) and three gating mechanisms (input gate, forget gate, and output gate). This ensures that important information is retained for a long time while less important information is forgotten. The introduction of memory cells, input gate, forget gate, and output gate enables LSTM to effectively process long sequence data and learn and preserve important historical information, thereby improving predictions.
[0136] In the embodiment of the present application, an LSTM model is built to use the characteristic values such as the frequency, amplitude and modulation mode of the historical intrinsic mode function sub-signals of different communication modes in the 6G mixed historical network signals extracted by VMD technology, as well as the parameters such as the transmission power and altitude of the drone as training data to train the LSTM model, construct a drone-assisted 6G air-ground integrated network resource prediction model, predict and allocate network resources, and then adjust the position of the drone according to the prediction results of the LSTM model to meet the demand of maximizing network capacity.
[0137] In S121 , data preprocessing is performed on the data of the one or more historical intrinsic mode function sub-signals to obtain a data set of the historical intrinsic mode function sub-signals after data preprocessing.
[0138] First, the following data preprocessing can be performed on the characteristic values such as frequency, amplitude, and modulation mode of the historical intrinsic mode function sub-signals of different communication formats in the 6G mixed signal extracted by VMD technology.
[0139] 1) Eliminate outliers (such as samples whose eigenvalues are outside a reasonable range).
[0140] 2) Fill in missing values (linear interpolation or mean filling).
[0141] 3) Min-Max normalize the frequency and amplitude to the interval [0, 1] to avoid gradient disappearance or explosion during model training, as shown in the following calculation formula (18).
[0142]
[0143] The modulation method uses one-hot coding (such as QPSK→[1,0,0], 16QAM→[0,1,0]).
[0144] After the above data preprocessing, a data set of historical intrinsic mode function sub-signals after data preprocessing can be obtained.
[0145] In S122, a training data matrix is constructed based on the data set of the historical intrinsic mode function sub-signals, a time series is constructed, and the data is divided into a training set, a validation set, and a test set.
[0146] The data set of the historical intrinsic mode function sub-signal after the data preprocessing unit, including characteristic values such as frequency, amplitude and modulation mode, as well as parameters such as the transmission power and altitude of the UAV, can be constructed into a training data matrix according to the historical time step T The historical time step T is the modulation period of the 6G signal, K is the number of communication formats contained in the 6G mixed signal, N is the number of drones, and the training data matrix can be expressed as the following calculation formula (19).
[0147]
[0148] As an example, the first row in the above training data matrix Each element may be, for example, time 1, signal 1 frequency, signal 1 amplitude, signal 1 modulation mode, signal 2 frequency, ..., signal K modulation mode, drone 1 power, ..., drone N power, drone 1 height, ..., drone N height.
[0149] Second row Each element may be, for example, time 2, signal 1 frequency, signal 1 amplitude, signal 1 modulation mode, signal 2 frequency, ..., signal K modulation mode, drone 1 power, ..., drone N power, drone 1 height, ..., drone N height.
[0150] And so on, row T Each element may be, for example, time T, signal 1 frequency, signal 1 amplitude, signal 1 modulation mode, signal 2 frequency, ..., signal K modulation mode, drone 1 power, ..., drone N power, drone 1 height, ..., drone N height.
[0151] Define the training data time window W (6G signal is a high-frequency signal, in this solution W = 20*T, T is the modulation period of the 6G signal).
[0152] Construct the input-output pairs of LSTM as shown in the following calculation formulas (20) and (21).
[0153] Input training dataset:
[0154] Output:
[0155] The input training data set is divided into a training set (about 70%), a validation set (about 20%), and a test set (about 10%). The training set is used for model parameter learning, the validation set is used to adjust hyperparameters (such as the number of LSTM layers and learning rate), and the test set is used to evaluate the model generalization ability.
[0156] In S123, an initial network resource prediction model is constructed, and the initial network resource prediction model is trained and optimized to obtain a network resource prediction model.
[0157] (1) Take the LSTM model as an example to construct the initial network resource prediction model.
[0158] Input layer: Receives the training dataset X output by the LSTM input training data encapsulation unit, with a dimension of 20T*(3K+2N+1).
[0159] LSTM layer: Three layers of LSTM units are stacked, with 128 hidden units in each layer to capture the long-term dependencies of the time series.
[0160] Fully connected layer: maps the LSTM output to the prediction target and predicts the network resource allocation plan (such as drone transmission power, drone flight altitude, etc.). The fully connected layer uses a linear activation function and outputs continuous values.
[0161] 2) Train and optimize the initial network resource prediction model.
[0162] The LSTM model is iteratively trained using the training set, with a batch size of 64 and a training epoch of 200.
[0163] When training the LSTM model, this solution uses the mean square error (MSE) as the loss function. The calculation method of the mean square error is shown in the following calculation formula (22).
[0164]
[0165] Among them, Yi Represents the i-th output value in the calculation formula (21).
[0166] The solution of the embodiment of the present application is multi-objective optimization, which uses weighted combination to achieve the goals, as shown in the following calculation formula (23).
[0167]
[0168] In formula (23), v1, v2, and v3 represent the weights of the mean square error of the UAV allocation bandwidth, UAV transmission power, and UAV height, respectively. The initial values are all set to 1 / 3. After each round of training of the LSTM model, the performance of the model is evaluated by formula (23). The hyperparameters such as the number of LSTM layers, the number of hidden units, and the weights are adjusted by grid search or random search. The goal is to make the formula (23) The value decreases until the value of formula (23) stops decreasing, at which point training is stopped to avoid overfitting. The resulting hyperparameter combination is the final parameter, and the resulting LSTM model is considered the optimal model. To prevent long training times, the number of training rounds can be set to 200. If this number is reached, training is forced to stop.
[0169] In one embodiment, the deep Q network reinforcement learning model in the above step S13 is obtained by training based on the following steps S131 and S132.
[0170] S131, initialize the main network, target network, experience pool, and exploration rate.
[0171] S132, training is performed using a loop training method to obtain a deep Q network reinforcement learning model, and the loop training includes the following steps S1321 to S1324.
[0172] S1321, select an action, randomly select an action based on the exploration rate to promote the algorithm's exploration of the environment, or select the action with the largest current Q value;
[0173] S1322, execute the action and store the experience, execute the selected action, obtain the corresponding reward and next state, and store the current state, the selected action, the reward, and the next state into the experience pool;
[0174] S1323, training the network, sampling a batch of data from the experience pool, calculating the target Q value and loss based on the sampled data, and using the optimizer to update the parameters of the main network so that the network continuously optimizes the prediction of the Q value;
[0175] S1324, update the target network, and copy the parameters of the main network to the target network.
[0176] Since the device location, business demand and channel environment parameters in the network will change in real time, there will be a prediction lag problem caused by the rapid change of 6G signal characteristics. The embodiment of the present application can establish a deep Q network (Deep Q-Network, referred to as DQN) reinforcement learning model, and use the real-time online learning capability of deep reinforcement learning to continuously optimize and adjust the deep learning decision-making process in real time according to the device location, business demand changes and changes in various channel environment parameters. According to the variability parameters in the above network environment and the deep Q network reinforcement learning model, the parameters of the network resource prediction model are adjusted, and the position of the drone and the allocation strategy of network resources are dynamically adjusted in real time to enable the algorithm to have environmental perception capabilities. The embodiment of the present application can autonomously make the best decision based on the real-time position, speed, battery life and surrounding environment information of the drone, so that it can predict network changes and adjust the position and resource allocation strategy of the drone in real time to solve the prediction lag problem caused by the rapid change of 6G signal characteristics, and respond to changes in user behavior patterns, environmental interference and the mobility of the drone.
[0177] This embodiment of the application uses a deep Q-network (DQN) as a reinforcement learning framework, combined with the VMD-LSTM model designed above for resource prediction and allocation decisions. By defining a suitable reward function, this function evaluates different resource allocation strategies based on network performance (such as throughput, latency, user satisfaction, etc.) and resource utilization. Through continuous trial and error and feedback, the reinforcement learning mechanism guides the DQN network to gradually adjust its output strategy to maximize long-term cumulative rewards.
[0178] First, define the problem and model the environment.
[0179] Define the state space S. The state space S includes the UAV's location (latitude, longitude, and altitude), the network resource status (the UAV's allocated bandwidth and the UAV's transmit power), and the signal characteristics (frequency, amplitude, and modulation mode) of the historical intrinsic mode function sub-signals extracted through variational mode decomposition (VMD). The state space S is calculated as follows (24).
[0180] s t =[x t ,y t ,z t ,B t,1 ,…,B t,N ,P t,1 ,…,P t,N ,
[0181] f t,1 ,…,f t,K ,ψ t,1 ,…,ψ t,K ,mod t,1 ,…,mod t,K] (twenty four)
[0182] where x t ,y t ,z t They represent the three-dimensional coordinates of the UAV at time t; B t,i ,P t,i represent the bandwidth and transmission power of UAV i respectively; f t,j ,ψ t,j ,mod t,j They represent the frequency, amplitude and modulation mode of the j-th signal at time t respectively.
[0183] Define the action space A. The action space A includes the UAV's movement direction (eight directions, such as north, northeast, east, etc.) and the network resource allocation strategy (such as the UAV bandwidth allocation ratio and transmission power). The action space A is shown in the following calculation formula (25).
[0184] at=[dir t ,l t,1 ,…,l t,N ,m t,1 ,…,m t,N ] (25)
[0185] where dir t Indicates the direction of the drone's movement (8 directions are represented by discrete values 0 to 7); l t,i and m t,i Represent the bandwidth allocation weight and transmission power weight of the i-th UAV, respectively, satisfying and
[0186] Define the reward function. The reward function is used to measure the contribution of the action to the goal of maximizing network capacity, and combines factors such as throughput, interference suppression, and drone energy consumption. The reward function is shown in the following calculation formula (26).
[0187] R t =q1Throughput t -q2Interference t -q3Energy t (26)
[0188] Throughput t Indicates the total throughput of the 6G air-ground integrated network; Interference t Indicates the sum of interference intensities; Energy t Represents the drone's power consumption. q1, q2, and q3 are weight coefficients. For example, q1 = 1, q2 = 0.5, and q3 = 0.1.
[0189] Next, build a deep Q network.
[0190] The network structure of the deep Q network includes input layer, hidden layer, and output layer.
[0191] Input layer: receives state s t , whose dimension is consistent with the size of the state space.
[0192] Hidden layer: The fully connected layer or convolutional layer (when the state contains graphical features) can be selected according to the state characteristics. This patent solution is set to 3 layers, and the number of neurons in each layer is set to 128.
[0193] Output layer: Outputs the Q value of each action, with the dimension being the size of the action space |A|, thus providing a value assessment for each possible action.
[0194] The Q-value of a deep Q-network is calculated as follows. The Q-value is used to evaluate the long-term value of taking an action in a specific state. The Q-value can be calculated based on the following formula (27).
[0195] Q(s t ,a;θ)=output(FC(ReLU(FC(s t ;θ1));θ2)) (27)
[0196] Where θ1 represents the first fully connected layer FC(s t ; θ1) parameter set (including weight matrix and bias vector), θ2 represents the second fully connected layer FC(ReLU(FC(s t ; θ1)); θ2) are the set of parameters (including the weight matrix and bias vector). θ1 and θ2 are updated via backpropagation, with a learning rate typically set to 0.001 to control the step size and speed of parameter updates. FC denotes a fully connected layer. ReLU denotes the activation function (Rectified Linear Unit), where ReLU(x) = max(0,x).
[0197] The training process includes experience replay, defining the target network, defining the loss function and optimizer.
[0198] Experience replay definition: The experience replay mechanism stores the state-action-reward-next state quadruple (s t ,a t ,r t ,s t+1 ) to the experience pool D, and randomly sample for training, effectively breaking the correlation between data and improving the stability of the algorithm. The specific steps include: first perform action a t , get reward r t and the next state s t+1Next, the quadruple is stored in the experience pool D. Finally, a batch of data (e.g., 32 samples) is randomly sampled from the experience pool D for network training.
[0199] Target network definition: The target network is used to calculate the target Q value by copying the parameters of the main network and updating it periodically (e.g., every 100 steps), thereby reducing instability during training. The target network update frequency refers to the interval between target network parameter updates and can range from 100 to 1000 steps. The target network can be obtained based on the following calculation formula (28).
[0200]
[0201] Where γ is a discount factor used to balance short-term and long-term rewards, and its value range can be 0.95 to 0.99, preferably 0.97; θ - are the target network parameters.
[0202] Loss function definition: The loss function uses the mean square error (MSE) to measure the difference between the predicted Q value and the target Q value. The loss function and optimization can be obtained based on the following calculation formula (29).
[0203]
[0204] The Adam optimizer can be used to update the parameters θ1 and θ2. It is an adaptive learning rate optimization algorithm that can automatically adjust the learning rate during training. The update method of the Adam optimizer is shown in the following calculation formula (30).
[0205] m t =β1m t-1 +(1-β1)g t
[0206]
[0207]
[0208] Among them, m t and v t are the first-order moment estimate and the second-order moment estimate of the gradient respectively; β1 and β2 are exponential decay rates, β1 = 0.9, β2 = 0.999; g t is the gradient at the current moment; α is the learning rate, which is used to control the parameter update step size, and the value range can be 0.001 to 0.0001, preferably, α = 0.001; ∈ is a small constant, ∈ = 10 -8 , which is used to avoid division by zero errors. By iteratively updating parameters, the network continuously learns and optimizes its decision-making ability.
[0209] S131, initialize the main network, target network, experience pool, and exploration rate.
[0210] Initialize the main network Q(s,a;θ) and the target network Q(s,a;θ - ), let θ - =θ.
[0211] Experience pool D: The experience pool capacity refers to the buffer size for storing experience quadruplets, and the value range can be 10000 to 100000. Preferably, the initialization capacity is set to 10000.
[0212] Initialize the exploration rate ∈, the initial value can be set to 0.9. The exploration rate can be used to balance exploration and utilization. A higher exploration rate at the beginning helps the algorithm to explore different actions and states extensively.
[0213] S132, training is performed using a loop training method to obtain a deep Q network reinforcement learning model, and the loop training includes the following steps S1321 to S1324.
[0214] S1321, select an action, randomly select an action based on the exploration rate to promote the algorithm's exploration of the environment, or select the action with the largest current Q value.
[0215] Randomly select actions with exploration rate ∈ to promote the algorithm's exploration of the environment. Alternatively, select the action with the largest current Q value To leverage the currently learned knowledge. At the same time, gradually reduce ∈ (e.g., multiply by 0.99 every 100 steps). As training progresses, exploration is reduced and utilization of the learned optimal actions is increased. The exploration rate can be initially 0.9, for example, and can be gradually reduced to 0.1.
[0216] S1322, execute the action and store the experience, execute the selected action, obtain the corresponding reward, next state and store the current state, selected action, reward and next state into the experience pool.
[0217] Perform the selected actiona t , get the corresponding reward r t and the next state s t+1 , and (s t ,a t ,r t ,s t+1 ) is stored in the experience pool D to provide data for subsequent training.
[0218] S1323, training the network, sampling a batch of data from the experience pool, calculating the target Q value and loss based on the sampled data, and using the optimizer to update the parameters of the main network so that the network continuously optimizes the prediction of the Q value.
[0219] Sample a batch of data from the experience pool D and calculate the target Q value y based on the sampled data t And loss L(θ), and then use the Adam optimizer to update the main network parameters θ, so that the network continuously optimizes the prediction of Q value.
[0220] S1324, update the target network, and copy the parameters of the main network to the target network.
[0221] Every 100 steps, the main network parameters are copied to the target network, θ - =θ, keeping the target network synchronized with the main network and stabilizing the training process.
[0222] After the deep Q network reinforcement learning model training is completed, the network parameters can be fixed, and the algorithm has learned a better decision-making strategy.
[0223] In actual application, select As the optimal action, the drone path and resource allocation are adjusted in real time to achieve efficient drone operation and resource management.
[0224] The embodiment of the present application uses the DQN online deep reinforcement learning mechanism to update the weights and parameters of the LSTM model (such as the number of LSTM layers, the number of hidden units, the weights, and other parameters) in real time, and dynamically adjust the position of the drone and the allocation strategy of network resources in real time to adapt to the dynamically changing network environment. The experience replay pool is used to store the tuple of state-action-reward-new state, and samples are randomly extracted from it to update the network. Repeating LSTM and DQN, the goal is to make the calculation formula (23) The value becomes smaller and smaller until the value of formula (23) stops decreasing or the upper limit of the number of iterations is reached.
[0225] In response to the problem of large data volume in the 6G air-ground integrated network, the embodiment of the present application designs an adaptive signal decomposition technology to decompose the received 6G mixed signal into independent signals of their respective frequencies, and extract the characteristic values such as frequency, amplitude and modulation mode of the signals of different communication standards. These characteristic values are then used together with parameters such as the transmission power and height of the aerial base station as training data to train network resource prediction models such as deep learning models. This can reduce the training complexity and training time of network resource prediction models such as deep learning models, and reduce the power consumption of aerial base stations.
[0226] In response to the problem that network resource prediction models such as deep learning models may find it difficult to capture the complex spatiotemporal characteristics in the data, the embodiments of the present application propose a long short-term memory network model. By building a long short-term memory network model, the characteristic values such as frequency, amplitude and modulation mode of different communication standard signals in the 6G mixed signal extracted by adaptive signal decomposition technology, as well as parameters such as the transmission power and height of the aerial base station are used as training data to train the long short-term memory network model, construct a 6G air-ground integrated network resource prediction model, predict and allocate network resources, and then adjust the position of the aerial base station according to the prediction results to meet the needs of maximizing network capacity.
[0227] To address the problem of deep learning model prediction lag caused by the rapid changes in 6G signal characteristics, the embodiments of the present application build a deep Q network reinforcement learning model and utilize the real-time online learning capabilities of deep reinforcement learning to continuously optimize and adjust the deep learning decision-making process in real time according to device location, business demand changes, and changes in various channel environment parameters. It dynamically adjusts the allocation strategy of network resources and the location of aerial base stations in real time, which can better meet real-time requirements.
[0228] like Figure 2 As shown, the network resource allocation device 2 includes:
[0229] a decomposition module 21 for decomposing the historical network signal received by the UAV into one or more historical intrinsic mode function sub-signals based on variational mode decomposition;
[0230] A training module 22 is configured to perform training based on the data of the one or more historical intrinsic mode function sub-signals to obtain a network resource prediction model;
[0231] A prediction module 23 is configured to input one or more real-time intrinsic mode function sub-signals into the network resource prediction model to obtain a network resource prediction value output by the network resource prediction model, wherein the one or more real-time intrinsic mode function sub-signals are obtained based on variational mode decomposition of the real-time network signal received by the UAV, and the parameters of the network resource prediction model are adjusted according to the variability parameters in the network environment and the deep Q network reinforcement learning model;
[0232] The allocation module 24 is configured to allocate network resources according to the predicted network resource values.
[0233] Those skilled in the art will appreciate that the operations performed by the decomposition module 21 described above may also be implemented by a sub-device, such as a data decomposition sub-device. The operations performed by the training module 22, prediction module 23, and allocation module 24 described above may also be implemented by a sub-device, such as a network resource prediction and allocation sub-device. The operation of adjusting the parameters of the network resource prediction model based on the variability parameters in the network environment and the deep Q network reinforcement learning model may be implemented by a network resource real-time dynamic update sub-device.
[0234] Another embodiment of the present invention discloses a network resource allocation device 2. Figure 2 Based on the corresponding embodiment, the decomposition module 21 is used to:
[0235] The goal of performing variational modal decomposition on the historical network signal is set to minimize bandwidth constraints and minimize reconstruction error constraints, wherein the bandwidth minimization constraint is used to constrain the instantaneous frequency fluctuation of each historical intrinsic mode function sub-signal to be minimized, and the reconstruction error minimization constraint is used to constrain the superposition of each historical intrinsic mode function sub-signal to be equal to the historical network signal;
[0236] Initialize the number of historical intrinsic mode function sub-signals, penalty factor, convergence threshold, historical intrinsic mode function sub-signals, center frequency, maximum frequency, and Lagrange multiplier;
[0237] Iteratively update the historical intrinsic mode function sub-signals, center frequencies, and Lagrange multipliers and perform convergence judgment based on the minimized bandwidth constraint and the minimized reconstruction error constraint. If the residual of one or more historical intrinsic mode function sub-signals is less than a convergence threshold or the number of iterative updates reaches a maximum number of iterations, terminate the iterative update and obtain one or more historical intrinsic mode function sub-signals corresponding to the historical network signal. Otherwise, return to iterative update.
[0238] Another embodiment of the present invention discloses a network resource allocation device 2. Figure 2 Based on the corresponding embodiment, the training module 22 is used to:
[0239] Performing data preprocessing on the data of the one or more historical intrinsic mode function sub-signals to obtain a data set of historical intrinsic mode function sub-signals after data preprocessing;
[0240] Constructing a training data matrix based on the data set of the historical intrinsic mode function sub-signal, constructing a time series, and dividing the data into a training set, a validation set, and a test set;
[0241] Construct an initial network resource prediction model, train and optimize the initial network resource prediction model, and obtain a network resource prediction model.
[0242] Another embodiment of the present invention discloses a network resource allocation device 2. Figure 2 Based on the corresponding embodiment, the deep Q network reinforcement learning model is trained based on the following steps:
[0243] Initialize the main network, target network, experience pool, and exploration rate;
[0244] A deep Q network reinforcement learning model is obtained by training in a cyclic training manner. The cyclic training includes:
[0245] Select actions, randomly select actions based on the exploration rate to promote the algorithm's exploration of the environment, or select the action with the largest current Q value;
[0246] Execute actions and store experience: execute the selected action, obtain the corresponding reward and next state, and store the current state, selected action, reward, and next state into the experience pool;
[0247] Training the network involves sampling a batch of data from the experience pool, calculating the target Q value and loss based on the sampled data, and using the optimizer to update the parameters of the main network so that the network continuously optimizes its prediction of the Q value.
[0248] Update the target network and copy the parameters of the main network to the target network.
[0249] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0250] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0251] The present application also provides an electronic device, such as Figure 3As shown, the electronic device 3 includes: at least one processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor 30. When the processor 30 executes the computer program 32, the steps in any of the above-mentioned method embodiments are implemented.
[0252] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0253] An embodiment of the present application provides a computer program product. When the computer program product is executed, the steps in the above-mentioned method embodiments are executed.
[0254] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0255] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0256] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0257] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0258] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0259] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A network resource allocation method, characterized in that: include: Based on variational mode decomposition, the historical network signal received by the UAV is decomposed into one or more historical intrinsic mode function sub-signals; Performing training based on the data of the one or more historical intrinsic mode function sub-signals to obtain a network resource prediction model; Inputting one or more real-time intrinsic mode function sub-signals into the network resource prediction model to obtain a network resource prediction value output by the network resource prediction model, wherein the one or more real-time intrinsic mode function sub-signals are obtained based on variational mode decomposition of real-time network signals received by the UAV, and the parameters of the network resource prediction model are adjusted according to variability parameters in the network environment and a deep Q network reinforcement learning model; Network resources are allocated according to the predicted value of the network resources.
2. The method according to claim 1, wherein The variational modal decomposition method decomposes the historical network signal received by the UAV into one or more historical intrinsic mode function sub-signals, including: The goal of performing variational modal decomposition on the historical network signal is set to minimize bandwidth constraints and minimize reconstruction error constraints, wherein the bandwidth minimization constraint is used to constrain the instantaneous frequency fluctuation of each historical intrinsic mode function sub-signal to be minimized, and the reconstruction error minimization constraint is used to constrain the superposition of each historical intrinsic mode function sub-signal to be equal to the historical network signal; Initialize the number of historical intrinsic mode function sub-signals, penalty factor, convergence threshold, historical intrinsic mode function sub-signals, center frequency, maximum frequency, and Lagrange multiplier; Iteratively update the historical intrinsic mode function sub-signals, center frequencies, and Lagrange multipliers and perform convergence judgment based on the minimized bandwidth constraint and the minimized reconstruction error constraint. If the residual of one or more historical intrinsic mode function sub-signals is less than a convergence threshold or the number of iterative updates reaches a maximum number of iterations, terminate the iterative update and obtain one or more historical intrinsic mode function sub-signals corresponding to the historical network signal. Otherwise, return to iterative update.
3. The method according to claim 1, wherein The training based on the data of the one or more historical intrinsic mode function sub-signals to obtain a network resource prediction model includes: Performing data preprocessing on the data of the one or more historical intrinsic mode function sub-signals to obtain a data set of historical intrinsic mode function sub-signals after data preprocessing; Constructing a training data matrix based on the data set of the historical intrinsic mode function sub-signal, constructing a time series, and dividing the data into a training set, a validation set, and a test set; Construct an initial network resource prediction model, train and optimize the initial network resource prediction model, and obtain a network resource prediction model.
4. The method according to claim 1, wherein The deep Q network reinforcement learning model is trained based on the following steps: Initialize the main network, target network, experience pool, and exploration rate; A deep Q network reinforcement learning model is obtained by training in a cyclic training manner. The cyclic training includes: Select actions, randomly select actions based on the exploration rate to promote the algorithm's exploration of the environment, or select the action with the largest current Q value; Execute actions and store experience: execute the selected action, obtain the corresponding reward and next state, and store the current state, selected action, reward, and next state into the experience pool; Training the network involves sampling a batch of data from the experience pool, calculating the target Q value and loss based on the sampled data, and using the optimizer to update the parameters of the main network so that the network continuously optimizes its prediction of the Q value. Update the target network and copy the parameters of the main network to the target network.
5. A network resource allocation device, characterized in that: include: a decomposition device for decomposing the historical network signal received by the UAV into one or more historical intrinsic mode function sub-signals based on variational mode decomposition; A training device, configured to perform training based on data of the one or more historical intrinsic mode function sub-signals to obtain a network resource prediction model; A prediction device, configured to input one or more real-time intrinsic mode function sub-signals into the network resource prediction model to obtain a network resource prediction value output by the network resource prediction model, wherein the one or more real-time intrinsic mode function sub-signals are obtained based on variational mode decomposition of real-time network signals received by the drone, and parameters of the network resource prediction model are adjusted based on variability parameters in the network environment and a deep Q network reinforcement learning model; The allocation device is used to allocate network resources according to the predicted value of network resources.
6. The device according to claim 5, characterized in that The decomposition device is used for: The goal of performing variational modal decomposition on the historical network signal is set to minimize bandwidth constraints and minimize reconstruction error constraints, wherein the bandwidth minimization constraint is used to constrain the instantaneous frequency fluctuation of each historical intrinsic mode function sub-signal to be minimized, and the reconstruction error minimization constraint is used to constrain the superposition of each historical intrinsic mode function sub-signal to be equal to the historical network signal; Initialize the number of historical intrinsic mode function sub-signals, penalty factor, convergence threshold, historical intrinsic mode function sub-signals, center frequency, maximum frequency, and Lagrange multiplier; Iteratively update the historical intrinsic mode function sub-signals, center frequencies, and Lagrange multipliers and perform convergence judgment based on the minimized bandwidth constraint and the minimized reconstruction error constraint. If the residual of one or more historical intrinsic mode function sub-signals is less than a convergence threshold or the number of iterative updates reaches a maximum number of iterations, terminate the iterative update and obtain one or more historical intrinsic mode function sub-signals corresponding to the historical network signal. Otherwise, return to iterative update.
7. The device according to claim 5, characterized in that The training device is used for: Performing data preprocessing on the data of the one or more historical intrinsic mode function sub-signals to obtain a data set of historical intrinsic mode function sub-signals after data preprocessing; Constructing a training data matrix based on the data set of the historical intrinsic mode function sub-signal, constructing a time series, and dividing the data into a training set, a validation set, and a test set; Construct an initial network resource prediction model, train and optimize the initial network resource prediction model, and obtain a network resource prediction model.
8. The device according to claim 5, wherein The deep Q network reinforcement learning model is trained based on the following steps: Initialize the main network, target network, experience pool, and exploration rate; A deep Q network reinforcement learning model is obtained by training in a cyclic training manner. The cyclic training includes: Select actions, randomly select actions based on the exploration rate to promote the algorithm's exploration of the environment, or select the action with the largest current Q value; Execute actions and store experience: execute the selected action, obtain the corresponding reward and next state, and store the current state, selected action, reward, and next state into the experience pool; Training the network involves sampling a batch of data from the experience pool, calculating the target Q value and loss based on the sampled data, and using the optimizer to update the parameters of the main network so that the network continuously optimizes its prediction of the Q value. Update the target network and copy the parameters of the main network to the target network.
9. An electronic device, characterized in that: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the method according to any one of claims 1 to 4.
10. A computer program product, characterized in that The invention comprises a computer program, which enables the method according to any one of claims 1 to 4 to be performed when the computer program is executed.